darianrosebrook/agent-agency

Arbitration agent and task runners working together locally to complete a goal.

Rust

1

892 commits

updated Feb 16, 2026

See the code

README

Agent Agency - AI Orchestration Platform

Critical Warning: Verification Requirements

BEFORE TRUSTING ANY STATUS CLAIMS OR ACHIEVEMENT REPORTS

All system status claims must be verified against our rigorous standards. AI coding agents frequently produce overly optimistic reports about system readiness when only stubs, placeholders, or mock implementations exist.

Required Reading: docs/architecture/coreml-first-decision.md

This document provides guidelines to avoid unverified claims and maintain realistic progress assessment.

Key Principle: Never trust claims of "production-ready", "fully functional", or "complete" without independent verification.


Overview

Agent Agency V3 is an AI orchestration platform that implements constitutional governance for autonomous agent operations. The system leverages CoreML-optimized Mistral models with Apple Neural Engine acceleration for high-performance local inference, using a council of specialized AI judges to provide real-time oversight, ensuring ethical compliance, technical quality, and system coherence through evidence-based decision making.

SYSTEM STATUS: Operational - Core Orchestration Functional

The V3 system provides a functional AI orchestration platform with constitutional governance, council-based oversight, and autonomous task execution. Core capabilities are operational with ongoing integration polish and feature enhancements.

Current Capabilities:

  • Core Orchestration: Task planning, execution, and coordination operational
  • Council Governance: Four-judge constitutional oversight framework functional
  • Git Worktree Integration: Parallel worker isolation implemented
  • CAWS Compliance: Quality gates and provenance tracking operational
  • CoreML Inference: Hardware-accelerated local model execution
  • API Server: REST API for task management and monitoring

Architecture: 17-crate modular architecture with clear separation of concerns. System implements contracts-first design with zero circular dependencies.

This mono-repo contains multiple iterations examining different approaches to AI agent systems:

  • iterations/v2/: TypeScript implementation investigating multi-component agent orchestration with external service integration
  • iterations/v3/: Rust implementation with constitutional council governance, multiple execution modes, and monitoring capabilities
  • iterations/poc/: Reference implementation examining multi-tenant memory systems and federated learning concepts
  • iterations/main/: Reserved for stable research artifacts

Project Structure

agent-agency/
├── iterations/
│   ├── v2/               # TypeScript multi-component agent orchestration
│   ├── v3/               # Rust-based advanced AI capabilities
│   ├── poc/              # Multi-tenant memory systems reference
│   ├── main/             # Reserved for stable research artifacts
│   └── arbiter-poc/      # Arbiter-specific research experiments
├── docs/                 # Research documentation and findings
├── scripts/              # Shared build and utility scripts
├── apps/                 # MCP tools and utilities
├── package.json          # Mono-repo dependency management
└── tsconfig.json         # Base TypeScript configuration

Core Capabilities

Agent Agency V3 provides an AI orchestration platform with core features implemented:

AI Inference Pipeline

Implemented - Hardware-accelerated inference with safety guarantees:

  • Core ML Integration: Safe Rust wrappers with Send/Sync thread safety and async execution
  • ONNX Runtime: Cross-platform model execution with device selection and tensor validation
  • ANE Acceleration: Apple Silicon optimization with performance improvements
  • Hardware Telemetry: Safe system tool integration for thermal and power monitoring

Learning & Adaptation

Implemented - Self-improving agents with durable persistence:

  • Deep Reinforcement Learning: Neural network-based policy execution and Q-value estimation
  • Learning State Persistence: Complete durability across system restarts and failures
  • Worker Performance Evolution: Skill development tracking and continuous improvement
  • Resource Intelligence: Learned optimal allocation patterns from historical data

Observability & Analytics

Implemented - Monitoring with comprehensive insights:

  • Redis Analytics: Connection pooling, health monitoring, and trend prediction
  • CPU Utilization Tracking: Historical data analysis with volatility smoothing
  • System Health Monitoring: SLA tracking, circuit breakers, and automated alerts
  • Business Intelligence: Task throughput, error rates, and performance analytics

Evidence Pipeline

Implemented - Multi-modal verification with compliance standards:

  • Evidence Correlation: Cross-modal analysis across text, code, data, and visual modalities
  • Standards Compliance: GDPR, CCPA, HIPAA, SOC2, ISO, PCI, OWASP, WCAG verification
  • Code Quality Assurance: Unit tests, coverage analysis, linting, and integration checks
  • Claim Verification: Evidence-based validation with confidence scoring

Reliability & Resilience

Implemented - Circuit breakers, monitoring, and automated recovery:

  • Health Monitoring: CPU/memory tracking with availability SLA enforcement
  • Agent Coordination: Performance tracking and inter-agent communication
  • Failure Isolation: Automatic circuit breaker patterns for service protection
  • Recovery Automation: Self-healing capabilities with minimal downtime

Distributed Caching Infrastructure

Implemented - Multi-level caching with intelligent invalidation:

  • Erased Serde Serialization: Type-safe operations with compression/decompression
  • Tag-Based Invalidation: Redis set tracking for efficient cache management
  • SQL Query Analysis: Table dependency extraction and query complexity scoring
  • Priority Ordering: Memory-first, then Redis, then disk fallback strategies

Hardware Compatibility Layer

Implemented - iOS telemetry without unsafe FFI:

  • Thermal Monitoring: CPU, ANE, and battery temperature tracking via powermetrics
  • Power Consumption: System power estimation with detailed metrics
  • Thermal Pressure: Speed limit monitoring for thermal management
  • Fan Detection: Intelligent fan speed monitoring for equipped Macs

Constitutional Council Governance

Core Framework Implemented - Four specialized AI judges provide oversight framework:

  • Constitutional Judge: Ethical compliance and CAWS governance (framework implemented)
  • Technical Auditor: Code quality and security validation (framework implemented)
  • Quality Evaluator: Requirements satisfaction and correctness (framework implemented)
  • Integration Validator: System coherence and architectural integrity (framework implemented)
  • Multiple Execution Modes: Strict, Auto, and Dry-Run modes supported

Task Execution Pipeline

Implemented - Full orchestration with core features operational:

  • Worker Orchestration: HTTP-based task distribution with circuit breaker patterns
  • Progress Tracking: Real-time task status and comprehensive metrics collection
  • Intervention API: Pause, resume, cancel operations with full lifecycle management
  • Learning Integration: Task execution feeds into learning algorithms for continuous improvement
  • Resource Optimization: Dynamic allocation based on learned patterns and current load
  • CLI intervention commands (implemented)
  • Web dashboard with metrics (implemented)
  • SLO monitoring framework (planned)
  • Provenance tracking (basic implementation)

MCP Tool Ecosystem

Fully Implemented - Comprehensive Model Context Protocol (MCP) server with 13 specialized tools:

  • Policy Tools (3): caws_policy_validator, waiver_auditor, budget_verifier - Governance and compliance
  • Conflict Resolution Tools (3): debate_orchestrator, consensus_builder, evidence_synthesizer - Arbitration and decision-making
  • Evidence Collection Tools (3): claim_extractor, fact_verifier, source_validator - Verification and validation
  • Governance Tools (3): audit_logger, provenance_tracker, compliance_reporter - Audit trails and compliance
  • Quality Gate Tools (3): code_analyzer, test_executor, performance_validator - Code quality and testing
  • Reasoning Tools (2): logic_validator, inference_engine - Logical reasoning and probabilistic inference
  • Workflow Tools (2): progress_tracker, resource_allocator - Project management and resource optimization

All tools leverage existing systems (claim extraction, council arbitration, provenance service, quality gates, reflexive learning) and are available via standardized MCP protocol for external AI model integration.

Research Iterations

V3: AI Orchestration Platform

The V3 iteration provides an AI orchestration platform with core features implemented:

CoreML-First AI System Operational

  • CoreML Mistral: Primary model for all constitutional reasoning, judge deliberations, and orchestration tasks with ANE acceleration (2.8x speedup)
  • CoreML Acceleration: Apple Silicon optimized models including FastViT T8 F16 for vision processing with thread-safe FFI integration
  • Model Hot-Swapping: Zero-downtime model replacement with performance tracking and A/B testing
  • Self-Prompting Loops: Autonomous agent that iteratively improves outputs until quality thresholds met
  • Model Registry: Performance-weighted routing with task-specific model affinities (all critical paths → CoreML Mistral)
  • Send/Sync Safety: NEW - CoreML operations safely integrated with async Rust runtime through thread confinement and channel-based communication

CoreML Safety Architecture Implemented

  • Thread-Confinement: CoreML raw pointers isolated to dedicated threads, preventing Send/Sync violations
  • Opaque Model References: ModelRef(u64) identifiers safely cross async boundaries
  • Channel-Based Communication: Async coordination between council and inference threads using crossbeam::channel
  • Memory Safety: Proper resource cleanup and leak prevention with Drop implementations
  • FFI Boundary Control: All unsafe CoreML operations quarantined with comprehensive validation

Core Execution Loop Operational

  • Task Submission: REST API and CLI interfaces for task creation
  • Worker Orchestration: HTTP-based task distribution with circuit breaker patterns
  • Progress Tracking: Real-time task status and intervention capabilities
  • Execution Modes: Strict, Auto, and Dry-Run modes supported
  • Intervention API: Pause, resume, cancel operations implemented

Governance Framework Core Implemented

  • Constitutional Council: Four-judge framework for oversight (logic partially implemented)
  • CAWS Compliance: Runtime validation with waiver system for exceptions
  • Provenance Tracking: Basic Git integration with cryptographic signing framework
  • Quality Gates: Automated testing and validation pipelines

Monitoring & Control Partially Implemented

  • Real-time Monitoring: Task progress and basic system metrics
  • CLI Intervention: Core intervention commands implemented
  • Web Dashboard: Basic metrics display and database exploration
  • SLO Monitoring: Framework implemented, comprehensive monitoring TODO
  • Alert Management: Basic alerting, advanced features TODO

Infrastructure Partially Implemented

  • Database Layer: PostgreSQL persistence with core task storage
  • API Server: RESTful API with authentication and basic endpoints
  • Task Persistence: Task lifecycle management implemented
  • Security: Basic API key authentication implemented
  • Deployment Ready: Basic Docker setup, production deployment TODO

Advanced Features Planned/Incomplete

  • Multimodal Processing: Framework exists, CoreML/FastViT vision processing framework exists but disabled due to dependency conflicts, advanced enrichers TODO
  • Apple Silicon Optimization: CoreML infrastructure exists but real inference disabled (returns mock responses), advanced thermal management TODO
  • Distributed Processing: Single-node only, distributed features TODO
  • Advanced Analytics: Basic metrics, comprehensive analytics TODO

V2: TypeScript Multi-Component Orchestration

The V2 iteration investigates multi-component agent orchestration, examining:

  • External Service Integration: Patterns for connecting AI agents with enterprise services
  • Component Architecture: Modular design for agent capabilities and coordination
  • Quality Assurance: Automated testing and validation approaches
  • Infrastructure Management: Resource allocation and monitoring strategies

POC: Multi-Tenant Memory Systems

The POC iteration explores foundational concepts for agent memory and learning:

  • Multi-Tenant Memory: Context isolation and sharing mechanisms
  • Federated Learning: Privacy-preserving cross-agent knowledge transfer
  • MCP Integration: Model Context Protocol for agent communication
  • Reinforcement Learning: Tool optimization and adaptive behavior patterns

Research Areas

This framework investigates approaches to multimodal AI systems and constitutional governance in several areas:

  • Multimodal RAG Systems: Processing and retrieval across text, image, audio, video, and document modalities
  • Constitutional Governance: Real-time decision-making with evidence-based validation and constraint enforcement
  • Vector-Based Knowledge Systems: High-performance semantic search with pgvector and HNSW indexing
  • Production AI Deployment: Scalable, monitored, and secure deployment of multimodal AI systems
  • Cross-Modal Validation: Ensuring consistency and accuracy across different content modalities
  • Hardware-Accelerated Processing: Leveraging Apple Silicon for efficient multimodal processing and governance

V3 System Characteristics

Ideal Use Cases

The V3 system is designed for environments requiring quality assurance with local execution:

  • Development Teams: CAWS governance ensures code generation with audit trails
  • Privacy-Sensitive Organizations: Local CoreML Mistral models prevent data leakage to cloud providers
  • Apple Silicon Ecosystems: Native CoreML/ANE acceleration provides exceptional performance on Mac hardware
  • Quality-Critical Workflows: Self-prompting loops with satisficing logic prevent over-optimization
  • Cost-Conscious Development: Eliminates per-API-call costs for high-volume AI-assisted tasks

System Limitations

While powerful for its target use cases, V3 has specific constraints:

  • Local Model Constraints: CoreML Mistral provides strong reasoning capabilities with Apple Silicon optimization, though may have different training data recency than GPT-4
  • Hardware Dependencies: CoreML optimizations are Apple Silicon-specific, limiting platform portability
  • Resource Requirements: Requires powerful local machines (32GB+ RAM, M-series chips) that most developers lack
  • Cold Start Times: Model loading and initialization can take 30-60 seconds, unsuitable for interactive workflows
  • Scalability Boundaries: Cannot scale across multiple machines like cloud-based systems

Comparative Advantages

AspectV3 SystemCloud API (GPT-4)Traditional IDE Tools
PrivacyExcellentPoorGood
SafetyThread-safe CoreMLVariableGood
CostLowHigh (scale)Low
QualitySelf-improvingHigh baselineVariable
SpeedGood (local)ExcellentFast
ComplexityHighLowLow
MaintenanceHighLowLow
ScalabilityLimitedHighHigh

Technical Approaches

Constitutional Governance

The framework investigates several approaches to constitutional AI governance:

  • Judge Model Architectures: Different patterns for specialized evaluation models
  • Evidence-Based Verification: Mechanisms for validating agent outputs against constitutional requirements
  • Runtime Constraint Enforcement: Approaches to enforcing governance rules during execution
  • Learning Judge Systems: How judge models can improve through experience

Multi-Agent Coordination

Research into coordination mechanisms beyond traditional hierarchies:

  • Constitutional Concurrency: Agent coordination through agreed-upon principles
  • Evidence-Based Arbitration: Decision-making based on verifiable evidence rather than authority
  • Scalable Agent Ecosystems: Patterns for managing large numbers of coordinated agents
  • Conflict Resolution: Approaches to handling conflicting agent outputs

Implementation Strategies

Different technical approaches to constitutional governance:

  • TypeScript Orchestration: Dynamic coordination with comprehensive type safety
  • Rust Governance: Memory-safe, high-performance governance operations
  • Hardware Acceleration: Leveraging specialized hardware for governance tasks
  • Federated Learning: Privacy-preserving knowledge sharing across agent boundaries

Architecture Patterns

Constitutional Governance Patterns

The framework explores several architectural patterns for implementing constitutional governance:

  • Judge Model Networks: Networks of specialized models that evaluate different aspects of agent behavior
  • Evidence Pipelines: Multi-stage verification systems that validate agent outputs against constitutional requirements
  • Runtime Enforcement: Mechanisms for applying constitutional constraints during agent execution
  • Feedback Learning Loops: Systems where governance decisions improve through experience

Agent Coordination Models

Research into different approaches to agent coordination:

  • Constitutional Concurrency: Agents coordinate through shared constitutional principles
  • Evidence-Based Arbitration: Decision-making based on verifiable evidence and constitutional compliance
  • Hierarchical Governance: Multi-level governance with different scopes of authority
  • Distributed Consensus: Agreement protocols for constitutional decision-making

Research Methodology

The project employs progressive research through multiple implementation iterations:

  • V2 (TypeScript): Explores multi-component orchestration patterns and external service integration
  • V3 (Rust): Investigates memory safety, performance characteristics, and hardware acceleration
  • POC: Examines foundational concepts in multi-tenant memory and federated learning

Implementation Strategy

  • Mono-repo Structure: Enables comparison of different implementation approaches
  • Progressive Research: Each iteration builds on findings from previous work
  • Cross-iteration Validation: Concepts tested across different technical stacks
  • Research Documentation: Findings documented in the docs/ directory

Getting Started

Prerequisites

For V3 (Constitutional AI System):

  • Rust 1.75+
  • Docker 20.10+ and Docker Compose 2.0+
  • PostgreSQL with pgvector extension
  • Node.js 18+ (for CAWS tools and web dashboard)
  • Apple Silicon recommended for optimal performance

Installation

# Clone the repository
git clone <repository-url>
cd agent-agency

# Install Node.js dependencies (for CAWS and dashboard)
npm install

# Install CAWS Git hooks for provenance tracking
cd iterations/v3
./scripts/install-git-hooks.sh

Quick Start - V3 System

cd iterations/v3

# 1. Verify compilation (includes CoreML safety checks)
cargo check -p agent-agency-council -p agent-agency-apple-silicon
# Should show 0 errors - Send/Sync violations resolved

# 2. Start the database (optional - system has in-memory fallback)
docker run -d --name postgres-v3 -e POSTGRES_PASSWORD=password -p 5432:5432 postgres:15
docker exec -it postgres-v3 psql -U postgres -c "CREATE DATABASE agent_agency_v3;"

# 3. Run database migrations (if using PostgreSQL)
cargo run --bin migrate

# 4. Start the API server
cargo run --bin api-server &
API_PID=$!

# 5. Start the worker service (in another terminal)
cargo run --bin agent-agency-worker &
WORKER_PID=$!

# 6. Execute a task (core execution loop is operational with thread-safe CoreML)
cargo run --bin agent-agency-cli execute "Test the execution pipeline" --mode dry-run

# 7. Monitor progress via CLI
cargo run --bin agent-agency-cli intervene status <task-id>

# Cleanup when done
kill $API_PID $WORKER_PID

Note: The core task execution pipeline is operational with thread-safe CoreML integration. Send/Sync violations have been resolved through proper FFI boundary control. Many advanced features remain as TODO implementations. Use dry-run mode for safe testing without filesystem changes.

CLI Usage Examples

# Dry-run mode (safe testing)
cargo run --bin agent-agency-cli execute "Add user registration" --mode dry-run

# Auto mode with quality gates
cargo run --bin agent-agency-cli execute "Implement payment system" --mode auto --risk-tier 1

# Strict mode with manual approval
cargo run --bin agent-agency-cli execute "Deploy to production" --mode strict --watch

# CLI intervention during execution
cargo run --bin agent-agency-cli intervene pause task-123
cargo run --bin agent-agency-cli intervene resume task-123
cargo run --bin agent-agency-cli intervene cancel task-123

Web Dashboard

# Start the monitoring dashboard
cd iterations/v3/apps/web-dashboard
npm run dev

# Access at http://localhost:3000
# Features:
# - Real-time task monitoring
# - System metrics and SLOs
# - Database exploration
# - Provenance tracking
# - Alert management

Development Testing

# Build all components
cd iterations/v3
cargo build --workspace

# Run comprehensive tests
cargo test --workspace

# Run CAWS validation
cd ../../apps/tools/caws
npm run validate -- --spec-file ../../../iterations/v3/.caws/working-spec.yaml

# Test end-to-end integration
cd ../../../iterations/v3
npm run test:integration

# Run integration tests (verify all modules work together)
./scripts/run-integration-tests.sh

Infrastructure Features

  • Modular Architecture: Independent components with clear interfaces
  • Comprehensive Testing: Unit and integration tests for all modules
  • Performance Benchmarks: Automated benchmarking for optimization components
  • Security Validation: Security testing and vulnerability scanning
  • Documentation: Complete API documentation and usage examples

Documentation

V3 System Documentation

Component Documentation

Research & Reference

Author

@darianrosebrook

Contributors

darianrosebrook

892 commits

darianrosebrook/agent-agency

Arbitration agent and task runners working together locally to complete a goal.

Rust

1

892 commits

updated Feb 16, 2026

See the code

README

Agent Agency - AI Orchestration Platform

Critical Warning: Verification Requirements

BEFORE TRUSTING ANY STATUS CLAIMS OR ACHIEVEMENT REPORTS

All system status claims must be verified against our rigorous standards. AI coding agents frequently produce overly optimistic reports about system readiness when only stubs, placeholders, or mock implementations exist.

Required Reading: docs/architecture/coreml-first-decision.md

This document provides guidelines to avoid unverified claims and maintain realistic progress assessment.

Key Principle: Never trust claims of "production-ready", "fully functional", or "complete" without independent verification.


Overview

Agent Agency V3 is an AI orchestration platform that implements constitutional governance for autonomous agent operations. The system leverages CoreML-optimized Mistral models with Apple Neural Engine acceleration for high-performance local inference, using a council of specialized AI judges to provide real-time oversight, ensuring ethical compliance, technical quality, and system coherence through evidence-based decision making.

SYSTEM STATUS: Operational - Core Orchestration Functional

The V3 system provides a functional AI orchestration platform with constitutional governance, council-based oversight, and autonomous task execution. Core capabilities are operational with ongoing integration polish and feature enhancements.

Current Capabilities:

  • Core Orchestration: Task planning, execution, and coordination operational
  • Council Governance: Four-judge constitutional oversight framework functional
  • Git Worktree Integration: Parallel worker isolation implemented
  • CAWS Compliance: Quality gates and provenance tracking operational
  • CoreML Inference: Hardware-accelerated local model execution
  • API Server: REST API for task management and monitoring

Architecture: 17-crate modular architecture with clear separation of concerns. System implements contracts-first design with zero circular dependencies.

This mono-repo contains multiple iterations examining different approaches to AI agent systems:

  • iterations/v2/: TypeScript implementation investigating multi-component agent orchestration with external service integration
  • iterations/v3/: Rust implementation with constitutional council governance, multiple execution modes, and monitoring capabilities
  • iterations/poc/: Reference implementation examining multi-tenant memory systems and federated learning concepts
  • iterations/main/: Reserved for stable research artifacts

Project Structure

agent-agency/
├── iterations/
│   ├── v2/               # TypeScript multi-component agent orchestration
│   ├── v3/               # Rust-based advanced AI capabilities
│   ├── poc/              # Multi-tenant memory systems reference
│   ├── main/             # Reserved for stable research artifacts
│   └── arbiter-poc/      # Arbiter-specific research experiments
├── docs/                 # Research documentation and findings
├── scripts/              # Shared build and utility scripts
├── apps/                 # MCP tools and utilities
├── package.json          # Mono-repo dependency management
└── tsconfig.json         # Base TypeScript configuration

Core Capabilities

Agent Agency V3 provides an AI orchestration platform with core features implemented:

AI Inference Pipeline

Implemented - Hardware-accelerated inference with safety guarantees:

  • Core ML Integration: Safe Rust wrappers with Send/Sync thread safety and async execution
  • ONNX Runtime: Cross-platform model execution with device selection and tensor validation
  • ANE Acceleration: Apple Silicon optimization with performance improvements
  • Hardware Telemetry: Safe system tool integration for thermal and power monitoring

Learning & Adaptation

Implemented - Self-improving agents with durable persistence:

  • Deep Reinforcement Learning: Neural network-based policy execution and Q-value estimation
  • Learning State Persistence: Complete durability across system restarts and failures
  • Worker Performance Evolution: Skill development tracking and continuous improvement
  • Resource Intelligence: Learned optimal allocation patterns from historical data

Observability & Analytics

Implemented - Monitoring with comprehensive insights:

  • Redis Analytics: Connection pooling, health monitoring, and trend prediction
  • CPU Utilization Tracking: Historical data analysis with volatility smoothing
  • System Health Monitoring: SLA tracking, circuit breakers, and automated alerts
  • Business Intelligence: Task throughput, error rates, and performance analytics

Evidence Pipeline

Implemented - Multi-modal verification with compliance standards:

  • Evidence Correlation: Cross-modal analysis across text, code, data, and visual modalities
  • Standards Compliance: GDPR, CCPA, HIPAA, SOC2, ISO, PCI, OWASP, WCAG verification
  • Code Quality Assurance: Unit tests, coverage analysis, linting, and integration checks
  • Claim Verification: Evidence-based validation with confidence scoring

Reliability & Resilience

Implemented - Circuit breakers, monitoring, and automated recovery:

  • Health Monitoring: CPU/memory tracking with availability SLA enforcement
  • Agent Coordination: Performance tracking and inter-agent communication
  • Failure Isolation: Automatic circuit breaker patterns for service protection
  • Recovery Automation: Self-healing capabilities with minimal downtime

Distributed Caching Infrastructure

Implemented - Multi-level caching with intelligent invalidation:

  • Erased Serde Serialization: Type-safe operations with compression/decompression
  • Tag-Based Invalidation: Redis set tracking for efficient cache management
  • SQL Query Analysis: Table dependency extraction and query complexity scoring
  • Priority Ordering: Memory-first, then Redis, then disk fallback strategies

Hardware Compatibility Layer

Implemented - iOS telemetry without unsafe FFI:

  • Thermal Monitoring: CPU, ANE, and battery temperature tracking via powermetrics
  • Power Consumption: System power estimation with detailed metrics
  • Thermal Pressure: Speed limit monitoring for thermal management
  • Fan Detection: Intelligent fan speed monitoring for equipped Macs

Constitutional Council Governance

Core Framework Implemented - Four specialized AI judges provide oversight framework:

  • Constitutional Judge: Ethical compliance and CAWS governance (framework implemented)
  • Technical Auditor: Code quality and security validation (framework implemented)
  • Quality Evaluator: Requirements satisfaction and correctness (framework implemented)
  • Integration Validator: System coherence and architectural integrity (framework implemented)
  • Multiple Execution Modes: Strict, Auto, and Dry-Run modes supported

Task Execution Pipeline

Implemented - Full orchestration with core features operational:

  • Worker Orchestration: HTTP-based task distribution with circuit breaker patterns
  • Progress Tracking: Real-time task status and comprehensive metrics collection
  • Intervention API: Pause, resume, cancel operations with full lifecycle management
  • Learning Integration: Task execution feeds into learning algorithms for continuous improvement
  • Resource Optimization: Dynamic allocation based on learned patterns and current load
  • CLI intervention commands (implemented)
  • Web dashboard with metrics (implemented)
  • SLO monitoring framework (planned)
  • Provenance tracking (basic implementation)

MCP Tool Ecosystem

Fully Implemented - Comprehensive Model Context Protocol (MCP) server with 13 specialized tools:

  • Policy Tools (3): caws_policy_validator, waiver_auditor, budget_verifier - Governance and compliance
  • Conflict Resolution Tools (3): debate_orchestrator, consensus_builder, evidence_synthesizer - Arbitration and decision-making
  • Evidence Collection Tools (3): claim_extractor, fact_verifier, source_validator - Verification and validation
  • Governance Tools (3): audit_logger, provenance_tracker, compliance_reporter - Audit trails and compliance
  • Quality Gate Tools (3): code_analyzer, test_executor, performance_validator - Code quality and testing
  • Reasoning Tools (2): logic_validator, inference_engine - Logical reasoning and probabilistic inference
  • Workflow Tools (2): progress_tracker, resource_allocator - Project management and resource optimization

All tools leverage existing systems (claim extraction, council arbitration, provenance service, quality gates, reflexive learning) and are available via standardized MCP protocol for external AI model integration.

Research Iterations

V3: AI Orchestration Platform

The V3 iteration provides an AI orchestration platform with core features implemented:

CoreML-First AI System Operational

  • CoreML Mistral: Primary model for all constitutional reasoning, judge deliberations, and orchestration tasks with ANE acceleration (2.8x speedup)
  • CoreML Acceleration: Apple Silicon optimized models including FastViT T8 F16 for vision processing with thread-safe FFI integration
  • Model Hot-Swapping: Zero-downtime model replacement with performance tracking and A/B testing
  • Self-Prompting Loops: Autonomous agent that iteratively improves outputs until quality thresholds met
  • Model Registry: Performance-weighted routing with task-specific model affinities (all critical paths → CoreML Mistral)
  • Send/Sync Safety: NEW - CoreML operations safely integrated with async Rust runtime through thread confinement and channel-based communication

CoreML Safety Architecture Implemented

  • Thread-Confinement: CoreML raw pointers isolated to dedicated threads, preventing Send/Sync violations
  • Opaque Model References: ModelRef(u64) identifiers safely cross async boundaries
  • Channel-Based Communication: Async coordination between council and inference threads using crossbeam::channel
  • Memory Safety: Proper resource cleanup and leak prevention with Drop implementations
  • FFI Boundary Control: All unsafe CoreML operations quarantined with comprehensive validation

Core Execution Loop Operational

  • Task Submission: REST API and CLI interfaces for task creation
  • Worker Orchestration: HTTP-based task distribution with circuit breaker patterns
  • Progress Tracking: Real-time task status and intervention capabilities
  • Execution Modes: Strict, Auto, and Dry-Run modes supported
  • Intervention API: Pause, resume, cancel operations implemented

Governance Framework Core Implemented

  • Constitutional Council: Four-judge framework for oversight (logic partially implemented)
  • CAWS Compliance: Runtime validation with waiver system for exceptions
  • Provenance Tracking: Basic Git integration with cryptographic signing framework
  • Quality Gates: Automated testing and validation pipelines

Monitoring & Control Partially Implemented

  • Real-time Monitoring: Task progress and basic system metrics
  • CLI Intervention: Core intervention commands implemented
  • Web Dashboard: Basic metrics display and database exploration
  • SLO Monitoring: Framework implemented, comprehensive monitoring TODO
  • Alert Management: Basic alerting, advanced features TODO

Infrastructure Partially Implemented

  • Database Layer: PostgreSQL persistence with core task storage
  • API Server: RESTful API with authentication and basic endpoints
  • Task Persistence: Task lifecycle management implemented
  • Security: Basic API key authentication implemented
  • Deployment Ready: Basic Docker setup, production deployment TODO

Advanced Features Planned/Incomplete

  • Multimodal Processing: Framework exists, CoreML/FastViT vision processing framework exists but disabled due to dependency conflicts, advanced enrichers TODO
  • Apple Silicon Optimization: CoreML infrastructure exists but real inference disabled (returns mock responses), advanced thermal management TODO
  • Distributed Processing: Single-node only, distributed features TODO
  • Advanced Analytics: Basic metrics, comprehensive analytics TODO

V2: TypeScript Multi-Component Orchestration

The V2 iteration investigates multi-component agent orchestration, examining:

  • External Service Integration: Patterns for connecting AI agents with enterprise services
  • Component Architecture: Modular design for agent capabilities and coordination
  • Quality Assurance: Automated testing and validation approaches
  • Infrastructure Management: Resource allocation and monitoring strategies

POC: Multi-Tenant Memory Systems

The POC iteration explores foundational concepts for agent memory and learning:

  • Multi-Tenant Memory: Context isolation and sharing mechanisms
  • Federated Learning: Privacy-preserving cross-agent knowledge transfer
  • MCP Integration: Model Context Protocol for agent communication
  • Reinforcement Learning: Tool optimization and adaptive behavior patterns

Research Areas

This framework investigates approaches to multimodal AI systems and constitutional governance in several areas:

  • Multimodal RAG Systems: Processing and retrieval across text, image, audio, video, and document modalities
  • Constitutional Governance: Real-time decision-making with evidence-based validation and constraint enforcement
  • Vector-Based Knowledge Systems: High-performance semantic search with pgvector and HNSW indexing
  • Production AI Deployment: Scalable, monitored, and secure deployment of multimodal AI systems
  • Cross-Modal Validation: Ensuring consistency and accuracy across different content modalities
  • Hardware-Accelerated Processing: Leveraging Apple Silicon for efficient multimodal processing and governance

V3 System Characteristics

Ideal Use Cases

The V3 system is designed for environments requiring quality assurance with local execution:

  • Development Teams: CAWS governance ensures code generation with audit trails
  • Privacy-Sensitive Organizations: Local CoreML Mistral models prevent data leakage to cloud providers
  • Apple Silicon Ecosystems: Native CoreML/ANE acceleration provides exceptional performance on Mac hardware
  • Quality-Critical Workflows: Self-prompting loops with satisficing logic prevent over-optimization
  • Cost-Conscious Development: Eliminates per-API-call costs for high-volume AI-assisted tasks

System Limitations

While powerful for its target use cases, V3 has specific constraints:

  • Local Model Constraints: CoreML Mistral provides strong reasoning capabilities with Apple Silicon optimization, though may have different training data recency than GPT-4
  • Hardware Dependencies: CoreML optimizations are Apple Silicon-specific, limiting platform portability
  • Resource Requirements: Requires powerful local machines (32GB+ RAM, M-series chips) that most developers lack
  • Cold Start Times: Model loading and initialization can take 30-60 seconds, unsuitable for interactive workflows
  • Scalability Boundaries: Cannot scale across multiple machines like cloud-based systems

Comparative Advantages

AspectV3 SystemCloud API (GPT-4)Traditional IDE Tools
PrivacyExcellentPoorGood
SafetyThread-safe CoreMLVariableGood
CostLowHigh (scale)Low
QualitySelf-improvingHigh baselineVariable
SpeedGood (local)ExcellentFast
ComplexityHighLowLow
MaintenanceHighLowLow
ScalabilityLimitedHighHigh

Technical Approaches

Constitutional Governance

The framework investigates several approaches to constitutional AI governance:

  • Judge Model Architectures: Different patterns for specialized evaluation models
  • Evidence-Based Verification: Mechanisms for validating agent outputs against constitutional requirements
  • Runtime Constraint Enforcement: Approaches to enforcing governance rules during execution
  • Learning Judge Systems: How judge models can improve through experience

Multi-Agent Coordination

Research into coordination mechanisms beyond traditional hierarchies:

  • Constitutional Concurrency: Agent coordination through agreed-upon principles
  • Evidence-Based Arbitration: Decision-making based on verifiable evidence rather than authority
  • Scalable Agent Ecosystems: Patterns for managing large numbers of coordinated agents
  • Conflict Resolution: Approaches to handling conflicting agent outputs

Implementation Strategies

Different technical approaches to constitutional governance:

  • TypeScript Orchestration: Dynamic coordination with comprehensive type safety
  • Rust Governance: Memory-safe, high-performance governance operations
  • Hardware Acceleration: Leveraging specialized hardware for governance tasks
  • Federated Learning: Privacy-preserving knowledge sharing across agent boundaries

Architecture Patterns

Constitutional Governance Patterns

The framework explores several architectural patterns for implementing constitutional governance:

  • Judge Model Networks: Networks of specialized models that evaluate different aspects of agent behavior
  • Evidence Pipelines: Multi-stage verification systems that validate agent outputs against constitutional requirements
  • Runtime Enforcement: Mechanisms for applying constitutional constraints during agent execution
  • Feedback Learning Loops: Systems where governance decisions improve through experience

Agent Coordination Models

Research into different approaches to agent coordination:

  • Constitutional Concurrency: Agents coordinate through shared constitutional principles
  • Evidence-Based Arbitration: Decision-making based on verifiable evidence and constitutional compliance
  • Hierarchical Governance: Multi-level governance with different scopes of authority
  • Distributed Consensus: Agreement protocols for constitutional decision-making

Research Methodology

The project employs progressive research through multiple implementation iterations:

  • V2 (TypeScript): Explores multi-component orchestration patterns and external service integration
  • V3 (Rust): Investigates memory safety, performance characteristics, and hardware acceleration
  • POC: Examines foundational concepts in multi-tenant memory and federated learning

Implementation Strategy

  • Mono-repo Structure: Enables comparison of different implementation approaches
  • Progressive Research: Each iteration builds on findings from previous work
  • Cross-iteration Validation: Concepts tested across different technical stacks
  • Research Documentation: Findings documented in the docs/ directory

Getting Started

Prerequisites

For V3 (Constitutional AI System):

  • Rust 1.75+
  • Docker 20.10+ and Docker Compose 2.0+
  • PostgreSQL with pgvector extension
  • Node.js 18+ (for CAWS tools and web dashboard)
  • Apple Silicon recommended for optimal performance

Installation

# Clone the repository
git clone <repository-url>
cd agent-agency

# Install Node.js dependencies (for CAWS and dashboard)
npm install

# Install CAWS Git hooks for provenance tracking
cd iterations/v3
./scripts/install-git-hooks.sh

Quick Start - V3 System

cd iterations/v3

# 1. Verify compilation (includes CoreML safety checks)
cargo check -p agent-agency-council -p agent-agency-apple-silicon
# Should show 0 errors - Send/Sync violations resolved

# 2. Start the database (optional - system has in-memory fallback)
docker run -d --name postgres-v3 -e POSTGRES_PASSWORD=password -p 5432:5432 postgres:15
docker exec -it postgres-v3 psql -U postgres -c "CREATE DATABASE agent_agency_v3;"

# 3. Run database migrations (if using PostgreSQL)
cargo run --bin migrate

# 4. Start the API server
cargo run --bin api-server &
API_PID=$!

# 5. Start the worker service (in another terminal)
cargo run --bin agent-agency-worker &
WORKER_PID=$!

# 6. Execute a task (core execution loop is operational with thread-safe CoreML)
cargo run --bin agent-agency-cli execute "Test the execution pipeline" --mode dry-run

# 7. Monitor progress via CLI
cargo run --bin agent-agency-cli intervene status <task-id>

# Cleanup when done
kill $API_PID $WORKER_PID

Note: The core task execution pipeline is operational with thread-safe CoreML integration. Send/Sync violations have been resolved through proper FFI boundary control. Many advanced features remain as TODO implementations. Use dry-run mode for safe testing without filesystem changes.

CLI Usage Examples

# Dry-run mode (safe testing)
cargo run --bin agent-agency-cli execute "Add user registration" --mode dry-run

# Auto mode with quality gates
cargo run --bin agent-agency-cli execute "Implement payment system" --mode auto --risk-tier 1

# Strict mode with manual approval
cargo run --bin agent-agency-cli execute "Deploy to production" --mode strict --watch

# CLI intervention during execution
cargo run --bin agent-agency-cli intervene pause task-123
cargo run --bin agent-agency-cli intervene resume task-123
cargo run --bin agent-agency-cli intervene cancel task-123

Web Dashboard

# Start the monitoring dashboard
cd iterations/v3/apps/web-dashboard
npm run dev

# Access at http://localhost:3000
# Features:
# - Real-time task monitoring
# - System metrics and SLOs
# - Database exploration
# - Provenance tracking
# - Alert management

Development Testing

# Build all components
cd iterations/v3
cargo build --workspace

# Run comprehensive tests
cargo test --workspace

# Run CAWS validation
cd ../../apps/tools/caws
npm run validate -- --spec-file ../../../iterations/v3/.caws/working-spec.yaml

# Test end-to-end integration
cd ../../../iterations/v3
npm run test:integration

# Run integration tests (verify all modules work together)
./scripts/run-integration-tests.sh

Infrastructure Features

  • Modular Architecture: Independent components with clear interfaces
  • Comprehensive Testing: Unit and integration tests for all modules
  • Performance Benchmarks: Automated benchmarking for optimization components
  • Security Validation: Security testing and vulnerability scanning
  • Documentation: Complete API documentation and usage examples

Documentation

V3 System Documentation

Component Documentation

Research & Reference

Author

@darianrosebrook

Contributors

darianrosebrook

892 commits

Languages

Rust

50.4%

TypeScript

36.3%

JavaScript

4.1%

Python

2.3%

SCSS

1.8%

Shell

1.7%

PLpgSQL

1.6%

Swift

1.3%