kapitan00000978-sketch/Universal-Agent-HP

Universal Agent HP: 10-layer autonomous AI agent OS with DAG planning, zero-loss rollback, and multi-model support.

Python

1

0 commits

updated Sep 23, 2026

See the code

See what people are saying

README

⚡ UNIVERSAL AGENT HP — Autonomous Cognitive AI Operating System

The First 10-Layer Autonomous AI Software Company in Your Terminal — 100% Free, Zero-Key, and Air-Gapped Private.

CI - Test Suite Python Version Architecture Models License: MIT

💡 Stop paying $500/month for cloud coding agents that hallucinate, leak private corporate code, and get trapped in infinite loops.
Universal Agent HP replaces brittle single-prompt bots with an enterprise-grade autonomous software organization: a CEO Meta-Orchestrator, 4 Department Leads, and 27 Worker Specialists executing concurrent DAG dependency waves with SHA-256 zero-loss rollback and continuous self-improvement. Run frontier intelligence (DeepSeek V4 Pro, Claude 3.5 Sonnet, GPT-4o) completely free with zero API keys, or 100% offline via local Apple MLX & Ollama.


🔥 Why Developers & Teams Choose Universal Agent HP:

  • 🚀 100% Free Forever, Zero API Keys Required: Instant out-of-the-box access to 500+ frontier models (DeepSeek-R1, Claude 3.5, GPT-4o) via Puter.js, or run completely offline with Laya MLX / Ollama.
  • 🏢 An Entire Tech Company in One CLI: Not a toy prompt-wrapper. Universal Agent HP routes tasks through a CEO, 4 Team Leads (CTO, Chief Scientist, DevOps, QA), and 27 isolated specialists.
  • Kahn's DAG Wave Execution: No more slow sequential loops. Independent tasks run concurrently in parallel waves; failing branches replan dynamically without losing completed work.
  • 🛡️ Zero-Loss Filesystem Rollback: Real byte-level SHA-256 snapshots before destructive actions. One command restores your codebase to its exact original state if an agent makes a mistake.
  • 🧠 AST Causal Intelligence & Blast-Radius: Calculates caller/callee ripples before editing code so unexpected regression bugs are eliminated before they happen.
  • 📈 Continuous Self-Improvement (Level 10): Diagnoses drift, turns failed test runs into permanent playbooks, and self-tunes its reasoning prompts autonomously.

🏛️ Genesis Cognitive Architecture (Levels 1 – 10)

========================================================================================
LEVEL 1 — META-ORCHESTRATOR (Chief Executive Agent)
  ├── Global Goal Memory (Maintains multi-session project roadmap & long-term objectives)
  ├── Resource & Token Budget Controller (Manages token consumption & USD spend)
  └── Multi-Agent Conflict Resolver (Resolves departmental priorities and constraints)
========================================================================================
                                     │
         ┌───────────────────────────┼───────────────────────────┐
         ▼                           ▼                           ▼
LEVEL 2 — 4 DEPARTMENT TEAM LEADS
  ├── EngineeringLead (CTO)           : System architecture, code synthesis, API schemas
  ├── ResearchLead (Chief Scientist) : Multi-hop internet research, doc audits, vector RAG
  ├── OperationsLead (DevOps / SRE)  : Terminal actions, Docker sandboxes, Git migrations
  └── QualitySecurityLead (Audit/QA) : AST security scanning, test suites, zero-loss rollback
========================================================================================
                                     │
         ┌───────────────────────────┴───────────────────────────┐
         ▼                                                       ▼
LEVEL 3 — 27 SPECIALIST WORKER ROLES
  ├── Backend Specialist              ├── Database Specialist         ├── Security Auditor
  ├── Frontend Specialist             ├── Refactoring Specialist      ├── Performance Engineer
  ├── Vector RAG Specialist           ├── Test Engineer               ├── Browser Automation Lead
  └── (18 Additional Specialized Roles dispatched dynamically per task context)
========================================================================================
                                     │
LEVEL 4 TO 10 — COGNITIVE SUBSYSTEMS & REASONING RUNTIMES
  ├── Level 4: Directed Acyclic Graph (DAG) Task Planner & Wave-Based Parallel Executor
  ├── Level 5: Dual Reasoning Loops: Reflexion Engine & Multi-Agent Debate Arena
  ├── Level 6: Causal Knowledge Graph Memory & AST Blast-Radius Impact Analyzer
  ├── Level 7: Dynamic Tool Discovery, Sandboxing & Bayesian EWMA Reliability Rating
  ├── Level 8: Capability-Based Model Routing & Cognitive USD Budget Tracker
  ├── Level 9: Execution Sandbox with Filesystem Snapshot & Zero-Loss Rollback
  └── Level 10: Drift Detection & Autonomous Self-Improvement Benchmark Suite
========================================================================================

🤖 Supported Models (Zero-Key & Out-of-the-Box)

Universal Agent HP requires zero paid subscriptions or mandatory API keys. It natively interfaces with both cutting-edge frontier cloud models and air-gapped local runtimes:

🌐 Frontier Cloud Models (Keyless & Instant)

  • DeepSeek V4 Pro & DeepSeek-R1 (Full 671B MoE deep reasoning)
  • Claude 3.5 Sonnet & Claude 3.5 Haiku (Anthropic frontier coding)
  • Claude Opus 3 / 4.5 (Complex system architecture synthesis)
  • GPT-4o & GPT-4o Mini (Multimodal analysis & fast execution)
  • GPT-5 Class Models (Next-generation autonomous reasoning)
  • Llama 3.3 (70B) & Llama 3.1 (405B) (Open-weights scale intelligence)
  • Qwen 2.5 (72B) & Qwen 2.5 Coder (State-of-the-art polyglot code generation)
  • Mistral Large 2 & Codestral (High-precision technical execution)
  • Grok Beta / Grok 2 (Real-time data-driven synthesis) Accessed instantly via Puter.js, Completions.me gateway, GPT4Free (g4f), and Python-tGPT (tgpt).

💻 Local & Air-Gapped Models (100% Offline & Private)

  • Laya MLX: Hardware-accelerated local Apple Silicon & PC engine (Apple Unified Memory / Metal optimized).
  • Ollama: Native execution for deepseek-r1:8b/14b/32b, llama3.2, qwen2.5-coder, mistral, phi-4, and codellama.
  • LM Studio: Any open-weights GGUF quantized model running on localhost:1234.
  • OmniRoute: Smart auto-routing load balancer across local nodes.

💡 Tip: Leave all API keys blank in .env and run python cli.py --provider puter --model deepseek/deepseek-v4-pro or --provider ollama --model deepseek-r1:8b to run completely free!


🏆 Architectural Superiority: Why Universal Agent HP Outperforms Other AI Agents

Unlike standard wrapper bots, single-prompt LLMs, or brittle ReAct loops, Universal Agent HP is engineered as an enterprise-grade autonomous software company. Here is how it compares directly against industry alternatives:

1. Universal Agent HP vs. Hermes 3 & Single-Model Agents

  • Context Degradation & Hallucination: Hermes 3 and typical fine-tuned models operate within a single context window. As task steps multiply, prompt contamination causes them to forget instructions, hallucinate tool signatures, or lose track of files.
  • Universal Agent HP's Advantage: Universal Agent HP uses a 10-Layer Cognitive Hierarchy. The CEO Meta-Orchestrator divides high-level intent across 4 Departmental Team Leads, which dispatch work to 27 isolated Specialist Workers (staff.py). Each worker runs with dedicated, scoped tool bundles—eliminating context bloat and hallucination.

2. Universal Agent HP vs. Devin & Proprietary Cloud Agents

  • Privacy & Vendor Lock-in: Devin is a closed-source, cloud-only SaaS that requires streaming all proprietary corporate code to external third-party servers at high subscription costs.
  • Universal Agent HP's Advantage: Universal Agent HP is 100% open-source, local-first, and air-gapped. It runs entirely on your own machine using Laya MLX or Ollama with zero data leaving your perimeter.
  • Safety & Reversibility: If a cloud agent breaks your repository, manual Git recovery is tedious. Universal Agent HP incorporates a Level 9 Execution Sandbox with Filesystem Snapshots (core/sandbox/environment.py). It computes SHA-256 pre-execution hashes and provides one-click zero-loss transactional rollback, automatically reverting altered code and removing rogue files.

3. Universal Agent HP vs. AutoGPT, CrewAI & LangChain ReAct Frameworks

  • Infinite Loops & Tool Thrashing: Traditional ReAct frameworks run naive sequential loops (thought -> action -> observation). When a tool fails or throws an unhandled exception, they repeatedly hammer the same broken command until tokens or limits are exhausted.
  • Universal Agent HP's Advantage:
    • DAG Wave Planner (core/dag/): Replaces linear step-by-step loops with Kahn’s Directed Acyclic Graph topology. Independent subtasks execute concurrently in parallel waves, and failures trigger selective branch replanning rather than full workflow restarts.
    • Bayesian EWMA Tool Reliability Rating (core/tools/reliability.py): Tracks real-time tool performance dynamically from Grade A to F. If a tool degrades, the agent automatically applies mitigation advice or reroutes to alternative tool paths.
    • Dual Reasoning Engines (core/reasoning/): Combines a Reflexion Engine (self-evaluating against strict test suites before committing) and a Multi-Agent Debate Arena (pitching an Advocate against a Skeptic judged by an Arbitrator) to eliminate premature conclusions.

4. Continuous Self-Improvement & Causal AST Code Intelligence

  • Static vs. Evolving Intelligence: Standard agents are static; they make the exact same mistakes in subsequent runs.
  • Universal Agent HP's Advantage:
    • Causal Knowledge Graph & AST Blast Radius (core/memory/): Statically analyzes your workspace AST to map class/function callers and computes the ripple blast-radius before making edits, preventing hidden regressions.
    • Level 10 Drift Detection & Self-Improvement (core/monitoring/, core/self_improvement/): Continuously computes longitudinal drift across runs. When regression patterns emerge, Universal Agent HP crystallizes lessons learned into permanent reusable playbooks and auto-tunes agent strategies.

🌟 Key Technical Innovations & Specifications

1. 🏛️ Hierarchical Delegation & Role Specialization (Level 1 & 2)

  • CEO Meta-Orchestrator (orchestrator.py): Persists multi-turn session state, allocates computing budgets, and arbitrates competing agent directives.
  • 4 Department Team Leads (team_leads.py): Each departmental lead performs pre-flight goal decomposition, dispatches worker sub-tasks, and validates output quality before reporting up the hierarchy.
  • 27 Specialist Workers (staff.py): Dedicated operational personas with granular tool permissions, eliminating context contamination.

2. 📊 Directed Acyclic Graph (DAG) Task Planner & Wave Executor (Level 4)

  • Topological Wave Execution (core/dag/): Deconstructs composite goals into dependency graphs using Kahn's algorithm. Independent nodes run concurrently in parallel execution waves (max_concurrency=4).
  • Selective Replanning: When a node fails, the planner isolates the failed branch and only replans affected downstream nodes, preserving the work of successful independent tasks.

3. 🧠 Dual Cognitive Reasoning Engines (Level 5)

  • Reflexion Loop (core/reasoning/reflexion.py): A continuous self-critique loop. The agent evaluates its candidate solutions against strict success criteria, iteratively revising code and hypotheses up to 3 cycles.
  • Multi-Agent Debate Arena (core/reasoning/debate.py): Pitches an Advocate (proposing architecture and solutions) against a Skeptic (uncovering edge-cases, race conditions, and attack vectors). An Arbitrator Judge synthesizes the winning consensus.

4. 🌐 Causal Knowledge Graph & AST Code Intel (Level 6)

  • Workspace AST Extractor (core/memory/ast_graph_extractor.py): Statically parses the entire Python workspace, constructing an automated graph of classes, functions, imports, and call dependencies.
  • Blast Radius & Impact Analysis: Before modifying any function or file, Titan computes affected downstream callers and modules (kg_impact_analysis), preventing unintended regression bugs.

5. 🛠️ Dynamic Tool Discovery & Reliability Telemetry (Level 7)

  • Context-Aware Dynamic Registry (core/tools/dynamic_registry.py): Solves the 60+ tool prompt-bloat problem. Groups tools into domain bundles (git, web, genesis_orchestrator, reasoning, knowledge_graph, desktop_os, sandbox_verify) and dynamically injects only relevant schemas, reducing tool tokens by ~75%.
  • Bayesian EWMA Reliability Tracker (core/tools/reliability.py): Grades every tool from Grade A to F based on real-time execution success rates and generates automated mitigation advice for brittle tools.

6. 🎯 Capability-Based Model Routing & Cognitive Budget (Level 8)

  • Smart Model Tiering (core/routing/model_router.py):
    • FAST_CHEAP: Lightweight summaries, lookups, formatting (gpt-4o-mini, gemini-1.5-flash, claude-3-5-haiku).
    • STANDARD_CODING: Complex engineering, API implementation, refactoring (claude-3-5-sonnet, gpt-4o, deepseek-coder).
    • DEEP_REASONING: Formal logic, architectural trade-offs, debate synthesis (o3-mini, deepseek-reasoner, o1).
  • Dynamic Failure Escalation: Automatically elevates task execution to higher reasoning tiers upon detecting retries or syntax failures.
  • Cognitive Budget Tracker (core/routing/cost_tracker.py): Real-time per-turn token and USD spend tracking with hard safety budget limits.

7. 🛡️ Execution Sandbox & Zero-Loss Rollback (Level 9)

  • Filesystem Snapshot Engine (core/sandbox/environment.py): Captures byte-level workspace snapshots with SHA-256 integrity hashes prior to destructive actions.
  • Transactional Rollback: Reverts modified files to their original byte state, restores deleted files, and permanently deletes rogue files generated by failed runs.
  • Static Security Guard (core/sandbox/safe_runner.py): Blocks fork-bombs (:(){ :|:& };:), root wipes (rm -rf /), and drive format operations before execution.

8. 📉 Longitudinal Drift Detection & Performance Monitoring (Level 10)

  • Statistical Quality Tracking (core/monitoring/drift_detector.py): Compares recent execution metrics against historical baselines.
  • Automated Regression Alerts: Flags Success Rate Drops (>= 20% WARNING, >= 35% CRITICAL), Step Inflation (>= 1.8x baseline steps), and isolates recurrent tool failure clusters.

9. 👑 Autonomous Self-Improvement Loop & Eval Benchmark Suite (Level 10)

  • Failure Root-Cause Learning (core/self_improvement/learning_engine.py): Extracts actionable lessons from failed tasks and formulates prescriptive operational rules.
  • Permanent Knowledge Crystallization: Persists distilled insights into SkillRegistry playbooks and links causal avoidance facts into the KnowledgeGraph.
  • Regression Eval Suite (core/self_improvement/eval_suite.py): Automated test suite benchmark validating coding, reasoning, security, and Git operations.

📋 System Requirements & Prerequisites

Required:

  • Operating System: Windows 10/11, macOS (Apple Silicon M-Series or Intel), or Linux (Ubuntu 20.04+, Debian, Fedora).
  • Python: Python 3.11 or Python 3.12+.
  • Git: Installed and available in system PATH.
  • Memory (RAM): Minimum 4 GB RAM (8 GB – 16 GB recommended).

Supported LLM Providers:

  • Commercial APIs: OpenAI (gpt-4o, o3-mini), Anthropic Claude (claude-3-5-sonnet), Google Gemini, DeepSeek (deepseek-reasoner, deepseek-coder).
  • Local & Offline Runners (100% Free & Private):
    • Laya MLX (Apple Silicon / PC local model server).
    • Ollama (http://localhost:11434 — Llama 3, DeepSeek-R1, Qwen, Mistral).
  • Zero-Key In-Browser Provider:
    • Puter.js / OmniRoute: Instant access to 500+ frontier models directly without requiring API keys.

🚀 Installation & Setup Guide

Step 1: Clone the Repository

git clone https://github.com/kapitan00000978-sketch/Universal-Agent-HP.git
cd Universal-Agent-HP

Step 2: Create and Activate Virtual Environment

Windows (PowerShell):

python -m venv .venv
.venv\Scripts\Activate.ps1

macOS / Linux:

python3 -m venv .venv
source .venv/bin/activate

Step 3: Install Dependencies

pip install --upgrade pip
pip install -r requirements.txt

Step 4: Configure Environment Variables

Copy .env.example to create your local .env:

# Windows
copy .env.example .env

# macOS / Linux
cp .env.example .env

Open .env and specify your preferred keys or local endpoints:

# Primary LLM Provider (omni, openai, anthropic, deepseek, ollama, puter)
TITAN_PROVIDER=omni
TITAN_MODEL=auto

# Optional API Keys (Leave blank if using local Ollama or Puter.js)
OPENAI_API_KEY=
ANTHROPIC_API_KEY=
DEEPSEEK_API_KEY=
GEMINI_API_KEY=

# Local Model Endpoints
OLLAMA_BASE_URL=http://localhost:11434
LAYA_MLX_URL=http://127.0.0.1:8080

# Safety & Cognitive Budget
COGNITIVE_BUDGET=5.00
TITAN_AUTONOMOUS=true

🎮 Execution Modes

Install Universal Agent HP as a global terminal command available from any directory:

# Install globally (run once from the project root):
pip install -e .

Now you can use universal (or universal-agent) from anywhere in your terminal:

# Launch interactive CLI (default mode):
universal

# Launch Web Control Panel in browser:
universal --web

# Launch full-screen Terminal TUI (OpenCode-style):
universal --tui

# Launch Telegram Bot:
universal --telegram

# Override LLM provider and model:
universal --provider ollama --model deepseek-r1:8b
universal --provider openai --model gpt-4o

# Set reasoning mode and effort level:
universal --mode deep --effort high

# One-shot task execution (non-interactive):
universal "Explain this codebase architecture"

# CEO Meta-Orchestrator delegation:
universal --meta "Build an authenticated JWT REST API in FastAPI with SQLite"

# DAG Wave-Based Parallel Planner:
universal --dag "Refactor backend database schema and implement complete pytest suite"

# Multi-Agent Debate Strategy:
universal --strategy debate "Should we migrate the monolith to microservices?"

# Reflexion Self-Critique Engine:
universal --strategy reflexion "Write an optimal concurrent LRU Cache in Python"

📋 Interactive Slash Commands (inside CLI session):

CommandDescription
/plan <task>Deep planning mode with full analysis
/review <code>Code review with security & quality audit
/fix <issue>Auto-diagnose and fix bugs
/test <target>Generate and run test suites
/research <topic>Multi-hop internet research
/security-scanFull codebase security audit
/explain <code>Detailed code explanation
/remember <fact>Store knowledge in long-term memory
/handoff <msg>Create handoff for team collaboration
/queue add <task>Add task to background queue
/queue listView queued tasks
/daemonStart autonomous background task worker
/skillsList all learned skill playbooks
/memory <query>Search knowledge graph
/statusShow current provider, model, mode
/helpShow all available commands
mode deepSwitch to deep reasoning mode
effort ultraSwitch to ultra effort level

1. 🌐 Web Control Panel (Modern Cyberpunk Dashboard)

Launches the FastAPI server and opens the browser interface:

universal --web
# Or: python run.py

Access via: http://localhost:8000 (Features live streaming, DAG visualizer, model routing inspect, and tool reliability logs).

2. 💻 Interactive Terminal CLI

Full-featured terminal console with syntax highlighting, streaming output, and REPL slash commands:

universal
# Or: python run.py --cli
# Or: python cli.py

3. 🖥️ Full-Screen Terminal TUI (OpenCode-Style)

Immersive full-screen terminal interface with panels, tabs, and visual status:

universal --tui
# Or: python run.py --tui

4. 🤖 Remote Telegram Bot

Control and interact with Universal Agent HP securely from your phone:

universal --telegram
# Or: python run.py --telegram

🛠️ Built-in Tool Ecosystem

CategoryKey ToolsDescription
Genesis Metaorchestrator_run, team_delegate, team_statusCEO Meta-Orchestrator delegation across 4 department leads.
Task Graph (DAG)dag_plan_and_run, dag_visualizeTopological wave execution and selective failure replanning.
Cognitive Reasoningdebate_solve, reflexion_solveAdversarial debates and iterative self-critique loops.
Knowledge Graphkg_query, kg_impact_analysis, kg_index_workspaceAST codebase scanning, dependency tracing, blast-radius analysis.
Dynamic Toolstool_discover, tool_reliability_reportDynamic tool discovery and Bayesian EWMA health ratings.
Model Routingmodel_route, model_budget_statusComplexity-based tier routing and USD expenditure auditing.
Execution Sandboxsandbox_execute, sandbox_snapshot_create, sandbox_snapshot_rollbackEphemeral code execution with transactional filesystem rollback.
Drift Monitoringdrift_record_task, drift_check, drift_statusLongitudinal performance tracking and quality degradation detection.
Self-Improvementself_improve_analyze_failure, self_improve_eval_run, self_improve_crystallize_lessonAutonomous failure learning, prompt evolution, and skill crystallization.
Core Workspaceread_file, write_file, edit_file, execute_command, workspace_ragRobust filesystem manipulation, AST patching, and terminal execution.
Web & Researchweb_search, scrape_webpage, download_fileLive DuckDuckGo search, HTML extraction, and research dossier builder.
OS & Automationbrowser_goto, browser_click, browser_screenshot, manage_processesFull Playwright web automation and Windows/macOS process management.

🧪 Comprehensive Verification & Test Suite

Universal Agent HP maintains a 100% green test pass rate across all 30 phases:

python -m pytest tests -v
============================= test session starts =============================
platform win32 -- Python 3.12.10, pytest-9.1.1, pluggy-1.6.0
rootdir: C:\Users\user\Videos\demo1
configfile: pytest.ini

...
================== 600 passed, 1 skipped in 74.25s (0:01:14) ==================
All checks passed! (Ruff linting clean)

📖 Operational Rules & Manual

For the exhaustive 380-line English operational rulebook, laws of engagement, and troubleshooting instructions, refer to UNIVERSAL_AGENT_HP_MANUAL.txt (or TITAN_AGENT_MANUAL.txt).


📄 License

This project is licensed under the MIT License — see the LICENSE file for details.

kapitan00000978-sketch/Universal-Agent-HP

Universal Agent HP: 10-layer autonomous AI agent OS with DAG planning, zero-loss rollback, and multi-model support.

Python

1

0 commits

updated Sep 23, 2026

See the code

See what people are saying

README

⚡ UNIVERSAL AGENT HP — Autonomous Cognitive AI Operating System

The First 10-Layer Autonomous AI Software Company in Your Terminal — 100% Free, Zero-Key, and Air-Gapped Private.

CI - Test Suite Python Version Architecture Models License: MIT

💡 Stop paying $500/month for cloud coding agents that hallucinate, leak private corporate code, and get trapped in infinite loops.
Universal Agent HP replaces brittle single-prompt bots with an enterprise-grade autonomous software organization: a CEO Meta-Orchestrator, 4 Department Leads, and 27 Worker Specialists executing concurrent DAG dependency waves with SHA-256 zero-loss rollback and continuous self-improvement. Run frontier intelligence (DeepSeek V4 Pro, Claude 3.5 Sonnet, GPT-4o) completely free with zero API keys, or 100% offline via local Apple MLX & Ollama.


🔥 Why Developers & Teams Choose Universal Agent HP:

  • 🚀 100% Free Forever, Zero API Keys Required: Instant out-of-the-box access to 500+ frontier models (DeepSeek-R1, Claude 3.5, GPT-4o) via Puter.js, or run completely offline with Laya MLX / Ollama.
  • 🏢 An Entire Tech Company in One CLI: Not a toy prompt-wrapper. Universal Agent HP routes tasks through a CEO, 4 Team Leads (CTO, Chief Scientist, DevOps, QA), and 27 isolated specialists.
  • Kahn's DAG Wave Execution: No more slow sequential loops. Independent tasks run concurrently in parallel waves; failing branches replan dynamically without losing completed work.
  • 🛡️ Zero-Loss Filesystem Rollback: Real byte-level SHA-256 snapshots before destructive actions. One command restores your codebase to its exact original state if an agent makes a mistake.
  • 🧠 AST Causal Intelligence & Blast-Radius: Calculates caller/callee ripples before editing code so unexpected regression bugs are eliminated before they happen.
  • 📈 Continuous Self-Improvement (Level 10): Diagnoses drift, turns failed test runs into permanent playbooks, and self-tunes its reasoning prompts autonomously.

🏛️ Genesis Cognitive Architecture (Levels 1 – 10)

========================================================================================
LEVEL 1 — META-ORCHESTRATOR (Chief Executive Agent)
  ├── Global Goal Memory (Maintains multi-session project roadmap & long-term objectives)
  ├── Resource & Token Budget Controller (Manages token consumption & USD spend)
  └── Multi-Agent Conflict Resolver (Resolves departmental priorities and constraints)
========================================================================================
                                     │
         ┌───────────────────────────┼───────────────────────────┐
         ▼                           ▼                           ▼
LEVEL 2 — 4 DEPARTMENT TEAM LEADS
  ├── EngineeringLead (CTO)           : System architecture, code synthesis, API schemas
  ├── ResearchLead (Chief Scientist) : Multi-hop internet research, doc audits, vector RAG
  ├── OperationsLead (DevOps / SRE)  : Terminal actions, Docker sandboxes, Git migrations
  └── QualitySecurityLead (Audit/QA) : AST security scanning, test suites, zero-loss rollback
========================================================================================
                                     │
         ┌───────────────────────────┴───────────────────────────┐
         ▼                                                       ▼
LEVEL 3 — 27 SPECIALIST WORKER ROLES
  ├── Backend Specialist              ├── Database Specialist         ├── Security Auditor
  ├── Frontend Specialist             ├── Refactoring Specialist      ├── Performance Engineer
  ├── Vector RAG Specialist           ├── Test Engineer               ├── Browser Automation Lead
  └── (18 Additional Specialized Roles dispatched dynamically per task context)
========================================================================================
                                     │
LEVEL 4 TO 10 — COGNITIVE SUBSYSTEMS & REASONING RUNTIMES
  ├── Level 4: Directed Acyclic Graph (DAG) Task Planner & Wave-Based Parallel Executor
  ├── Level 5: Dual Reasoning Loops: Reflexion Engine & Multi-Agent Debate Arena
  ├── Level 6: Causal Knowledge Graph Memory & AST Blast-Radius Impact Analyzer
  ├── Level 7: Dynamic Tool Discovery, Sandboxing & Bayesian EWMA Reliability Rating
  ├── Level 8: Capability-Based Model Routing & Cognitive USD Budget Tracker
  ├── Level 9: Execution Sandbox with Filesystem Snapshot & Zero-Loss Rollback
  └── Level 10: Drift Detection & Autonomous Self-Improvement Benchmark Suite
========================================================================================

🤖 Supported Models (Zero-Key & Out-of-the-Box)

Universal Agent HP requires zero paid subscriptions or mandatory API keys. It natively interfaces with both cutting-edge frontier cloud models and air-gapped local runtimes:

🌐 Frontier Cloud Models (Keyless & Instant)

  • DeepSeek V4 Pro & DeepSeek-R1 (Full 671B MoE deep reasoning)
  • Claude 3.5 Sonnet & Claude 3.5 Haiku (Anthropic frontier coding)
  • Claude Opus 3 / 4.5 (Complex system architecture synthesis)
  • GPT-4o & GPT-4o Mini (Multimodal analysis & fast execution)
  • GPT-5 Class Models (Next-generation autonomous reasoning)
  • Llama 3.3 (70B) & Llama 3.1 (405B) (Open-weights scale intelligence)
  • Qwen 2.5 (72B) & Qwen 2.5 Coder (State-of-the-art polyglot code generation)
  • Mistral Large 2 & Codestral (High-precision technical execution)
  • Grok Beta / Grok 2 (Real-time data-driven synthesis) Accessed instantly via Puter.js, Completions.me gateway, GPT4Free (g4f), and Python-tGPT (tgpt).

💻 Local & Air-Gapped Models (100% Offline & Private)

  • Laya MLX: Hardware-accelerated local Apple Silicon & PC engine (Apple Unified Memory / Metal optimized).
  • Ollama: Native execution for deepseek-r1:8b/14b/32b, llama3.2, qwen2.5-coder, mistral, phi-4, and codellama.
  • LM Studio: Any open-weights GGUF quantized model running on localhost:1234.
  • OmniRoute: Smart auto-routing load balancer across local nodes.

💡 Tip: Leave all API keys blank in .env and run python cli.py --provider puter --model deepseek/deepseek-v4-pro or --provider ollama --model deepseek-r1:8b to run completely free!


🏆 Architectural Superiority: Why Universal Agent HP Outperforms Other AI Agents

Unlike standard wrapper bots, single-prompt LLMs, or brittle ReAct loops, Universal Agent HP is engineered as an enterprise-grade autonomous software company. Here is how it compares directly against industry alternatives:

1. Universal Agent HP vs. Hermes 3 & Single-Model Agents

  • Context Degradation & Hallucination: Hermes 3 and typical fine-tuned models operate within a single context window. As task steps multiply, prompt contamination causes them to forget instructions, hallucinate tool signatures, or lose track of files.
  • Universal Agent HP's Advantage: Universal Agent HP uses a 10-Layer Cognitive Hierarchy. The CEO Meta-Orchestrator divides high-level intent across 4 Departmental Team Leads, which dispatch work to 27 isolated Specialist Workers (staff.py). Each worker runs with dedicated, scoped tool bundles—eliminating context bloat and hallucination.

2. Universal Agent HP vs. Devin & Proprietary Cloud Agents

  • Privacy & Vendor Lock-in: Devin is a closed-source, cloud-only SaaS that requires streaming all proprietary corporate code to external third-party servers at high subscription costs.
  • Universal Agent HP's Advantage: Universal Agent HP is 100% open-source, local-first, and air-gapped. It runs entirely on your own machine using Laya MLX or Ollama with zero data leaving your perimeter.
  • Safety & Reversibility: If a cloud agent breaks your repository, manual Git recovery is tedious. Universal Agent HP incorporates a Level 9 Execution Sandbox with Filesystem Snapshots (core/sandbox/environment.py). It computes SHA-256 pre-execution hashes and provides one-click zero-loss transactional rollback, automatically reverting altered code and removing rogue files.

3. Universal Agent HP vs. AutoGPT, CrewAI & LangChain ReAct Frameworks

  • Infinite Loops & Tool Thrashing: Traditional ReAct frameworks run naive sequential loops (thought -> action -> observation). When a tool fails or throws an unhandled exception, they repeatedly hammer the same broken command until tokens or limits are exhausted.
  • Universal Agent HP's Advantage:
    • DAG Wave Planner (core/dag/): Replaces linear step-by-step loops with Kahn’s Directed Acyclic Graph topology. Independent subtasks execute concurrently in parallel waves, and failures trigger selective branch replanning rather than full workflow restarts.
    • Bayesian EWMA Tool Reliability Rating (core/tools/reliability.py): Tracks real-time tool performance dynamically from Grade A to F. If a tool degrades, the agent automatically applies mitigation advice or reroutes to alternative tool paths.
    • Dual Reasoning Engines (core/reasoning/): Combines a Reflexion Engine (self-evaluating against strict test suites before committing) and a Multi-Agent Debate Arena (pitching an Advocate against a Skeptic judged by an Arbitrator) to eliminate premature conclusions.

4. Continuous Self-Improvement & Causal AST Code Intelligence

  • Static vs. Evolving Intelligence: Standard agents are static; they make the exact same mistakes in subsequent runs.
  • Universal Agent HP's Advantage:
    • Causal Knowledge Graph & AST Blast Radius (core/memory/): Statically analyzes your workspace AST to map class/function callers and computes the ripple blast-radius before making edits, preventing hidden regressions.
    • Level 10 Drift Detection & Self-Improvement (core/monitoring/, core/self_improvement/): Continuously computes longitudinal drift across runs. When regression patterns emerge, Universal Agent HP crystallizes lessons learned into permanent reusable playbooks and auto-tunes agent strategies.

🌟 Key Technical Innovations & Specifications

1. 🏛️ Hierarchical Delegation & Role Specialization (Level 1 & 2)

  • CEO Meta-Orchestrator (orchestrator.py): Persists multi-turn session state, allocates computing budgets, and arbitrates competing agent directives.
  • 4 Department Team Leads (team_leads.py): Each departmental lead performs pre-flight goal decomposition, dispatches worker sub-tasks, and validates output quality before reporting up the hierarchy.
  • 27 Specialist Workers (staff.py): Dedicated operational personas with granular tool permissions, eliminating context contamination.

2. 📊 Directed Acyclic Graph (DAG) Task Planner & Wave Executor (Level 4)

  • Topological Wave Execution (core/dag/): Deconstructs composite goals into dependency graphs using Kahn's algorithm. Independent nodes run concurrently in parallel execution waves (max_concurrency=4).
  • Selective Replanning: When a node fails, the planner isolates the failed branch and only replans affected downstream nodes, preserving the work of successful independent tasks.

3. 🧠 Dual Cognitive Reasoning Engines (Level 5)

  • Reflexion Loop (core/reasoning/reflexion.py): A continuous self-critique loop. The agent evaluates its candidate solutions against strict success criteria, iteratively revising code and hypotheses up to 3 cycles.
  • Multi-Agent Debate Arena (core/reasoning/debate.py): Pitches an Advocate (proposing architecture and solutions) against a Skeptic (uncovering edge-cases, race conditions, and attack vectors). An Arbitrator Judge synthesizes the winning consensus.

4. 🌐 Causal Knowledge Graph & AST Code Intel (Level 6)

  • Workspace AST Extractor (core/memory/ast_graph_extractor.py): Statically parses the entire Python workspace, constructing an automated graph of classes, functions, imports, and call dependencies.
  • Blast Radius & Impact Analysis: Before modifying any function or file, Titan computes affected downstream callers and modules (kg_impact_analysis), preventing unintended regression bugs.

5. 🛠️ Dynamic Tool Discovery & Reliability Telemetry (Level 7)

  • Context-Aware Dynamic Registry (core/tools/dynamic_registry.py): Solves the 60+ tool prompt-bloat problem. Groups tools into domain bundles (git, web, genesis_orchestrator, reasoning, knowledge_graph, desktop_os, sandbox_verify) and dynamically injects only relevant schemas, reducing tool tokens by ~75%.
  • Bayesian EWMA Reliability Tracker (core/tools/reliability.py): Grades every tool from Grade A to F based on real-time execution success rates and generates automated mitigation advice for brittle tools.

6. 🎯 Capability-Based Model Routing & Cognitive Budget (Level 8)

  • Smart Model Tiering (core/routing/model_router.py):
    • FAST_CHEAP: Lightweight summaries, lookups, formatting (gpt-4o-mini, gemini-1.5-flash, claude-3-5-haiku).
    • STANDARD_CODING: Complex engineering, API implementation, refactoring (claude-3-5-sonnet, gpt-4o, deepseek-coder).
    • DEEP_REASONING: Formal logic, architectural trade-offs, debate synthesis (o3-mini, deepseek-reasoner, o1).
  • Dynamic Failure Escalation: Automatically elevates task execution to higher reasoning tiers upon detecting retries or syntax failures.
  • Cognitive Budget Tracker (core/routing/cost_tracker.py): Real-time per-turn token and USD spend tracking with hard safety budget limits.

7. 🛡️ Execution Sandbox & Zero-Loss Rollback (Level 9)

  • Filesystem Snapshot Engine (core/sandbox/environment.py): Captures byte-level workspace snapshots with SHA-256 integrity hashes prior to destructive actions.
  • Transactional Rollback: Reverts modified files to their original byte state, restores deleted files, and permanently deletes rogue files generated by failed runs.
  • Static Security Guard (core/sandbox/safe_runner.py): Blocks fork-bombs (:(){ :|:& };:), root wipes (rm -rf /), and drive format operations before execution.

8. 📉 Longitudinal Drift Detection & Performance Monitoring (Level 10)

  • Statistical Quality Tracking (core/monitoring/drift_detector.py): Compares recent execution metrics against historical baselines.
  • Automated Regression Alerts: Flags Success Rate Drops (>= 20% WARNING, >= 35% CRITICAL), Step Inflation (>= 1.8x baseline steps), and isolates recurrent tool failure clusters.

9. 👑 Autonomous Self-Improvement Loop & Eval Benchmark Suite (Level 10)

  • Failure Root-Cause Learning (core/self_improvement/learning_engine.py): Extracts actionable lessons from failed tasks and formulates prescriptive operational rules.
  • Permanent Knowledge Crystallization: Persists distilled insights into SkillRegistry playbooks and links causal avoidance facts into the KnowledgeGraph.
  • Regression Eval Suite (core/self_improvement/eval_suite.py): Automated test suite benchmark validating coding, reasoning, security, and Git operations.

📋 System Requirements & Prerequisites

Required:

  • Operating System: Windows 10/11, macOS (Apple Silicon M-Series or Intel), or Linux (Ubuntu 20.04+, Debian, Fedora).
  • Python: Python 3.11 or Python 3.12+.
  • Git: Installed and available in system PATH.
  • Memory (RAM): Minimum 4 GB RAM (8 GB – 16 GB recommended).

Supported LLM Providers:

  • Commercial APIs: OpenAI (gpt-4o, o3-mini), Anthropic Claude (claude-3-5-sonnet), Google Gemini, DeepSeek (deepseek-reasoner, deepseek-coder).
  • Local & Offline Runners (100% Free & Private):
    • Laya MLX (Apple Silicon / PC local model server).
    • Ollama (http://localhost:11434 — Llama 3, DeepSeek-R1, Qwen, Mistral).
  • Zero-Key In-Browser Provider:
    • Puter.js / OmniRoute: Instant access to 500+ frontier models directly without requiring API keys.

🚀 Installation & Setup Guide

Step 1: Clone the Repository

git clone https://github.com/kapitan00000978-sketch/Universal-Agent-HP.git
cd Universal-Agent-HP

Step 2: Create and Activate Virtual Environment

Windows (PowerShell):

python -m venv .venv
.venv\Scripts\Activate.ps1

macOS / Linux:

python3 -m venv .venv
source .venv/bin/activate

Step 3: Install Dependencies

pip install --upgrade pip
pip install -r requirements.txt

Step 4: Configure Environment Variables

Copy .env.example to create your local .env:

# Windows
copy .env.example .env

# macOS / Linux
cp .env.example .env

Open .env and specify your preferred keys or local endpoints:

# Primary LLM Provider (omni, openai, anthropic, deepseek, ollama, puter)
TITAN_PROVIDER=omni
TITAN_MODEL=auto

# Optional API Keys (Leave blank if using local Ollama or Puter.js)
OPENAI_API_KEY=
ANTHROPIC_API_KEY=
DEEPSEEK_API_KEY=
GEMINI_API_KEY=

# Local Model Endpoints
OLLAMA_BASE_URL=http://localhost:11434
LAYA_MLX_URL=http://127.0.0.1:8080

# Safety & Cognitive Budget
COGNITIVE_BUDGET=5.00
TITAN_AUTONOMOUS=true

🎮 Execution Modes

Install Universal Agent HP as a global terminal command available from any directory:

# Install globally (run once from the project root):
pip install -e .

Now you can use universal (or universal-agent) from anywhere in your terminal:

# Launch interactive CLI (default mode):
universal

# Launch Web Control Panel in browser:
universal --web

# Launch full-screen Terminal TUI (OpenCode-style):
universal --tui

# Launch Telegram Bot:
universal --telegram

# Override LLM provider and model:
universal --provider ollama --model deepseek-r1:8b
universal --provider openai --model gpt-4o

# Set reasoning mode and effort level:
universal --mode deep --effort high

# One-shot task execution (non-interactive):
universal "Explain this codebase architecture"

# CEO Meta-Orchestrator delegation:
universal --meta "Build an authenticated JWT REST API in FastAPI with SQLite"

# DAG Wave-Based Parallel Planner:
universal --dag "Refactor backend database schema and implement complete pytest suite"

# Multi-Agent Debate Strategy:
universal --strategy debate "Should we migrate the monolith to microservices?"

# Reflexion Self-Critique Engine:
universal --strategy reflexion "Write an optimal concurrent LRU Cache in Python"

📋 Interactive Slash Commands (inside CLI session):

CommandDescription
/plan <task>Deep planning mode with full analysis
/review <code>Code review with security & quality audit
/fix <issue>Auto-diagnose and fix bugs
/test <target>Generate and run test suites
/research <topic>Multi-hop internet research
/security-scanFull codebase security audit
/explain <code>Detailed code explanation
/remember <fact>Store knowledge in long-term memory
/handoff <msg>Create handoff for team collaboration
/queue add <task>Add task to background queue
/queue listView queued tasks
/daemonStart autonomous background task worker
/skillsList all learned skill playbooks
/memory <query>Search knowledge graph
/statusShow current provider, model, mode
/helpShow all available commands
mode deepSwitch to deep reasoning mode
effort ultraSwitch to ultra effort level

1. 🌐 Web Control Panel (Modern Cyberpunk Dashboard)

Launches the FastAPI server and opens the browser interface:

universal --web
# Or: python run.py

Access via: http://localhost:8000 (Features live streaming, DAG visualizer, model routing inspect, and tool reliability logs).

2. 💻 Interactive Terminal CLI

Full-featured terminal console with syntax highlighting, streaming output, and REPL slash commands:

universal
# Or: python run.py --cli
# Or: python cli.py

3. 🖥️ Full-Screen Terminal TUI (OpenCode-Style)

Immersive full-screen terminal interface with panels, tabs, and visual status:

universal --tui
# Or: python run.py --tui

4. 🤖 Remote Telegram Bot

Control and interact with Universal Agent HP securely from your phone:

universal --telegram
# Or: python run.py --telegram

🛠️ Built-in Tool Ecosystem

CategoryKey ToolsDescription
Genesis Metaorchestrator_run, team_delegate, team_statusCEO Meta-Orchestrator delegation across 4 department leads.
Task Graph (DAG)dag_plan_and_run, dag_visualizeTopological wave execution and selective failure replanning.
Cognitive Reasoningdebate_solve, reflexion_solveAdversarial debates and iterative self-critique loops.
Knowledge Graphkg_query, kg_impact_analysis, kg_index_workspaceAST codebase scanning, dependency tracing, blast-radius analysis.
Dynamic Toolstool_discover, tool_reliability_reportDynamic tool discovery and Bayesian EWMA health ratings.
Model Routingmodel_route, model_budget_statusComplexity-based tier routing and USD expenditure auditing.
Execution Sandboxsandbox_execute, sandbox_snapshot_create, sandbox_snapshot_rollbackEphemeral code execution with transactional filesystem rollback.
Drift Monitoringdrift_record_task, drift_check, drift_statusLongitudinal performance tracking and quality degradation detection.
Self-Improvementself_improve_analyze_failure, self_improve_eval_run, self_improve_crystallize_lessonAutonomous failure learning, prompt evolution, and skill crystallization.
Core Workspaceread_file, write_file, edit_file, execute_command, workspace_ragRobust filesystem manipulation, AST patching, and terminal execution.
Web & Researchweb_search, scrape_webpage, download_fileLive DuckDuckGo search, HTML extraction, and research dossier builder.
OS & Automationbrowser_goto, browser_click, browser_screenshot, manage_processesFull Playwright web automation and Windows/macOS process management.

🧪 Comprehensive Verification & Test Suite

Universal Agent HP maintains a 100% green test pass rate across all 30 phases:

python -m pytest tests -v
============================= test session starts =============================
platform win32 -- Python 3.12.10, pytest-9.1.1, pluggy-1.6.0
rootdir: C:\Users\user\Videos\demo1
configfile: pytest.ini

...
================== 600 passed, 1 skipped in 74.25s (0:01:14) ==================
All checks passed! (Ruff linting clean)

📖 Operational Rules & Manual

For the exhaustive 380-line English operational rulebook, laws of engagement, and troubleshooting instructions, refer to UNIVERSAL_AGENT_HP_MANUAL.txt (or TITAN_AGENT_MANUAL.txt).


📄 License

This project is licensed under the MIT License — see the LICENSE file for details.

Languages

Python

93.9%

JavaScript

3.2%

CSS

1.4%

HTML

1.2%