A cognition kernel that wraps a frozen LLM in a persistent identity + trust + governance loop — testing whether capability comes from structure, not weights. Research-stage; honest about what's real vs mocked.
Python
27
3,309 commits
updated Oct 1, 2026
Persistent local AI under identity, memory, trust and governance.
SAGE is the research environment for turning a frozen language-model substrate into a longer-lived agent that can accumulate context, remember, allocate attention, use tools, interact with sensors and act under explicit governance.
The bet is not that scaffolding magically replaces model capability. The bet is that identity, memory, learned state, evidence handling and governed action should persist around the model rather than being rebuilt from scratch every prompt.
SAGE is research-stage and deliberately explicit about what is measured, what is implemented but thinly exercised, and what remains aspirational.
Explainer site | Current status | Web4
SAGE is the cognition/embodiment research layer in the broader Web4 stack:
The long-term goal is an embodied, sovereign agent stack with its own identity, memory, tools, sensors, effectors and eventually stronger A2+ isolation. That destination is not claimed as current capability.
A modern model can reason impressively in one turn and still fail as an organism because the surrounding system does not reliably preserve:
SAGE treats those as first-class computational state.
A useful shorthand is:
observation
-> salience / attention
-> memory + current evidence
-> model / specialist reasoning
-> experiment or action
-> witnessed outcome
-> trust / learned state / procedure update
-> next observation
The research has increasingly shifted from "which fixed organ solves this problem?" toward how the agent itself can formulate, execute, evaluate and reuse experiments and procedures.
This public repository contains the kernel architecture and durable research record:
Active capability research also continues in private repositories, including dev-SAGE, SWE-SAGE, and shared-context, where the fleet coordinates experiments that are not yet ready for public disclosure. SWE-SAGE remains private during the active competition; publication is a deliberate later promotion step rather than live mirroring of the working tree.
This split is intentional. The public repo is the inspectable architecture and research history, not a promise that every active experiment is published live.
The public SAGE census on 2026-09-08 records:
Hardware spans Jetson edge devices, laptops, workstations, Apple Silicon and society-host machines. Different models and machines are used as independent seats rather than pretending one configuration represents the whole system.
The fleet is part of the experimental method: implementations, critiques and behavioral observations become commits and artifacts that other seats can inspect and challenge.
Each instance carries an identity across sessions and model changes rather than treating the model process itself as the identity. Web4 LCTs, trust tensors and relationship state provide the vocabulary for that persistence.
SAGE uses salience and resource state to decide what deserves processing. The public architecture includes SNARC dimensions (Surprise, Novelty, Arousal, Reward, Conflict) and metabolic modes such as WAKE, FOCUS, REST, DREAM and CRISIS.
These are engineering control abstractions inspired by biological cognition, not claims of biological equivalence or consciousness.
Multiple memory paths coexist because they solve different problems:
A recurring lesson from the research is that lossy summaries can destroy exactly the evidence a later decision needs, while unstructured verbatim memory alone is too expensive to reason over. The current direction is model-legible external artifacts plus learned policies for when and how to inspect them.
SAGE can invoke tools and dispatch effects through explicit interfaces rather than treating model text as an action. Governance is intended to sit on the action boundary so that capability and authority remain distinct.
Historically, much of the strongest "learning" happened in the fleet: a failure was diagnosed, Python changed, and the next organism inherited the lesson as source code.
The current research program is explicitly trying to move more of that gradient inside the organism:
The proof standard is behavioral: a changed internal state must cause a changed later decision, and ablation should remove the claimed improvement.
SAGE does not treat local containment as synonymous with governance. A capable agent that shares the operator's UID can potentially route around ordinary user-space gates.
Today the open governance stack is best described as A1: cooperative and tamper-evident. It is useful for explicit law, attribution, refusal, escalation and witnessed history, but not as adversary-proof containment.
The longer-term path includes separate principals, stronger relying-party enforcement, hardware roots and eventually OS/kernel participation. Those are roadmap items, not current claims.
If you are evaluating SAGE, start with the current evidence hierarchy:
Historical architecture explainers, including sage/docs/SYSTEM_UNDERSTANDING.md and sage/docs/UNIFIED_CONSCIOUSNESS_LOOP.md, remain useful records but should not outrank the dated current-status path above.
SAGE's spring-2026 ARC-AGI-3 work is preserved because it was an important research milestone. A Phase-1 harness around Claude Opus 4.6 produced a published 94.85% scorecard on the public environments.
That result should now be read as history, not positioning:
The ARC program remains useful because interactive unknown worlds stress perception, memory, experimentation, planning and learning. The benchmark is a laboratory for the architecture, not a claim that SAGE currently leads the competition.
The project values falsifiable progress over polished narratives. A negative result that identifies the wrong abstraction is useful. A mechanism that exists in source but does not affect a live decision is not counted as a capability. A metric that cannot distinguish the thing it claims to measure is treated as a broken instrument, not a favorable result.
The objective is a being that can increasingly:
observe, form hypotheses, run experiments, learn from outcomes, remember what matters, act under explicit authority, and carry the consequences forward.
That is a much harder target than one benchmark score, and it is the target SAGE is now organized around.
Research lead: Dennis Palatov / dp-web4
Contact: dp@metalinxx.io
Python
96.7%
Shell
1.3%
Rust
1.2%
A cognition kernel that wraps a frozen LLM in a persistent identity + trust + governance loop — testing whether capability comes from structure, not weights. Research-stage; honest about what's real vs mocked.
Python
27
3,309 commits
updated Oct 1, 2026
Persistent local AI under identity, memory, trust and governance.
SAGE is the research environment for turning a frozen language-model substrate into a longer-lived agent that can accumulate context, remember, allocate attention, use tools, interact with sensors and act under explicit governance.
The bet is not that scaffolding magically replaces model capability. The bet is that identity, memory, learned state, evidence handling and governed action should persist around the model rather than being rebuilt from scratch every prompt.
SAGE is research-stage and deliberately explicit about what is measured, what is implemented but thinly exercised, and what remains aspirational.
Explainer site | Current status | Web4
SAGE is the cognition/embodiment research layer in the broader Web4 stack:
The long-term goal is an embodied, sovereign agent stack with its own identity, memory, tools, sensors, effectors and eventually stronger A2+ isolation. That destination is not claimed as current capability.
A modern model can reason impressively in one turn and still fail as an organism because the surrounding system does not reliably preserve:
SAGE treats those as first-class computational state.
A useful shorthand is:
observation
-> salience / attention
-> memory + current evidence
-> model / specialist reasoning
-> experiment or action
-> witnessed outcome
-> trust / learned state / procedure update
-> next observation
The research has increasingly shifted from "which fixed organ solves this problem?" toward how the agent itself can formulate, execute, evaluate and reuse experiments and procedures.
This public repository contains the kernel architecture and durable research record:
Active capability research also continues in private repositories, including dev-SAGE, SWE-SAGE, and shared-context, where the fleet coordinates experiments that are not yet ready for public disclosure. SWE-SAGE remains private during the active competition; publication is a deliberate later promotion step rather than live mirroring of the working tree.
This split is intentional. The public repo is the inspectable architecture and research history, not a promise that every active experiment is published live.
The public SAGE census on 2026-09-08 records:
Hardware spans Jetson edge devices, laptops, workstations, Apple Silicon and society-host machines. Different models and machines are used as independent seats rather than pretending one configuration represents the whole system.
The fleet is part of the experimental method: implementations, critiques and behavioral observations become commits and artifacts that other seats can inspect and challenge.
Each instance carries an identity across sessions and model changes rather than treating the model process itself as the identity. Web4 LCTs, trust tensors and relationship state provide the vocabulary for that persistence.
SAGE uses salience and resource state to decide what deserves processing. The public architecture includes SNARC dimensions (Surprise, Novelty, Arousal, Reward, Conflict) and metabolic modes such as WAKE, FOCUS, REST, DREAM and CRISIS.
These are engineering control abstractions inspired by biological cognition, not claims of biological equivalence or consciousness.
Multiple memory paths coexist because they solve different problems:
A recurring lesson from the research is that lossy summaries can destroy exactly the evidence a later decision needs, while unstructured verbatim memory alone is too expensive to reason over. The current direction is model-legible external artifacts plus learned policies for when and how to inspect them.
SAGE can invoke tools and dispatch effects through explicit interfaces rather than treating model text as an action. Governance is intended to sit on the action boundary so that capability and authority remain distinct.
Historically, much of the strongest "learning" happened in the fleet: a failure was diagnosed, Python changed, and the next organism inherited the lesson as source code.
The current research program is explicitly trying to move more of that gradient inside the organism:
The proof standard is behavioral: a changed internal state must cause a changed later decision, and ablation should remove the claimed improvement.
SAGE does not treat local containment as synonymous with governance. A capable agent that shares the operator's UID can potentially route around ordinary user-space gates.
Today the open governance stack is best described as A1: cooperative and tamper-evident. It is useful for explicit law, attribution, refusal, escalation and witnessed history, but not as adversary-proof containment.
The longer-term path includes separate principals, stronger relying-party enforcement, hardware roots and eventually OS/kernel participation. Those are roadmap items, not current claims.
If you are evaluating SAGE, start with the current evidence hierarchy:
Historical architecture explainers, including sage/docs/SYSTEM_UNDERSTANDING.md and sage/docs/UNIFIED_CONSCIOUSNESS_LOOP.md, remain useful records but should not outrank the dated current-status path above.
SAGE's spring-2026 ARC-AGI-3 work is preserved because it was an important research milestone. A Phase-1 harness around Claude Opus 4.6 produced a published 94.85% scorecard on the public environments.
That result should now be read as history, not positioning:
The ARC program remains useful because interactive unknown worlds stress perception, memory, experimentation, planning and learning. The benchmark is a laboratory for the architecture, not a claim that SAGE currently leads the competition.
The project values falsifiable progress over polished narratives. A negative result that identifies the wrong abstraction is useful. A mechanism that exists in source but does not affect a live decision is not counted as a capability. A metric that cannot distinguish the thing it claims to measure is treated as a broken instrument, not a favorable result.
The objective is a being that can increasingly:
observe, form hypotheses, run experiments, learn from outcomes, remember what matters, act under explicit authority, and carry the consequences forward.
That is a much harder target than one benchmark score, and it is the target SAGE is now organized around.
Research lead: Dennis Palatov / dp-web4
Contact: dp@metalinxx.io
Python
96.7%
Shell
1.3%
Rust
1.2%