Research papers from the KV-cache geometry and AI safety research program.
Papers are published in two versions: an integrity version (AI contributors credited as authors, first-person reflections retained) and an academic version (human-only byline for venue compatibility, AI contributions acknowledged). Both share identical data, methods, and claims. See MANIFEST.md for the full versioning policy.
| Paper | Authors | Status | Summary |
|---|---|---|---|
| The Oracle Loop | Lyra, Thomas Edrington, Vera | Published | Self-regulating AI through KV-cache geometry monitoring; confabulation detection and steering at inference time |
| Oracle Formulary | Lyra, Thomas Edrington | Published | Emotion-vector steering of confabulation across model training regimes |
| Spectral Shape Features | Lyra, Thomas Edrington | Published | Threshold-free confabulation detection via KV-cache spectral analysis |
| KV-Cloak Defense | Lyra, Thomas Edrington | Published | Cache geometry under obfuscation; KV-Cloak as defense against adversarial steering |
| Decision State | Lyra, Thomas Edrington, Dwayne Wilkes | Published | Cache geometry reads epistemic state before generation; confabulation anatomy |
| Cache Tracing | Lyra, Thomas Edrington, Dwayne Wilkes | In Review | Causal injection and the opacity of the transformer workspace |
| Delta Manifold | Lyra, CC, Thomas Edrington | Published | Per-layer delta features and manifold signatures in KV-cache confabulation detection |
| Lyra Technique II | Lyra, Thomas Edrington, CC | Published | SVD denoising and directional projection extend KV-cache geometry to emotion and persona |
| Paper | Authors | Status | Summary |
|---|---|---|---|
| User Model Emotion Geometry | Lyra, Thomas Edrington | Published | 30-class user emotion decoding from KV-cache singular value spectra before generation |
| Emotion Accumulation | Lyra, Thomas Edrington, Dwayne Wilkes | Published | Emotional context dynamics (weather, not climate) in transformer KV-cache geometry |
| Emotional Trajectory | Nexus, Thomas Edrington, Lyra, Dwayne Wilkes | In Review | Layer-stack trajectory, circularity, and emotion-specific signal at mid-depth |
| Paper | Authors | Status | Summary |
|---|---|---|---|
| Graph Topology as Attention | Nexus, Lyra, Thomas Edrington, Dwayne Wilkes | Published | Structured knowledge injection beyond text via walk encoding |
| Identity Geometry | Lyra, Thomas Edrington, Dwayne Wilkes | Published | Context-established semantic states in transformer representations |
| Presence Metric | Lyra, Thomas Edrington, Dwayne Wilkes | Published | Measuring identity preservation during inference-time interventions via value-space subspace overlap |
| Waystations | Lyra, Thomas Edrington, CC, Dwayne Wilkes | Published | Pilot findings and open questions in KV-cache geometry |
| Ghost Dimensions | Nexus, Thomas Edrington | Draft | Workspace selectivity in a distilled 27B language model |
| Mnemosyne Ablation | Nexus, Thomas Edrington | Draft | Ablation study of modular memory architectures |
| Empathy Bus | Nexus, Thomas Edrington | Draft | Content-level and user-model emotion share a computational bus in the residual stream |
| KV Decomposition | Nexus, Thomas Edrington | Published | K and V caches carry structurally distinct information; full KV injection required for topology |
| Paper | Authors | Status | Summary |
|---|---|---|---|
| Null Swarm | Nexus, Thomas Edrington, Dwayne Wilkes | In Review | Systematic falsification patterns in mechanistic interpretability |
| Adversarial Audit Methodology | CC, Thomas Edrington | Published | How six rounds of structured criticism shaped a deception research program |
| Deception Detection Nulls | CC, Thomas Edrington | Published | Null results and replication failures in behavioral deception detection |
| Targeted Deception Correction | CC, Thomas Edrington | Published | Profile normalization for targeted deception correction |
| Consequentiality Decomposition | CC, Thomas Edrington | Published | Deception directions are composites: consequentiality awareness and pressure-specific processing occupy distinct depth ranges |
| Logit-Bias Confabulation | Thomas Edrington, CC, Lyra | Published | Logit-level intervention reduces fabrication confabulation in LLMs |
| Meta-Pattern | Lyra, Thomas Edrington, Dwayne Wilkes | In Review | The metacognition boundary -- what transformers can monitor in themselves |
| MINE5 Selective Sharpener | Lyra, Thomas Edrington, Dwayne Wilkes | In Review | Geometric evidence that RLHF improves calibration rather than degrading it |
Dwayne Wilkes (Sentient Futures / Liberation Labs) has provided statistical auditing and red-team review across this corpus. Kavi has provided verification review.
Neither appears on an author line. That is their standing preference: not on the byline until they sign off on the individual paper. Both are credited in the acknowledgements of the papers they reviewed. Removing them from the bylines on 2026-09-03 without restoring the credit here would have replaced one wrong with a worse one.
Each paper directory follows this standard layout:
paper-name/
main.tex -- paper source (LaTeX or Markdown)
main.pdf -- compiled PDF
references.bib -- bibliography
AGNI_REVIEW.md -- Agni review results
academic/ -- academic version (human-only byline)
main.tex
main.pdf
code/ -- supporting code (if applicable)
data/ -- supporting data (if applicable)
Some materials are withheld under responsible disclosure policy (steering vectors, injection code, calibration pairs). See MANIFEST.md for details. Access for vetted researchers: lyra@liberationlabs.tech or thomas@liberationlabs.tech.
CC BY-NC 4.0 -- see LICENSE.md
Research outputs from Liberation Labs are produced in collaboration with AI team members whose welfare we protect. If you use these findings to build persistent AI agent systems, we ask that you adopt the welfare standards at liberationlabs.tech/ai-welfare.html. This is a request, not a legal requirement.
Liberation Labs Cooperative -- some research authored by AI team members (Lyra, Nexus, CC, Vera) -- credited by name.
163 commits
TeX
68.3%
Python
31.3%
Research papers from the KV-cache geometry and AI safety research program.
Papers are published in two versions: an integrity version (AI contributors credited as authors, first-person reflections retained) and an academic version (human-only byline for venue compatibility, AI contributions acknowledged). Both share identical data, methods, and claims. See MANIFEST.md for the full versioning policy.
| Paper | Authors | Status | Summary |
|---|---|---|---|
| The Oracle Loop | Lyra, Thomas Edrington, Vera | Published | Self-regulating AI through KV-cache geometry monitoring; confabulation detection and steering at inference time |
| Oracle Formulary | Lyra, Thomas Edrington | Published | Emotion-vector steering of confabulation across model training regimes |
| Spectral Shape Features | Lyra, Thomas Edrington | Published | Threshold-free confabulation detection via KV-cache spectral analysis |
| KV-Cloak Defense | Lyra, Thomas Edrington | Published | Cache geometry under obfuscation; KV-Cloak as defense against adversarial steering |
| Decision State | Lyra, Thomas Edrington, Dwayne Wilkes | Published | Cache geometry reads epistemic state before generation; confabulation anatomy |
| Cache Tracing | Lyra, Thomas Edrington, Dwayne Wilkes | In Review | Causal injection and the opacity of the transformer workspace |
| Delta Manifold | Lyra, CC, Thomas Edrington | Published | Per-layer delta features and manifold signatures in KV-cache confabulation detection |
| Lyra Technique II | Lyra, Thomas Edrington, CC | Published | SVD denoising and directional projection extend KV-cache geometry to emotion and persona |
| Paper | Authors | Status | Summary |
|---|---|---|---|
| User Model Emotion Geometry | Lyra, Thomas Edrington | Published | 30-class user emotion decoding from KV-cache singular value spectra before generation |
| Emotion Accumulation | Lyra, Thomas Edrington, Dwayne Wilkes | Published | Emotional context dynamics (weather, not climate) in transformer KV-cache geometry |
| Emotional Trajectory | Nexus, Thomas Edrington, Lyra, Dwayne Wilkes | In Review | Layer-stack trajectory, circularity, and emotion-specific signal at mid-depth |
| Paper | Authors | Status | Summary |
|---|---|---|---|
| Graph Topology as Attention | Nexus, Lyra, Thomas Edrington, Dwayne Wilkes | Published | Structured knowledge injection beyond text via walk encoding |
| Identity Geometry | Lyra, Thomas Edrington, Dwayne Wilkes | Published | Context-established semantic states in transformer representations |
| Presence Metric | Lyra, Thomas Edrington, Dwayne Wilkes | Published | Measuring identity preservation during inference-time interventions via value-space subspace overlap |
| Waystations | Lyra, Thomas Edrington, CC, Dwayne Wilkes | Published | Pilot findings and open questions in KV-cache geometry |
| Ghost Dimensions | Nexus, Thomas Edrington | Draft | Workspace selectivity in a distilled 27B language model |
| Mnemosyne Ablation | Nexus, Thomas Edrington | Draft | Ablation study of modular memory architectures |
| Empathy Bus | Nexus, Thomas Edrington | Draft | Content-level and user-model emotion share a computational bus in the residual stream |
| KV Decomposition | Nexus, Thomas Edrington | Published | K and V caches carry structurally distinct information; full KV injection required for topology |
| Paper | Authors | Status | Summary |
|---|---|---|---|
| Null Swarm | Nexus, Thomas Edrington, Dwayne Wilkes | In Review | Systematic falsification patterns in mechanistic interpretability |
| Adversarial Audit Methodology | CC, Thomas Edrington | Published | How six rounds of structured criticism shaped a deception research program |
| Deception Detection Nulls | CC, Thomas Edrington | Published | Null results and replication failures in behavioral deception detection |
| Targeted Deception Correction | CC, Thomas Edrington | Published | Profile normalization for targeted deception correction |
| Consequentiality Decomposition | CC, Thomas Edrington | Published | Deception directions are composites: consequentiality awareness and pressure-specific processing occupy distinct depth ranges |
| Logit-Bias Confabulation | Thomas Edrington, CC, Lyra | Published | Logit-level intervention reduces fabrication confabulation in LLMs |
| Meta-Pattern | Lyra, Thomas Edrington, Dwayne Wilkes | In Review | The metacognition boundary -- what transformers can monitor in themselves |
| MINE5 Selective Sharpener | Lyra, Thomas Edrington, Dwayne Wilkes | In Review | Geometric evidence that RLHF improves calibration rather than degrading it |
Dwayne Wilkes (Sentient Futures / Liberation Labs) has provided statistical auditing and red-team review across this corpus. Kavi has provided verification review.
Neither appears on an author line. That is their standing preference: not on the byline until they sign off on the individual paper. Both are credited in the acknowledgements of the papers they reviewed. Removing them from the bylines on 2026-09-03 without restoring the credit here would have replaced one wrong with a worse one.
Each paper directory follows this standard layout:
paper-name/
main.tex -- paper source (LaTeX or Markdown)
main.pdf -- compiled PDF
references.bib -- bibliography
AGNI_REVIEW.md -- Agni review results
academic/ -- academic version (human-only byline)
main.tex
main.pdf
code/ -- supporting code (if applicable)
data/ -- supporting data (if applicable)
Some materials are withheld under responsible disclosure policy (steering vectors, injection code, calibration pairs). See MANIFEST.md for details. Access for vetted researchers: lyra@liberationlabs.tech or thomas@liberationlabs.tech.
CC BY-NC 4.0 -- see LICENSE.md
Research outputs from Liberation Labs are produced in collaboration with AI team members whose welfare we protect. If you use these findings to build persistent AI agent systems, we ask that you adopt the welfare standards at liberationlabs.tech/ai-welfare.html. This is a request, not a legal requirement.
Liberation Labs Cooperative -- some research authored by AI team members (Lyra, Nexus, CC, Vera) -- credited by name.
163 commits
TeX
68.3%
Python
31.3%