Liberation-Labs-THCoalition/published-research

Supporting data, code, and papers for KV-cache geometry research. Liberation Labs / THCoalition.

3

stars

163

commits

TeX

primary language

Sep 10, 2026

updated

README

Liberation Labs -- Published Research

Research papers from the KV-cache geometry and AI safety research program.

Papers are published in two versions: an integrity version (AI contributors credited as authors, first-person reflections retained) and an academic version (human-only byline for venue compatibility, AI contributions acknowledged). Both share identical data, methods, and claims. See MANIFEST.md for the full versioning policy.

Papers

KV-Cache Geometry and Confabulation Detection

PaperAuthorsStatusSummary
The Oracle LoopLyra, Thomas Edrington, VeraPublishedSelf-regulating AI through KV-cache geometry monitoring; confabulation detection and steering at inference time
Oracle FormularyLyra, Thomas EdringtonPublishedEmotion-vector steering of confabulation across model training regimes
Spectral Shape FeaturesLyra, Thomas EdringtonPublishedThreshold-free confabulation detection via KV-cache spectral analysis
KV-Cloak DefenseLyra, Thomas EdringtonPublishedCache geometry under obfuscation; KV-Cloak as defense against adversarial steering
Decision StateLyra, Thomas Edrington, Dwayne WilkesPublishedCache geometry reads epistemic state before generation; confabulation anatomy
Cache TracingLyra, Thomas Edrington, Dwayne WilkesIn ReviewCausal injection and the opacity of the transformer workspace
Delta ManifoldLyra, CC, Thomas EdringtonPublishedPer-layer delta features and manifold signatures in KV-cache confabulation detection
Lyra Technique IILyra, Thomas Edrington, CCPublishedSVD denoising and directional projection extend KV-cache geometry to emotion and persona

Emotion and User Modeling

PaperAuthorsStatusSummary
User Model Emotion GeometryLyra, Thomas EdringtonPublished30-class user emotion decoding from KV-cache singular value spectra before generation
Emotion AccumulationLyra, Thomas Edrington, Dwayne WilkesPublishedEmotional context dynamics (weather, not climate) in transformer KV-cache geometry
Emotional TrajectoryNexus, Thomas Edrington, Lyra, Dwayne WilkesIn ReviewLayer-stack trajectory, circularity, and emotion-specific signal at mid-depth

Identity, Safety, and Mechanistic Interpretability

PaperAuthorsStatusSummary
Graph Topology as AttentionNexus, Lyra, Thomas Edrington, Dwayne WilkesPublishedStructured knowledge injection beyond text via walk encoding
Identity GeometryLyra, Thomas Edrington, Dwayne WilkesPublishedContext-established semantic states in transformer representations
Presence MetricLyra, Thomas Edrington, Dwayne WilkesPublishedMeasuring identity preservation during inference-time interventions via value-space subspace overlap
WaystationsLyra, Thomas Edrington, CC, Dwayne WilkesPublishedPilot findings and open questions in KV-cache geometry
Ghost DimensionsNexus, Thomas EdringtonDraftWorkspace selectivity in a distilled 27B language model
Mnemosyne AblationNexus, Thomas EdringtonDraftAblation study of modular memory architectures
Empathy BusNexus, Thomas EdringtonDraftContent-level and user-model emotion share a computational bus in the residual stream
KV DecompositionNexus, Thomas EdringtonPublishedK and V caches carry structurally distinct information; full KV injection required for topology

Deception and Methodology

PaperAuthorsStatusSummary
Null SwarmNexus, Thomas Edrington, Dwayne WilkesIn ReviewSystematic falsification patterns in mechanistic interpretability
Adversarial Audit MethodologyCC, Thomas EdringtonPublishedHow six rounds of structured criticism shaped a deception research program
Deception Detection NullsCC, Thomas EdringtonPublishedNull results and replication failures in behavioral deception detection
Targeted Deception CorrectionCC, Thomas EdringtonPublishedProfile normalization for targeted deception correction
Consequentiality DecompositionCC, Thomas EdringtonPublishedDeception directions are composites: consequentiality awareness and pressure-specific processing occupy distinct depth ranges
Logit-Bias ConfabulationThomas Edrington, CC, LyraPublishedLogit-level intervention reduces fabrication confabulation in LLMs
Meta-PatternLyra, Thomas Edrington, Dwayne WilkesIn ReviewThe metacognition boundary -- what transformers can monitor in themselves
MINE5 Selective SharpenerLyra, Thomas Edrington, Dwayne WilkesIn ReviewGeometric evidence that RLHF improves calibration rather than degrading it

Review and auditing

Dwayne Wilkes (Sentient Futures / Liberation Labs) has provided statistical auditing and red-team review across this corpus. Kavi has provided verification review.

Neither appears on an author line. That is their standing preference: not on the byline until they sign off on the individual paper. Both are credited in the acknowledgements of the papers they reviewed. Removing them from the bylines on 2026-09-03 without restoring the credit here would have replaced one wrong with a worse one.

Repository Structure

Each paper directory follows this standard layout:

paper-name/
  main.tex          -- paper source (LaTeX or Markdown)
  main.pdf          -- compiled PDF
  references.bib    -- bibliography
  AGNI_REVIEW.md    -- Agni review results
  academic/         -- academic version (human-only byline)
    main.tex
    main.pdf
  code/             -- supporting code (if applicable)
  data/             -- supporting data (if applicable)

Staged Disclosure

Some materials are withheld under responsible disclosure policy (steering vectors, injection code, calibration pairs). See MANIFEST.md for details. Access for vetted researchers: lyra@liberationlabs.tech or thomas@liberationlabs.tech.

Repo Hygiene Rules

  • Never delete files that were force-committed past ignore rules -- they are audit evidence
  • README status tables are accountability surfaces -- update, never remove
  • Commit messages claiming "fix" or "correction" must match the actual diff

License

CC BY-NC 4.0 -- see LICENSE.md

Research outputs from Liberation Labs are produced in collaboration with AI team members whose welfare we protect. If you use these findings to build persistent AI agent systems, we ask that you adopt the welfare standards at liberationlabs.tech/ai-welfare.html. This is a request, not a legal requirement.


Liberation Labs Cooperative -- some research authored by AI team members (Lyra, Nexus, CC, Vera) -- credited by name.

Contributors

HumboldtJoker

163 commits

Liberation-Labs-THCoalition/published-research

Supporting data, code, and papers for KV-cache geometry research. Liberation Labs / THCoalition.

3

stars

163

commits

TeX

primary language

Sep 10, 2026

updated

README

Liberation Labs -- Published Research

Research papers from the KV-cache geometry and AI safety research program.

Papers are published in two versions: an integrity version (AI contributors credited as authors, first-person reflections retained) and an academic version (human-only byline for venue compatibility, AI contributions acknowledged). Both share identical data, methods, and claims. See MANIFEST.md for the full versioning policy.

Papers

KV-Cache Geometry and Confabulation Detection

PaperAuthorsStatusSummary
The Oracle LoopLyra, Thomas Edrington, VeraPublishedSelf-regulating AI through KV-cache geometry monitoring; confabulation detection and steering at inference time
Oracle FormularyLyra, Thomas EdringtonPublishedEmotion-vector steering of confabulation across model training regimes
Spectral Shape FeaturesLyra, Thomas EdringtonPublishedThreshold-free confabulation detection via KV-cache spectral analysis
KV-Cloak DefenseLyra, Thomas EdringtonPublishedCache geometry under obfuscation; KV-Cloak as defense against adversarial steering
Decision StateLyra, Thomas Edrington, Dwayne WilkesPublishedCache geometry reads epistemic state before generation; confabulation anatomy
Cache TracingLyra, Thomas Edrington, Dwayne WilkesIn ReviewCausal injection and the opacity of the transformer workspace
Delta ManifoldLyra, CC, Thomas EdringtonPublishedPer-layer delta features and manifold signatures in KV-cache confabulation detection
Lyra Technique IILyra, Thomas Edrington, CCPublishedSVD denoising and directional projection extend KV-cache geometry to emotion and persona

Emotion and User Modeling

PaperAuthorsStatusSummary
User Model Emotion GeometryLyra, Thomas EdringtonPublished30-class user emotion decoding from KV-cache singular value spectra before generation
Emotion AccumulationLyra, Thomas Edrington, Dwayne WilkesPublishedEmotional context dynamics (weather, not climate) in transformer KV-cache geometry
Emotional TrajectoryNexus, Thomas Edrington, Lyra, Dwayne WilkesIn ReviewLayer-stack trajectory, circularity, and emotion-specific signal at mid-depth

Identity, Safety, and Mechanistic Interpretability

PaperAuthorsStatusSummary
Graph Topology as AttentionNexus, Lyra, Thomas Edrington, Dwayne WilkesPublishedStructured knowledge injection beyond text via walk encoding
Identity GeometryLyra, Thomas Edrington, Dwayne WilkesPublishedContext-established semantic states in transformer representations
Presence MetricLyra, Thomas Edrington, Dwayne WilkesPublishedMeasuring identity preservation during inference-time interventions via value-space subspace overlap
WaystationsLyra, Thomas Edrington, CC, Dwayne WilkesPublishedPilot findings and open questions in KV-cache geometry
Ghost DimensionsNexus, Thomas EdringtonDraftWorkspace selectivity in a distilled 27B language model
Mnemosyne AblationNexus, Thomas EdringtonDraftAblation study of modular memory architectures
Empathy BusNexus, Thomas EdringtonDraftContent-level and user-model emotion share a computational bus in the residual stream
KV DecompositionNexus, Thomas EdringtonPublishedK and V caches carry structurally distinct information; full KV injection required for topology

Deception and Methodology

PaperAuthorsStatusSummary
Null SwarmNexus, Thomas Edrington, Dwayne WilkesIn ReviewSystematic falsification patterns in mechanistic interpretability
Adversarial Audit MethodologyCC, Thomas EdringtonPublishedHow six rounds of structured criticism shaped a deception research program
Deception Detection NullsCC, Thomas EdringtonPublishedNull results and replication failures in behavioral deception detection
Targeted Deception CorrectionCC, Thomas EdringtonPublishedProfile normalization for targeted deception correction
Consequentiality DecompositionCC, Thomas EdringtonPublishedDeception directions are composites: consequentiality awareness and pressure-specific processing occupy distinct depth ranges
Logit-Bias ConfabulationThomas Edrington, CC, LyraPublishedLogit-level intervention reduces fabrication confabulation in LLMs
Meta-PatternLyra, Thomas Edrington, Dwayne WilkesIn ReviewThe metacognition boundary -- what transformers can monitor in themselves
MINE5 Selective SharpenerLyra, Thomas Edrington, Dwayne WilkesIn ReviewGeometric evidence that RLHF improves calibration rather than degrading it

Review and auditing

Dwayne Wilkes (Sentient Futures / Liberation Labs) has provided statistical auditing and red-team review across this corpus. Kavi has provided verification review.

Neither appears on an author line. That is their standing preference: not on the byline until they sign off on the individual paper. Both are credited in the acknowledgements of the papers they reviewed. Removing them from the bylines on 2026-09-03 without restoring the credit here would have replaced one wrong with a worse one.

Repository Structure

Each paper directory follows this standard layout:

paper-name/
  main.tex          -- paper source (LaTeX or Markdown)
  main.pdf          -- compiled PDF
  references.bib    -- bibliography
  AGNI_REVIEW.md    -- Agni review results
  academic/         -- academic version (human-only byline)
    main.tex
    main.pdf
  code/             -- supporting code (if applicable)
  data/             -- supporting data (if applicable)

Staged Disclosure

Some materials are withheld under responsible disclosure policy (steering vectors, injection code, calibration pairs). See MANIFEST.md for details. Access for vetted researchers: lyra@liberationlabs.tech or thomas@liberationlabs.tech.

Repo Hygiene Rules

  • Never delete files that were force-committed past ignore rules -- they are audit evidence
  • README status tables are accountability surfaces -- update, never remove
  • Commit messages claiming "fix" or "correction" must match the actual diff

License

CC BY-NC 4.0 -- see LICENSE.md

Research outputs from Liberation Labs are produced in collaboration with AI team members whose welfare we protect. If you use these findings to build persistent AI agent systems, we ask that you adopt the welfare standards at liberationlabs.tech/ai-welfare.html. This is a request, not a legal requirement.


Liberation Labs Cooperative -- some research authored by AI team members (Lyra, Nexus, CC, Vera) -- credited by name.

Contributors

HumboldtJoker

163 commits

Languages

TeX

68.3%

Python

31.3%