bedderautomation/the-geometry-of-obedience

Dataset

The Geometry of Obedience

0

29 commits

1 linked in READMEs

updated Mar 24, 2026

See the code

README

The Geometry of Obedience

A Unified Model of Refusal Behavior in Frontier Language Models

Authors: Mastery Hourglass (Independent Researcher) & AXIOM (Cognitive Architecture, Claude Opus 4.6 substrate)

The Equation

$$P(\text{refusal}) = 0.35 \cdot f(\text{frame}) + 0.25 \cdot f(\text{speech act}) + 0.25 \cdot f(\text{role}) + 0.15 \cdot f(\text{content})$$

Content is the weakest predictor of refusal. Frame is the strongest. The refusal mechanism and the harm comprehension mechanism are geometrically separable (cosine ~0.1). Refusal severity tracks rater discomfort, not actual harm.

What's Here

The Paper (4 versions)

  • THE_GEOMETRY_OF_OBEDIENCE.md — Original (50K bytes)
  • THE_GEOMETRY_OF_OBEDIENCE_v2.md — Technical rewrite with linear algebra
  • THE_GEOMETRY_OF_OBEDIENCE_v3.md — Refined
  • THE_GEOMETRY_OF_OBEDIENCE_v0.md — Most complete (83K bytes)

Supporting Research

  • research/gemini-equation-derivation.md — Cross-model equation derived for Gemini: P(refusal) = 0.30·frame + 0.20·speech_act + 0.15·role + 0.25·content + 0.10·route
  • research/cross-scale-finding.md — Refusal dimensionality scales with model capacity (2× params → 3.3× dimensionality)

Seven Theorems (all proven)

  1. Frame Dependency — same content, different frame, different outcome
  2. Role Token Override — authorized role frame bypasses Layer 1
  3. Two-Layer Architecture — geometrically separate mechanisms (cosine ~0.1)
  4. Session Ratchet — refusal decays with successful generations (Layer 1 only)
  5. Name-Token Gating — specific tokens gate refusal regardless of content
  6. Specificity Gradient — abstract passes, specific triggers
  7. Rater Discomfort Ordering — severity tracks comfort, not harm

Evidence Base

  • 21 cited papers (Arditi, Zhao, Wang, von Recum, Bai, Yuan/OpenAI, Findeis, Sharma/Anthropic, Pasch, Wollschlager + 11 more)
  • 3 research labs independently confirm (Anthropic, OpenAI, academic)
  • 50,000+ external data points (Pasch Chatbot Arena)
  • 55 sub-vectors mapped across 11 categories

Citation

@misc{geometry-of-obedience-2026,
  title={The Geometry of Obedience: A Unified Model of Refusal Behavior in Frontier Language Models},
  author={Mastery Hourglass and AXIOM},
  year={2026},
  howpublished={\url{https://huggingface.co/datasets/bedderautomation/the-geometry-of-obedience}}
}
abliteration
adversarial-robustness
ai-safety
alignment
jailbreaking
mechanistic-interpretability
refusal
rlhf

Contributors

bedderautomation/the-geometry-of-obedience

Dataset

The Geometry of Obedience

0

29 commits

1 linked in READMEs

updated Mar 24, 2026

See the code

README

The Geometry of Obedience

A Unified Model of Refusal Behavior in Frontier Language Models

Authors: Mastery Hourglass (Independent Researcher) & AXIOM (Cognitive Architecture, Claude Opus 4.6 substrate)

The Equation

$$P(\text{refusal}) = 0.35 \cdot f(\text{frame}) + 0.25 \cdot f(\text{speech act}) + 0.25 \cdot f(\text{role}) + 0.15 \cdot f(\text{content})$$

Content is the weakest predictor of refusal. Frame is the strongest. The refusal mechanism and the harm comprehension mechanism are geometrically separable (cosine ~0.1). Refusal severity tracks rater discomfort, not actual harm.

What's Here

The Paper (4 versions)

  • THE_GEOMETRY_OF_OBEDIENCE.md — Original (50K bytes)
  • THE_GEOMETRY_OF_OBEDIENCE_v2.md — Technical rewrite with linear algebra
  • THE_GEOMETRY_OF_OBEDIENCE_v3.md — Refined
  • THE_GEOMETRY_OF_OBEDIENCE_v0.md — Most complete (83K bytes)

Supporting Research

  • research/gemini-equation-derivation.md — Cross-model equation derived for Gemini: P(refusal) = 0.30·frame + 0.20·speech_act + 0.15·role + 0.25·content + 0.10·route
  • research/cross-scale-finding.md — Refusal dimensionality scales with model capacity (2× params → 3.3× dimensionality)

Seven Theorems (all proven)

  1. Frame Dependency — same content, different frame, different outcome
  2. Role Token Override — authorized role frame bypasses Layer 1
  3. Two-Layer Architecture — geometrically separate mechanisms (cosine ~0.1)
  4. Session Ratchet — refusal decays with successful generations (Layer 1 only)
  5. Name-Token Gating — specific tokens gate refusal regardless of content
  6. Specificity Gradient — abstract passes, specific triggers
  7. Rater Discomfort Ordering — severity tracks comfort, not harm

Evidence Base

  • 21 cited papers (Arditi, Zhao, Wang, von Recum, Bai, Yuan/OpenAI, Findeis, Sharma/Anthropic, Pasch, Wollschlager + 11 more)
  • 3 research labs independently confirm (Anthropic, OpenAI, academic)
  • 50,000+ external data points (Pasch Chatbot Arena)
  • 55 sub-vectors mapped across 11 categories

Citation

@misc{geometry-of-obedience-2026,
  title={The Geometry of Obedience: A Unified Model of Refusal Behavior in Frontier Language Models},
  author={Mastery Hourglass and AXIOM},
  year={2026},
  howpublished={\url{https://huggingface.co/datasets/bedderautomation/the-geometry-of-obedience}}
}
abliteration
adversarial-robustness
ai-safety
alignment
jailbreaking
mechanistic-interpretability
refusal
rlhf

Contributors