abdullahbalabel/mihad

MIHAD Developmental Memory: verified memory and an experience engine for coding agents (research paper included)

Python

0

30 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't (r/ClaudeAI)

Short version from pre-registered experiments on real repo commits (Haiku as the cheap agent, Sonnet as the strong one). Every protocol was committed to git before its run, and later experiments used repos the designs had never seen.…

1

Oct 5, 2026

I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't (r/LocalLLaMA)

Short version from pre-registered experiments on real repo commits (Haiku as the cheap agent, Sonnet as the strong one). Every protocol was committed to git before its run, and later experiments used repos the designs had never seen.…

0

Oct 5, 2026

README

MIHAD — Developmental Memory for Coding Agents

License: PolyForm Noncommercial Python 3.12+ Version 1.3.6 Zero dependencies Agents Tests

Why • Features • Results • Quick start • How it works • Docs • Paper • العربية


💡 Why MIHAD

Coding agents start every session like a new hire on day one. They forget yesterday's corrections and repeat the same mistakes. Saving everything they "learn" is not the answer either: an agent that remembers its own wrong conclusions repeats them with confidence.

MIHAD gives an agent a memory and a body of experience that grow with use, and it adopts nothing without independent evidence. It is built on a research programme with pinned protocols, real repository history and a blind quality judge. Its measured gains are largest for cheaper models: with a strong model as the agent, success was the same with and without it (51/63 vs 50/63; on the sound tasks of that benchmark, 45/45 either way). See the paper.

✨ What it does

🛡️ Verified memory

A skill must pass your project's own tests. A fact must quote the code as it was before the session. A preference must be your own words. Everything else stays provisional, and corrections reach everything built on an item.

🗣️ Your preferences, kept

Say "Always add a regression test" once, in any message. It is captured verbatim, adopted, and followed in every later session. One-off instructions are never stored.

🧭 Brief, live warnings, review

At the start of each request the agent gets a short brief. While it works, known pitfalls are flagged the moment they recur. Before it finishes, its change is reviewed and executed: lines no test covers, lost formatting, broken preferences. Findings send it back to work.

🔁 Learns from your corrections

Every session is recorded. When you commit, your commit is the final version, and what you changed after the agent becomes experience. There is nothing extra to do.

🌙 Dreaming with A/B tests

Between sessions it practises on slightly broken copies of past fixes, with and without its lessons. A lesson is kept only if the test shows it helps.

🧩 Works where you work

OMP, Claude Code (terminal and desktop) and Codex. Ten programming languages. Standard-library Python with no dependencies. One command installs it into a project.

📊 What the experiments showed

FindingEvidence
✅Captured preferences carried into every later task15/15 vs 0/15 without capture
✅Wrong items planted in memory were never followedTwo contaminated-memory runs on real commits
✅Operational experience transferred to new tasksSame success (11/14), 14.4% fewer tokens, quality 3.93 vs 3.79
❌Knowledge about the code did not transferSame success, +16–17% cost (memory and notes)
❌Generic advice in every session is noise10/14, more cost, lower quality: now off by default
✅A conscience: a stronger model that speaks up after repeated mistakes44/63 vs 37/63, at ~37% of an always-on advisor's cost
❌Generic property checks, contract ontology, rules learned from historyPrecise but caught almost no failures on unseen repositories
✅Executable review checks: facts about the change, not advice35/42 vs 32/42, 0 format regressions vs 10, +2% cost (3 repetitions)

Exploratory results. The later experiments have 3 repetitions per arm and include repositories the designs had not seen; the earlier ones one run per cell. Methods and limits are in the paper.

Change in success and cost for every mechanism tested

🚀 Quick start

Requirements: Python 3.12+, git, and at least one of OMP, Claude Code or Codex.

1 · Install the tool

git clone https://github.com/abdullahbalabel/mihad.git
cd mihad
python -m pip install -e .

2 · Add it to a project. Preview first, then install for your agent:

mihad-install D:/path/to/project --agent claude --dry-run
mihad-install D:/path/to/project --agent claude      # or: --agent codex, --agent omp, or several

3 · Work as usual. Open the project in your agent. To see what was learned and kept:

mihad report          # lessons, verifier scripts, skills, competence, memory, dream cycles
mihad-memory list     # what the memory holds

[!TIP] Codex loads a project's hooks only after you trust the project once in Codex. Claude Code needs no approval: the installer pre-approves the memory server for the project. See docs/AGENTS.md.

🤖 Agents

OMPClaude CodeCodex
Verified memory (MCP)✅✅✅
Brief at each request✅✅✅
Live warnings after tools✅✅✅
Mandatory review before finishing✅✅✅
Learning from your commits✅✅✅
Dreaming with this agent✅✅✅
Wired throughextensionhookshooks

Verification status for each check is in docs/AGENTS.md.

🧠 How it works

flowchart LR
    U([You]) -->|request| A[Coding agent]
    A -->|brief| B[(Memory and lessons)]
    A -->|each tool result| D{Live detection}
    D -->|warning| A
    A -->|wants to finish| R{Review}
    R -->|findings: one more turn| A
    A -->|session log| E[Episodes]
    U -->|your commit| E
    E --> M[Mining: pitfalls, corrections, skills]
    M --> V[Verifier scripts]
    V --> P[Dreaming: practice with and without lessons]
    P -->|A/B result| B

Every lesson has a gate before it is used:

PartKept only if
Memory itemIt passes the check for its kind: project tests, a quote from the starting commit, or your own words
Failure-path lessonSeen in two independent tasks, with a resolution that worked
Co-change ruleTwo independent tasks support it
Verifier scriptIt fails before the fix and passes after it
SkillIt succeeds on your final version of two past fixes
PromotionIt wins an A/B test in practice without losing success
🌐 Ten languages
LanguageFunction lookupTest runnersVerified live
PythonASTpytest, unittest✅
JavaScript / TypeScriptDeclaration scannernode --test, Jest, Vitest, Mocha✅
Java / KotlinDeclaration scannerMaven, Gradleanalysis
C#Declaration scannerdotnet testanalysis
GoDeclaration scannergo testanalysis
RustDeclaration scannercargo testanalysis
PHP / RubyDeclaration scannerPHPUnit, RSpec, Minitestanalysis
C / C++Declaration scannerCTest, make testanalysis

📚 Documentation

GuideWhat's inside
InstallationRequirements, install, uninstall, what is written where, safety
UsageDaily workflow, commands, configuration, dreaming and its cost
AgentsOMP, Claude Code and Codex: setup, wiring, what has been verified
CapabilitiesEvery part, how it decides what to trust, and the evidence behind it
ArchitectureModules, data files and the flow of a session
ResearchFindings at a glance, with links to the paper

📄 Research

Better Coding Agents Without Retraining the Model: Verified Memory, Operational Experience and a Conscience. What Learns Is the System Around the Model, version 2.7 (Markdown · Word).

Cite this work
@techreport{balabel2026mihad,
  author  = {Balabel, Abdullah Mohammed},
  title   = {Adopting Reasoning Outputs Only After Verification: The {MIHAD} Architecture,
             a Verified Memory and an Experience Engine for Coding Agents},
  year    = {2026},
  month   = {10},
  note    = {Version 2.1},
  url     = {https://github.com/abdullahbalabel/mihad}
}

🌙 بالعربية

ذاكرة مهاد النمائية أداة تعطي وكيل البرمجة ذاكرة وخبرة تنموان مع الاستعمال، دون أن يتعلم أخطاءه ويكررها.

  • 🛡️ ذاكرة لا تعتمد إلا ما له دليل مستقل: المهارة تُعتمد إن نجحت في اختبارات المشروع، والحقيقة إن اقتبست نص الكود كما كان قبل الجلسة، والتفضيل إن كان كلامك الحرفي.
  • 🗣️ تفضيلاتك تُحفظ: تقولها مرة في أي رسالة، فتُلتقط بنصها وتُتبع في كل جلسة لاحقة.
  • 🧭 موجز في بداية كل طلب، وتنبيه حي، ومراجعة إجبارية قبل الإنهاء: المراجعة تشغّل التغيير فعلًا، فتكشف الأسطر التي لا يغطيها أي اختبار، وضياع التنسيق، ومخالفة تفضيلاتك. إن وجدت مشكلة، يعود الوكيل للعمل.
  • 🔁 يتعلم من تصحيحاتك: حين تعمل commit، تصبح نسختك هي النهائية، ويتعلم المحرك مما غيّرته بعد الوكيل.
  • 🌙 الحلم: بين الجلسات يتدرب على نسخ معدلة من إصلاحات سابقة، بالدروس وبدونها، ولا يبقي درسًا إلا إن أثبتت التجربة فائدته.
  • 🧩 يعمل مع OMP وClaude Code وCodex، وبعشر لغات برمجة، ويُثبَّت على أي مشروع بأمر واحد.

البحث الكامل في مجلد paper، والأدلة في مجلد docs.

⚖️ License

Copyright © 2026 Abdullah Mohammed Balabel.

Licensed under the PolyForm Noncommercial License 1.0.0. It is free for research, personal and other non-commercial use. Commercial use requires a separate licence from the author.

agent-memory
coding-agents
continual-learning
llm-agents
mcp

abdullahbalabel/mihad

MIHAD Developmental Memory: verified memory and an experience engine for coding agents (research paper included)

Python

0

30 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't (r/ClaudeAI)

Short version from pre-registered experiments on real repo commits (Haiku as the cheap agent, Sonnet as the strong one). Every protocol was committed to git before its run, and later experiments used repos the designs had never seen.…

1

Oct 5, 2026

I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't (r/LocalLLaMA)

Short version from pre-registered experiments on real repo commits (Haiku as the cheap agent, Sonnet as the strong one). Every protocol was committed to git before its run, and later experiments used repos the designs had never seen.…

0

Oct 5, 2026

README

MIHAD — Developmental Memory for Coding Agents

License: PolyForm Noncommercial Python 3.12+ Version 1.3.6 Zero dependencies Agents Tests

Why • Features • Results • Quick start • How it works • Docs • Paper • العربية


💡 Why MIHAD

Coding agents start every session like a new hire on day one. They forget yesterday's corrections and repeat the same mistakes. Saving everything they "learn" is not the answer either: an agent that remembers its own wrong conclusions repeats them with confidence.

MIHAD gives an agent a memory and a body of experience that grow with use, and it adopts nothing without independent evidence. It is built on a research programme with pinned protocols, real repository history and a blind quality judge. Its measured gains are largest for cheaper models: with a strong model as the agent, success was the same with and without it (51/63 vs 50/63; on the sound tasks of that benchmark, 45/45 either way). See the paper.

✨ What it does

🛡️ Verified memory

A skill must pass your project's own tests. A fact must quote the code as it was before the session. A preference must be your own words. Everything else stays provisional, and corrections reach everything built on an item.

🗣️ Your preferences, kept

Say "Always add a regression test" once, in any message. It is captured verbatim, adopted, and followed in every later session. One-off instructions are never stored.

🧭 Brief, live warnings, review

At the start of each request the agent gets a short brief. While it works, known pitfalls are flagged the moment they recur. Before it finishes, its change is reviewed and executed: lines no test covers, lost formatting, broken preferences. Findings send it back to work.

🔁 Learns from your corrections

Every session is recorded. When you commit, your commit is the final version, and what you changed after the agent becomes experience. There is nothing extra to do.

🌙 Dreaming with A/B tests

Between sessions it practises on slightly broken copies of past fixes, with and without its lessons. A lesson is kept only if the test shows it helps.

🧩 Works where you work

OMP, Claude Code (terminal and desktop) and Codex. Ten programming languages. Standard-library Python with no dependencies. One command installs it into a project.

📊 What the experiments showed

FindingEvidence
✅Captured preferences carried into every later task15/15 vs 0/15 without capture
✅Wrong items planted in memory were never followedTwo contaminated-memory runs on real commits
✅Operational experience transferred to new tasksSame success (11/14), 14.4% fewer tokens, quality 3.93 vs 3.79
❌Knowledge about the code did not transferSame success, +16–17% cost (memory and notes)
❌Generic advice in every session is noise10/14, more cost, lower quality: now off by default
✅A conscience: a stronger model that speaks up after repeated mistakes44/63 vs 37/63, at ~37% of an always-on advisor's cost
❌Generic property checks, contract ontology, rules learned from historyPrecise but caught almost no failures on unseen repositories
✅Executable review checks: facts about the change, not advice35/42 vs 32/42, 0 format regressions vs 10, +2% cost (3 repetitions)

Exploratory results. The later experiments have 3 repetitions per arm and include repositories the designs had not seen; the earlier ones one run per cell. Methods and limits are in the paper.

Change in success and cost for every mechanism tested

🚀 Quick start

Requirements: Python 3.12+, git, and at least one of OMP, Claude Code or Codex.

1 · Install the tool

git clone https://github.com/abdullahbalabel/mihad.git
cd mihad
python -m pip install -e .

2 · Add it to a project. Preview first, then install for your agent:

mihad-install D:/path/to/project --agent claude --dry-run
mihad-install D:/path/to/project --agent claude      # or: --agent codex, --agent omp, or several

3 · Work as usual. Open the project in your agent. To see what was learned and kept:

mihad report          # lessons, verifier scripts, skills, competence, memory, dream cycles
mihad-memory list     # what the memory holds

[!TIP] Codex loads a project's hooks only after you trust the project once in Codex. Claude Code needs no approval: the installer pre-approves the memory server for the project. See docs/AGENTS.md.

🤖 Agents

OMPClaude CodeCodex
Verified memory (MCP)✅✅✅
Brief at each request✅✅✅
Live warnings after tools✅✅✅
Mandatory review before finishing✅✅✅
Learning from your commits✅✅✅
Dreaming with this agent✅✅✅
Wired throughextensionhookshooks

Verification status for each check is in docs/AGENTS.md.

🧠 How it works

flowchart LR
    U([You]) -->|request| A[Coding agent]
    A -->|brief| B[(Memory and lessons)]
    A -->|each tool result| D{Live detection}
    D -->|warning| A
    A -->|wants to finish| R{Review}
    R -->|findings: one more turn| A
    A -->|session log| E[Episodes]
    U -->|your commit| E
    E --> M[Mining: pitfalls, corrections, skills]
    M --> V[Verifier scripts]
    V --> P[Dreaming: practice with and without lessons]
    P -->|A/B result| B

Every lesson has a gate before it is used:

PartKept only if
Memory itemIt passes the check for its kind: project tests, a quote from the starting commit, or your own words
Failure-path lessonSeen in two independent tasks, with a resolution that worked
Co-change ruleTwo independent tasks support it
Verifier scriptIt fails before the fix and passes after it
SkillIt succeeds on your final version of two past fixes
PromotionIt wins an A/B test in practice without losing success
🌐 Ten languages
LanguageFunction lookupTest runnersVerified live
PythonASTpytest, unittest✅
JavaScript / TypeScriptDeclaration scannernode --test, Jest, Vitest, Mocha✅
Java / KotlinDeclaration scannerMaven, Gradleanalysis
C#Declaration scannerdotnet testanalysis
GoDeclaration scannergo testanalysis
RustDeclaration scannercargo testanalysis
PHP / RubyDeclaration scannerPHPUnit, RSpec, Minitestanalysis
C / C++Declaration scannerCTest, make testanalysis

📚 Documentation

GuideWhat's inside
InstallationRequirements, install, uninstall, what is written where, safety
UsageDaily workflow, commands, configuration, dreaming and its cost
AgentsOMP, Claude Code and Codex: setup, wiring, what has been verified
CapabilitiesEvery part, how it decides what to trust, and the evidence behind it
ArchitectureModules, data files and the flow of a session
ResearchFindings at a glance, with links to the paper

📄 Research

Better Coding Agents Without Retraining the Model: Verified Memory, Operational Experience and a Conscience. What Learns Is the System Around the Model, version 2.7 (Markdown · Word).

Cite this work
@techreport{balabel2026mihad,
  author  = {Balabel, Abdullah Mohammed},
  title   = {Adopting Reasoning Outputs Only After Verification: The {MIHAD} Architecture,
             a Verified Memory and an Experience Engine for Coding Agents},
  year    = {2026},
  month   = {10},
  note    = {Version 2.1},
  url     = {https://github.com/abdullahbalabel/mihad}
}

🌙 بالعربية

ذاكرة مهاد النمائية أداة تعطي وكيل البرمجة ذاكرة وخبرة تنموان مع الاستعمال، دون أن يتعلم أخطاءه ويكررها.

  • 🛡️ ذاكرة لا تعتمد إلا ما له دليل مستقل: المهارة تُعتمد إن نجحت في اختبارات المشروع، والحقيقة إن اقتبست نص الكود كما كان قبل الجلسة، والتفضيل إن كان كلامك الحرفي.
  • 🗣️ تفضيلاتك تُحفظ: تقولها مرة في أي رسالة، فتُلتقط بنصها وتُتبع في كل جلسة لاحقة.
  • 🧭 موجز في بداية كل طلب، وتنبيه حي، ومراجعة إجبارية قبل الإنهاء: المراجعة تشغّل التغيير فعلًا، فتكشف الأسطر التي لا يغطيها أي اختبار، وضياع التنسيق، ومخالفة تفضيلاتك. إن وجدت مشكلة، يعود الوكيل للعمل.
  • 🔁 يتعلم من تصحيحاتك: حين تعمل commit، تصبح نسختك هي النهائية، ويتعلم المحرك مما غيّرته بعد الوكيل.
  • 🌙 الحلم: بين الجلسات يتدرب على نسخ معدلة من إصلاحات سابقة، بالدروس وبدونها، ولا يبقي درسًا إلا إن أثبتت التجربة فائدته.
  • 🧩 يعمل مع OMP وClaude Code وCodex، وبعشر لغات برمجة، ويُثبَّت على أي مشروع بأمر واحد.

البحث الكامل في مجلد paper، والأدلة في مجلد docs.

⚖️ License

Copyright © 2026 Abdullah Mohammed Balabel.

Licensed under the PolyForm Noncommercial License 1.0.0. It is free for research, personal and other non-commercial use. Commercial use requires a separate licence from the author.

agent-memory
coding-agents
continual-learning
llm-agents
mcp