BrainStem is a biologically inspired, Real Neuro-Symbolic (RNS-AI) cognitive architecture for lifelong learning. It is designed to learn models of the structures and dynamics of language and text through context hypotheses, uncertainty, contradiction, revision, neuromodulation, replay, and consolidation rather than by merely storing isolated facts.
Python
37
397 commits
updated Sep 28, 2026
BrainStem is a biologically inspired, Real Neuro-Symbolic (RNS-AI) cognitive architecture for lifelong learning. It is designed to learn models of the structures and dynamics of language and text through context hypotheses, uncertainty, contradiction, revision, neuromodulation, replay, and consolidation rather than by merely storing isolated facts. A second, character-level observation layer discovers word boundaries directly from unsegmented text, without any predefined notion of "word." Since 25–28 September 2026, a third, fourth, and fifth observation layer additionally discover relations between already-observed elements, categories emerging from the resulting relation graph, and durable questions emerging from persistent, unresolved information gaps.
One CPU Core / No GPU
[!IMPORTANT] BrainStem is a research and calibration system, not a production-ready assistant. As of the current project stage, all previously closed productive write paths (facts, relations, ontology categories, questions, fact/relation/ontology/question promotion, gap/contradiction/revision writes) have been deliberately opened as an explicitly framed, ongoing experiment, with a full project backup taken beforehand as a fallback point.
YouTube - AI conversation about BrainStem ProjectThe system has completed a full read-through of its current two-source corpus (German Wikipedia Physics and Computer categories, 167,661 chunks, 100%) and has since run well beyond 11,500 real learning cycles in both active-learning and replay/consolidation-only modes without data loss, corruption, or GUI failure. A dedicated 1,500-cycle Etappe-A drift validation (sensory-deprivation mode, input disabled, inner dynamics only) completed with an overall result of "konvergiert": all 20 evaluated signals were classified either stabil or konvergiert, with zero signals flagged as divergent.
Cortisol Stage 2 is active and functionally verified. stage=2 is set in the production database. Under real production conditions it has not been triggered live (allostatic_load has stayed well below the 0.6 threshold — a sign of a consistently calm system, not a defect). Its guarded intervention logic (per-value cap 0.01, per-cycle budget 0.03, 3-cycle cooldown) has been independently confirmed correct under simulated stress.
Hypothesis graduation (uncertain_hypothesis → stable_hypothesis, and — since the Relations/Ontology work below — uncertain_lexical_boundary → stable_lexical_boundary and uncertain_relation_hypothesis → stable_relation_hypothesis) is active, gated by the consolidation-survival criterion (≥3 confirmed Phase-7d survival cycles), an available critic-gate cross-check, warm-up dampening, a budget of one graduation per cycle, and — new — neuromodulator-coupled divisive-normalization role selection among the three competing roles (see below).
As of 25–28 September 2026, the complete Relations/Ontology/Questions emergence chain (originally planned as three slices) is fully implemented, wired into the runtime phase chain, and verified through repeated real, multi-hundred-cycle end-to-end runs. The runtime phase registry now holds 39 entries (previously 34). Additionally, eleven of the pipeline's own calibrated thresholds and budgets are now directly coupled to one or more of the six core digital neuromodulators, each grounded in a specific piece of neuroscience literature (see Neuromodulator-Coupled Thresholds).
Direct Facts, Relations, Ontology, and Questions writes, Fact/Relation/Ontology/Question promotion, Attention writes, and productive Phase-5f/5g/5i experiments are all currently active, alongside the full nine-module Stage-B chain described below.
A Phase 0 module (v8_phase0_lexical_boundary_observation_release) is registered at the front of the runtime phase chain, running alongside the existing sentence-level context_observation_learning entry point. It reads the raw, unsegmented character stream of already-imported chunks, maintains a pure frequency table of "which character follows this preceding context of length k," and computes the local branching entropy at each character position:
H(context) = − Σ P(next_char | context) · log2 P(next_char | context)
A pronounced spike in this entropy at a given position is treated as a boundary candidate and creates or re-observes a context_hypotheses row using role='uncertain_lexical_boundary' — reusing exactly the same insert-or-reobserve mechanism, consolidation path (Phase 7d), Stage-B graduation, and fact promotion already used for every other hypothesis. Once a lexical-boundary candidate's own evidence_count clears a small threshold, it additionally becomes eligible as an input element for the Relations emergence layer below — deliberately before full Stage-B graduation, a design decision grounded in evidence that word segmentation and relational learning proceed in parallel in human learners rather than sequentially (see Academic References).
Three additional observation/promotion layers, completed this session, extend the system beyond isolated sentence- and word-level hypotheses toward structured, relational, categorical, and curiosity-driven knowledge:
Relations (Modul A). v8_phase0b_relational_binding_observation_release binds pairs of already-stable elements that co-occur within the same sentence into a candidate relation, via a strict two-stage statistical process: a symmetric Pointwise Mutual Information existence gate (PMI(A;B) = log2(P(A,B)/(P(A)·P(B)))), followed — only once enough evidence exists — by a purely positional direction signal, calibrated separately per corpus language (German requires a stronger positional skew than English, reflecting German's weaker "position implies grammatical role" signal). v8_stageb_relation_promotion_release promotes graduated relation hypotheses into the relations table, fully retractable on later revision, exactly like fact promotion.
Ontology (Modul B). v8_stageb_ontology_cluster_observation_release builds an undirected graph over the relations table and clusters it via Label Propagation, anchoring cluster identity across cycles to each cluster's own centrality-based prototype (Rosch prototype theory) rather than to Label Propagation's own non-deterministic per-cycle label. v8_stageb_ontology_promotion_release promotes reconfirmed-stable clusters into the ontology table as child → parent rows, with the fixed, honest placeholder relation "is_a" — the system has structurally identified a graded group-membership relationship via connectivity and centrality, without parsing or inferring any actual linguistic category name.
Questions (Modul C). v8_stageb_question_promotion_release promotes a persistent, repeatedly reconfirmed entry from the existing internal_learning_gaps table into the questions table once it clears a calibrated persistence threshold — directly implementing Loewenstein's Information-Gap Theory of curiosity. Question text is never a generated sentence, only a fixed, non-linguistic marker followed by the underlying hypothesis's own verbatim, already-observed surface text. v8_stageb_question_chunk_feedback_release closes the loop back into reading behavior: for every still-open question, it boosts both the exact source chunk where the underlying gap was first observed and searches the full corpus via SQLite FTS5 for other, not-yet-read chunks that might resolve it — the first and only module in the codebase that reads questions to influence real system behavior.
Habituation. Gap detection now additionally retires a persistent-but-unproductive gap (no genuine evidence growth for 64 consecutive reconfirmation cycles) into a reversible habituated state, implementing the empirically established inverted-U relationship between resolvability and curiosity.
Nine modules now form a single, ordered, same-cycle-visible chain per real learning cycle:
facts, fully retractable on later revision.The chain is ordered so that a contradiction resolved this cycle already causes its corresponding fact/relation to be retracted within the same cycle, and a relation/cluster/question promoted this cycle is already visible to the next module in the same chain.
Beyond the pre-existing six-core neuromodulator engine, eleven of the pipeline's own calibrated thresholds and budgets are now directly modulated by one or more of the six core messengers, each using the same symmetric, self-regulating gain (exactly the unmodulated calibrated value at the messenger's own neutral value of 0.5):
| Coupling | Messenger(s) | Grounding |
|---|---|---|
| Gap detection existence gate | Noradrenaline | Aston-Jones & Cohen (2005) adaptive gain theory |
| Gap habituation threshold | Serotonin | Grossman, Bari & Cohen (2022); Hochner et al. (1986) |
| Relational binding existence gate | Noradrenaline | Aston-Jones & Cohen (2005) |
| Relational binding direction threshold | Acetylcholine | Project-internal structural-revision role |
| Hypothesis revision budget | Acetylcholine | Project-internal structural-revision role |
| Contradiction resolution ratio | Noradrenaline | Aston-Jones & Cohen (2005) |
| Graduation role-selection pressure | Acetylcholine | Douchamps et al. (2013); Gómez-Ocádiz et al. (2022) |
| Graduation divisive-normalization temperature | GABA | Katzner, Busse & Carandini (2011) |
| Ontology stability streak | Serotonin | Grossman, Bari & Cohen (2022) |
| Ontology minimum cluster size | Noradrenaline (self-regulating direction) | Aston-Jones & Cohen (2005) vs. Shine et al. (2018) |
| Ontology overlap threshold | Dopamine | Kahnt & Tobler (2016); Novicky et al. (2023) |
| Ontology promotion budget | Acetylcholine | Project-internal structural-revision role |
| Question existence gate | Noradrenaline | Aston-Jones & Cohen (2005) |
| Question priority | Dopamine, Acetylcholine | Established gap-closure/curiosity roles |
| Question-chunk feedback boost magnitude | (inherits question priority) | Consistency with the upstream signal |
| Phase 7d consolidation gate | Acetylcholine | Gais & Born (2004); Hasselmo & McGaughy (2004) |
| Stage | Status | Gating condition to proceed |
|---|---|---|
| **Stage A — Core stability** | Complete: 100% corpus completion, 11,500+ real cycles, clean 1,500-cycle formal drift run ("konvergiert", zero divergent signals) | Complete |
| **Stage B — Guarded graduation** | Active: Cortisol Stage 2 applied, hypothesis graduation active for all three eligible roles, neuromodulator-coupled role selection | Ongoing observation |
| **Autonomous Lexical Emergence (Phase 0)** | Active, feeding both fact promotion and the Relations layer below once its own evidence threshold clears | Calibration against real corpus ongoing |
| **Relations Emergence (Modul A)** | Active and productive — Gate-and-Direction binding, promotion, retraction all verified over real multi-cycle runs | Ongoing observation at real production scale |
| **Ontology Emergence (Modul B)** | Active and productive — Label Propagation clustering, prototype-anchored stability, promotion/retraction, three neuromodulator couplings (one self-regulating) all verified | Ongoing observation at real production scale |
| **Questions Emergence (Modul C)** | Active and productive — persistence-gated promotion, habituation-aware retraction, full-corpus chunk feedback all verified | Ongoing observation at real production scale |
| **Neuromodulator coupling of pipeline thresholds** | Eleven couplings active across seven modules, each individually literature-grounded and baseline-invariance-verified | Ongoing; re-calibration against real production corpus pending |
| **Opening the productive write locks** (Facts / Relations / Ontology / Questions / all four Promotion paths) | Open (experimental project position) | Ongoing observation of the full chain; every promoted artifact remains traceable to, and retractable from, its source hypothesis/gap |
| **Full corpus scaling** (complete German Wikipedia) | Deliberately deferred | Order-of-magnitude estimate only, not a commitment |
| **Vector database evaluation** | Deliberately deferred per project's own architecture-checkpoint rule | Only once stable hypothesis identities exist, a concrete semantic-retrieval use case is identified, requirements are measurable, and a read-only/shadow comparison against the SQLite baseline is performed |
| **Symbolic reasoning plugin** (deterministic, non-LLM rule engine) | Concept documented, not scheduled | Gated behind Stage-B graduation and the write-lock roadmap; a validated symbolic rule becoming productive is itself a new class of productive write and must be gated at least as strictly as fact promotion |
| **Multi-core / multi-process learning** | Explicitly out of scope for now | Not currently planned |
Traditional semantic systems often focus on the what: storing and retrieving content. BrainStem focuses on the how: learning how context, uncertainty, evidence, contradiction, revision, and consolidation interact over time — extended to a character-level layer that learns how recurring units emerge from raw text, a relational layer that learns how those units connect to each other, a categorical layer that learns how connected units group into categories, and a curiosity layer that learns which of its own unresolved gaps are durable enough to motivate further reading. A corpus is treated as training substrate rather than as a static knowledge base.
Core principles:
BrainStem is an autonomous software architecture designed for continuous, self-improving data processing and knowledge management. At its core, the system operates through an Autonomous Loop that orchestrates a chain of 39 learning phases to ingest, analyze, and refine information without manual intervention.
The biological terminology used throughout the project's technical documentation, including terms such as "neuromodulators," "sleep," or "homeostasis," is not decorative. These labels are functional designators for mathematical state variables and algorithmic control mechanisms. The values are floats, not molecules. The behavior is biologically inspired, but the implementation is strictly mathematical.
The system's primary mechanics include:
Dynamic Steering Variables: Digital messenger substances are dynamic meta-parameters that adjust the system's learning rate, error weighting, and exploration strategies in real time, and — as of this session — directly gate eleven specific thresholds and budgets throughout the emergence pipeline.
Active versus Offline Processing: The system cycles between active ingestion and an offline optimization phase ("sleep") in which recorded hypotheses are re-evaluated through batch replay and consolidation.
Character-Level Emergence: Alongside sentence-level hypotheses, a branching-entropy-driven process observes the raw character stream to propose, and — through the same consolidation/graduation machinery — stabilize, candidate word-boundary units.
Relational, Categorical, and Curiosity-Driven Emergence: Beyond individual hypotheses, the system now binds pairs of stable elements into relations via a statistically principled existence-and-direction gate, clusters the resulting relation graph into emergent categories anchored by centrality rather than arbitrary labels, and promotes its own persistent, unresolved information gaps into durable questions that measurably redirect its own reading attention.
Knowledge Distillation: By comparing new data against existing stable records, the system filters out inconsistencies and promotes reliable information into its long-term fact/relation/ontology store — while retaining the ability to retract any of them if the underlying hypothesis is revised.
Equilibrium Control: Stability monitoring routines act as a feedback mechanism that pulls meta-parameters back into a functional range when the system detects a performance plateau or excessive variance.
Adaptive Boundaries: The limits within which the system operates are not hardcoded but self-regulating, expanding or contracting processing thresholds based on the complexity of the data encountered — and, for at least one specific coupling, self-regulating even in which direction a neuromodulator should push a threshold, learned from real outcomes rather than assumed.
In summary, the project is a recursive learning engine that uses biologically derived control logic to implement a highly flexible, self-governing system for automated knowledge acquisition, now spanning word-, relation-, category-, and question-level structure discovery.
BrainStem does not operate as a continuously coupled system of differential equations. Instead, it traverses a cyclic state graph: each phase activates at most 2–3 dominant neuromodulators, while the remainder are kept inactive or passive. This sequential architecture prevents interaction cascades and enables deterministic debugging.
| Stage | Name | Description |
|---|---|---|
| 1 | Inference-free pre-parsing | A raw corpus such as a Wikipedia ZIM file is extracted, structured, and partitioned into the chunk store before autonomous learning begins. |
| 2 | Autonomous learning | AutonomousLoop processes prepared chunks while the neuromodulatory, consolidation, and full write-path chain (facts, relations, ontology, questions) reacts to the evolving internal state. |
BrainStem currently uses 12 digital neuromodulators. Their values are normalized to [0.0, 1.0] and derived from internal system state under bounded, biologically inspired dynamics and homeostatic constraints.
| Neuromodulator | Current engineering role |
|---|---|
| Dopamine | outcome and gap-closure signal; now also couples question priority and the ontology layer's overlap tolerance |
| Serotonin | consolidation and stability signal; now also couples gap habituation resistance and ontology cluster-stability requirements |
| Glutamate | excitatory drive associated with exploration and learning activity |
| GABA | global inhibition and E/I-balance signal; now also couples the sharpness of divisive-normalization role competition at graduation |
| Noradrenaline | error, alarm, and persistent-pressure signal; now also couples five separate existence/resolution gates across the pipeline, one of them (ontology minimum cluster size) with a self-regulating, outcome-learned direction |
| Acetylcholine | novelty, attention, and structural-revision signal; now also couples relation direction commitment, hypothesis revision budget, graduation role-novelty pressure, ontology promotion budget, and Phase 7d consolidation gating |
| Adenosine | sleep-pressure homeostat |
| Endocannabinoids | retrograde gain control |
| Cortisol | top-level stability watcher and guarded soft regulator (Stage 2 active) |
| Histamine | wake and arousal signal |
| Orexin | reading-endurance and curiosity-related drive |
| BDNF | activity-dependent growth and consolidation substrate |
Phase 6a performs offline-style replay after the wake path. Phase 6b evaluates replay effectiveness and plasticity adjustments. The critic gate checks whether proposed changes remain consistent enough to be retained; rejected or unstable material remains available as error and revision evidence. The same critic gate is reused, unmodified, by Stage-B hypothesis graduation.
Phase 7d adds sub-1-Hz up/down-state processing with stochastic reactivation, adaptive thresholds, activity-dependent participation, survivor and weakening statistics, anchor interleaving, and self-regulating down-selection. Sentence-level, lexical-boundary, and relation hypotheses all share this same candidate pool, distinguished only by their role value. Consolidation activity is now additionally gated by acetylcholine, reflecting the causal, pharmacologically demonstrated role of low cholinergic tone during slow-wave sleep in permitting declarative memory consolidation.
As of the current experimental project position, the following paths are open:
BrainStem uses ki_memory.sqlite3 in the project root. The database is created automatically when absent. Schema rules: schema changes must be reflected in the bootstrap in the same delivery; ensure_schema must be idempotent; _self_check_schema must run before writes; every written column must already be declared in SCHEMA_TABLES; compile checks, smoke tests, and intermediate checks are required before delivery; structural changes require a full backup first.
Learning state can be reset without re-importing the corpus. Preserved content includes documents, chunks, FTS data, import state, and configuration. The reset workflow performs a dry run and creates a timestamped database backup before applying changes.
From the project root: python main.py --gui
| Step | GUI action | Purpose |
|---|---|---|
| 1 | Export / Configuration | Configure the maximum number of articles before import. |
| 2 | Import & Jobs → ZIM Einlesen | Extract and pre-parse the corpus. |
| 3 | Import & Jobs → Autonom dauerhaft starten | Start autonomous learning. |
| 4 | Import & Jobs → Autonom stoppen | Stop autonomous learning cooperatively. |
| 5 | Close the GUI normally | Wait for active workers instead of terminating them abruptly. |
This project builds upon concepts, algorithms, and theoretical frameworks established in the following academic literature:
Word segmentation and statistical language acquisition
Curiosity, information gaps, and question formation
Habituation, dishabituation, and information-gain-driven decay
Noradrenaline, arousal, and network topology
Dopamine, generalization, and precision
Acetylcholine, novelty, and encoding-versus-retrieval balance
Serotonin and meta-learning under uncertainty
GABA, gain, and stimulus selectivity
Divisive normalization
Category and prototype formation
Complementary learning systems and structure extraction
Sleep-dependent consolidation, homeostatic plasticity, and previously referenced principles (unchanged)
@article{aston-jones2005integrative,
title={An Integrative Theory of Locus Coeruleus-Norepinephrine Function: Adaptive Gain and Optimal Performance},
author={Aston-Jones, Gary and Cohen, Jonathan D.},
journal={Annual Review of Neuroscience},
year={2005}
}
@article{kahnt2016dopamine,
title={Dopamine regulates stimulus generalization in the human hippocampus},
author={Kahnt, Thorsten and Tobler, Philippe N.},
journal={eLife},
year={2016}
}
@article{grossman2022serotonin,
title={Serotonin neurons modulate learning rate through uncertainty},
author={Grossman, Cooper D. and Bari, Bilal A. and Cohen, Jeremiah Y.},
journal={Current Biology},
year={2022}
}
@article{katzner2011gaba,
title={GABA\_A Inhibition Controls Response Gain in Visual Cortex},
author={Katzner, Steffen and Busse, Laura and Carandini, Matteo},
journal={Journal of Neuroscience},
year={2011}
}
@article{carandini2012normalization,
title={Normalization as a canonical neural computation},
author={Carandini, Matteo and Heeger, David J.},
journal={Nature Reviews Neuroscience},
year={2012}
}
@article{loewenstein1994psychology,
title={The Psychology of Curiosity: A Review and Reinterpretation},
author={Loewenstein, George},
journal={Psychological Bulletin},
year={1994}
}
@article{shine2018modulation,
title={The modulation of neural gain facilitates a transition between functional segregation and integration in the brain},
author={Shine, James M. and Aburn, Matthew J. and Breakspear, Michael and Poldrack, Russell A.},
journal={eLife},
year={2018}
}
@article{hamilton2020graph,
title={Graph Representation Learning},
author={Hamilton, William L.},
year={2020},
publisher={Morgan \& Claypool Publishers}
}
@article{watkins2020using,
title={Using Sinusoidally-Modulated Noise as a Surrogate for Slow-Wave Sleep to Accomplish Stable Unsupervised Dictionary Learning in a Spike-Based Sparse Coding Model},
author={Watkins, Yijing and Kim, Edward and Kenyon, Garrett T.},
journal={Frontiers in Computational Neuroscience},
year={2020}
}
@article{tadros2022biologically,
title={Biologically Inspired Sleep Algorithm for Reducing Catastrophic Forgetting in Neural Networks},
author={Tadros, Timothy and Tran, Gia-Bao M. and Krishnan, Giri P. and Bazhenov, Maxim},
year={2022}
}
@article{fischbacher2020intelligent,
title={Intelligent Matrix Exponentiation},
author={Fischbacher, Thomas and Comsa, Iulia M. and Potempa, Krzysztof and Firsching, Moritz and Versari, Luca and Alakuijala, Jyrki},
journal={arXiv preprint arXiv:2008.03926},
year={2020}
}
@article{butz2013homeostatic,
title={Homeostatic structural plasticity--a key to neuronal network formation and repair},
author={Butz, Markus and van Ooyen, Arjen},
journal={PLoS Computational Biology},
year={2013}
}
@article{lee2019mechanisms,
title={Mechanisms of Homeostatic Synaptic Plasticity In Vivo},
author={Lee, Kea-Joo K. and Kirkwood, Alfredo},
journal={Frontiers in Cellular Neuroscience},
year={2019}
}
@article{parker2020nonlinear,
title={Nonlinear Time Series Classification Using Bispectrum-based Deep Convolutional Neural Networks},
author={Parker, Paul A. and Holan, Scott H. and Ravishanker, Nalini},
journal={arXiv preprint arXiv:2003.02353},
year={2020}
}
@article{roncevic2023molecule,
title={Supplementary Materials for A molecule with half-M{\"o}bius topology},
author={Ron{\v{c}}evi{\'c}, Igor and others},
journal={Nature Chemistry},
year={2023}
}
BrainStem is an experimental cognitive-architecture research project. Biological terminology is used as an engineering analogy and design inspiration. The software is not a biological simulation and does not claim neuroscientific equivalence.
Python
100.0%
BrainStem is a biologically inspired, Real Neuro-Symbolic (RNS-AI) cognitive architecture for lifelong learning. It is designed to learn models of the structures and dynamics of language and text through context hypotheses, uncertainty, contradiction, revision, neuromodulation, replay, and consolidation rather than by merely storing isolated facts.
Python
37
397 commits
updated Sep 28, 2026
BrainStem is a biologically inspired, Real Neuro-Symbolic (RNS-AI) cognitive architecture for lifelong learning. It is designed to learn models of the structures and dynamics of language and text through context hypotheses, uncertainty, contradiction, revision, neuromodulation, replay, and consolidation rather than by merely storing isolated facts. A second, character-level observation layer discovers word boundaries directly from unsegmented text, without any predefined notion of "word." Since 25–28 September 2026, a third, fourth, and fifth observation layer additionally discover relations between already-observed elements, categories emerging from the resulting relation graph, and durable questions emerging from persistent, unresolved information gaps.
One CPU Core / No GPU
[!IMPORTANT] BrainStem is a research and calibration system, not a production-ready assistant. As of the current project stage, all previously closed productive write paths (facts, relations, ontology categories, questions, fact/relation/ontology/question promotion, gap/contradiction/revision writes) have been deliberately opened as an explicitly framed, ongoing experiment, with a full project backup taken beforehand as a fallback point.
YouTube - AI conversation about BrainStem ProjectThe system has completed a full read-through of its current two-source corpus (German Wikipedia Physics and Computer categories, 167,661 chunks, 100%) and has since run well beyond 11,500 real learning cycles in both active-learning and replay/consolidation-only modes without data loss, corruption, or GUI failure. A dedicated 1,500-cycle Etappe-A drift validation (sensory-deprivation mode, input disabled, inner dynamics only) completed with an overall result of "konvergiert": all 20 evaluated signals were classified either stabil or konvergiert, with zero signals flagged as divergent.
Cortisol Stage 2 is active and functionally verified. stage=2 is set in the production database. Under real production conditions it has not been triggered live (allostatic_load has stayed well below the 0.6 threshold — a sign of a consistently calm system, not a defect). Its guarded intervention logic (per-value cap 0.01, per-cycle budget 0.03, 3-cycle cooldown) has been independently confirmed correct under simulated stress.
Hypothesis graduation (uncertain_hypothesis → stable_hypothesis, and — since the Relations/Ontology work below — uncertain_lexical_boundary → stable_lexical_boundary and uncertain_relation_hypothesis → stable_relation_hypothesis) is active, gated by the consolidation-survival criterion (≥3 confirmed Phase-7d survival cycles), an available critic-gate cross-check, warm-up dampening, a budget of one graduation per cycle, and — new — neuromodulator-coupled divisive-normalization role selection among the three competing roles (see below).
As of 25–28 September 2026, the complete Relations/Ontology/Questions emergence chain (originally planned as three slices) is fully implemented, wired into the runtime phase chain, and verified through repeated real, multi-hundred-cycle end-to-end runs. The runtime phase registry now holds 39 entries (previously 34). Additionally, eleven of the pipeline's own calibrated thresholds and budgets are now directly coupled to one or more of the six core digital neuromodulators, each grounded in a specific piece of neuroscience literature (see Neuromodulator-Coupled Thresholds).
Direct Facts, Relations, Ontology, and Questions writes, Fact/Relation/Ontology/Question promotion, Attention writes, and productive Phase-5f/5g/5i experiments are all currently active, alongside the full nine-module Stage-B chain described below.
A Phase 0 module (v8_phase0_lexical_boundary_observation_release) is registered at the front of the runtime phase chain, running alongside the existing sentence-level context_observation_learning entry point. It reads the raw, unsegmented character stream of already-imported chunks, maintains a pure frequency table of "which character follows this preceding context of length k," and computes the local branching entropy at each character position:
H(context) = − Σ P(next_char | context) · log2 P(next_char | context)
A pronounced spike in this entropy at a given position is treated as a boundary candidate and creates or re-observes a context_hypotheses row using role='uncertain_lexical_boundary' — reusing exactly the same insert-or-reobserve mechanism, consolidation path (Phase 7d), Stage-B graduation, and fact promotion already used for every other hypothesis. Once a lexical-boundary candidate's own evidence_count clears a small threshold, it additionally becomes eligible as an input element for the Relations emergence layer below — deliberately before full Stage-B graduation, a design decision grounded in evidence that word segmentation and relational learning proceed in parallel in human learners rather than sequentially (see Academic References).
Three additional observation/promotion layers, completed this session, extend the system beyond isolated sentence- and word-level hypotheses toward structured, relational, categorical, and curiosity-driven knowledge:
Relations (Modul A). v8_phase0b_relational_binding_observation_release binds pairs of already-stable elements that co-occur within the same sentence into a candidate relation, via a strict two-stage statistical process: a symmetric Pointwise Mutual Information existence gate (PMI(A;B) = log2(P(A,B)/(P(A)·P(B)))), followed — only once enough evidence exists — by a purely positional direction signal, calibrated separately per corpus language (German requires a stronger positional skew than English, reflecting German's weaker "position implies grammatical role" signal). v8_stageb_relation_promotion_release promotes graduated relation hypotheses into the relations table, fully retractable on later revision, exactly like fact promotion.
Ontology (Modul B). v8_stageb_ontology_cluster_observation_release builds an undirected graph over the relations table and clusters it via Label Propagation, anchoring cluster identity across cycles to each cluster's own centrality-based prototype (Rosch prototype theory) rather than to Label Propagation's own non-deterministic per-cycle label. v8_stageb_ontology_promotion_release promotes reconfirmed-stable clusters into the ontology table as child → parent rows, with the fixed, honest placeholder relation "is_a" — the system has structurally identified a graded group-membership relationship via connectivity and centrality, without parsing or inferring any actual linguistic category name.
Questions (Modul C). v8_stageb_question_promotion_release promotes a persistent, repeatedly reconfirmed entry from the existing internal_learning_gaps table into the questions table once it clears a calibrated persistence threshold — directly implementing Loewenstein's Information-Gap Theory of curiosity. Question text is never a generated sentence, only a fixed, non-linguistic marker followed by the underlying hypothesis's own verbatim, already-observed surface text. v8_stageb_question_chunk_feedback_release closes the loop back into reading behavior: for every still-open question, it boosts both the exact source chunk where the underlying gap was first observed and searches the full corpus via SQLite FTS5 for other, not-yet-read chunks that might resolve it — the first and only module in the codebase that reads questions to influence real system behavior.
Habituation. Gap detection now additionally retires a persistent-but-unproductive gap (no genuine evidence growth for 64 consecutive reconfirmation cycles) into a reversible habituated state, implementing the empirically established inverted-U relationship between resolvability and curiosity.
Nine modules now form a single, ordered, same-cycle-visible chain per real learning cycle:
facts, fully retractable on later revision.The chain is ordered so that a contradiction resolved this cycle already causes its corresponding fact/relation to be retracted within the same cycle, and a relation/cluster/question promoted this cycle is already visible to the next module in the same chain.
Beyond the pre-existing six-core neuromodulator engine, eleven of the pipeline's own calibrated thresholds and budgets are now directly modulated by one or more of the six core messengers, each using the same symmetric, self-regulating gain (exactly the unmodulated calibrated value at the messenger's own neutral value of 0.5):
| Coupling | Messenger(s) | Grounding |
|---|---|---|
| Gap detection existence gate | Noradrenaline | Aston-Jones & Cohen (2005) adaptive gain theory |
| Gap habituation threshold | Serotonin | Grossman, Bari & Cohen (2022); Hochner et al. (1986) |
| Relational binding existence gate | Noradrenaline | Aston-Jones & Cohen (2005) |
| Relational binding direction threshold | Acetylcholine | Project-internal structural-revision role |
| Hypothesis revision budget | Acetylcholine | Project-internal structural-revision role |
| Contradiction resolution ratio | Noradrenaline | Aston-Jones & Cohen (2005) |
| Graduation role-selection pressure | Acetylcholine | Douchamps et al. (2013); Gómez-Ocádiz et al. (2022) |
| Graduation divisive-normalization temperature | GABA | Katzner, Busse & Carandini (2011) |
| Ontology stability streak | Serotonin | Grossman, Bari & Cohen (2022) |
| Ontology minimum cluster size | Noradrenaline (self-regulating direction) | Aston-Jones & Cohen (2005) vs. Shine et al. (2018) |
| Ontology overlap threshold | Dopamine | Kahnt & Tobler (2016); Novicky et al. (2023) |
| Ontology promotion budget | Acetylcholine | Project-internal structural-revision role |
| Question existence gate | Noradrenaline | Aston-Jones & Cohen (2005) |
| Question priority | Dopamine, Acetylcholine | Established gap-closure/curiosity roles |
| Question-chunk feedback boost magnitude | (inherits question priority) | Consistency with the upstream signal |
| Phase 7d consolidation gate | Acetylcholine | Gais & Born (2004); Hasselmo & McGaughy (2004) |
| Stage | Status | Gating condition to proceed |
|---|---|---|
| **Stage A — Core stability** | Complete: 100% corpus completion, 11,500+ real cycles, clean 1,500-cycle formal drift run ("konvergiert", zero divergent signals) | Complete |
| **Stage B — Guarded graduation** | Active: Cortisol Stage 2 applied, hypothesis graduation active for all three eligible roles, neuromodulator-coupled role selection | Ongoing observation |
| **Autonomous Lexical Emergence (Phase 0)** | Active, feeding both fact promotion and the Relations layer below once its own evidence threshold clears | Calibration against real corpus ongoing |
| **Relations Emergence (Modul A)** | Active and productive — Gate-and-Direction binding, promotion, retraction all verified over real multi-cycle runs | Ongoing observation at real production scale |
| **Ontology Emergence (Modul B)** | Active and productive — Label Propagation clustering, prototype-anchored stability, promotion/retraction, three neuromodulator couplings (one self-regulating) all verified | Ongoing observation at real production scale |
| **Questions Emergence (Modul C)** | Active and productive — persistence-gated promotion, habituation-aware retraction, full-corpus chunk feedback all verified | Ongoing observation at real production scale |
| **Neuromodulator coupling of pipeline thresholds** | Eleven couplings active across seven modules, each individually literature-grounded and baseline-invariance-verified | Ongoing; re-calibration against real production corpus pending |
| **Opening the productive write locks** (Facts / Relations / Ontology / Questions / all four Promotion paths) | Open (experimental project position) | Ongoing observation of the full chain; every promoted artifact remains traceable to, and retractable from, its source hypothesis/gap |
| **Full corpus scaling** (complete German Wikipedia) | Deliberately deferred | Order-of-magnitude estimate only, not a commitment |
| **Vector database evaluation** | Deliberately deferred per project's own architecture-checkpoint rule | Only once stable hypothesis identities exist, a concrete semantic-retrieval use case is identified, requirements are measurable, and a read-only/shadow comparison against the SQLite baseline is performed |
| **Symbolic reasoning plugin** (deterministic, non-LLM rule engine) | Concept documented, not scheduled | Gated behind Stage-B graduation and the write-lock roadmap; a validated symbolic rule becoming productive is itself a new class of productive write and must be gated at least as strictly as fact promotion |
| **Multi-core / multi-process learning** | Explicitly out of scope for now | Not currently planned |
Traditional semantic systems often focus on the what: storing and retrieving content. BrainStem focuses on the how: learning how context, uncertainty, evidence, contradiction, revision, and consolidation interact over time — extended to a character-level layer that learns how recurring units emerge from raw text, a relational layer that learns how those units connect to each other, a categorical layer that learns how connected units group into categories, and a curiosity layer that learns which of its own unresolved gaps are durable enough to motivate further reading. A corpus is treated as training substrate rather than as a static knowledge base.
Core principles:
BrainStem is an autonomous software architecture designed for continuous, self-improving data processing and knowledge management. At its core, the system operates through an Autonomous Loop that orchestrates a chain of 39 learning phases to ingest, analyze, and refine information without manual intervention.
The biological terminology used throughout the project's technical documentation, including terms such as "neuromodulators," "sleep," or "homeostasis," is not decorative. These labels are functional designators for mathematical state variables and algorithmic control mechanisms. The values are floats, not molecules. The behavior is biologically inspired, but the implementation is strictly mathematical.
The system's primary mechanics include:
Dynamic Steering Variables: Digital messenger substances are dynamic meta-parameters that adjust the system's learning rate, error weighting, and exploration strategies in real time, and — as of this session — directly gate eleven specific thresholds and budgets throughout the emergence pipeline.
Active versus Offline Processing: The system cycles between active ingestion and an offline optimization phase ("sleep") in which recorded hypotheses are re-evaluated through batch replay and consolidation.
Character-Level Emergence: Alongside sentence-level hypotheses, a branching-entropy-driven process observes the raw character stream to propose, and — through the same consolidation/graduation machinery — stabilize, candidate word-boundary units.
Relational, Categorical, and Curiosity-Driven Emergence: Beyond individual hypotheses, the system now binds pairs of stable elements into relations via a statistically principled existence-and-direction gate, clusters the resulting relation graph into emergent categories anchored by centrality rather than arbitrary labels, and promotes its own persistent, unresolved information gaps into durable questions that measurably redirect its own reading attention.
Knowledge Distillation: By comparing new data against existing stable records, the system filters out inconsistencies and promotes reliable information into its long-term fact/relation/ontology store — while retaining the ability to retract any of them if the underlying hypothesis is revised.
Equilibrium Control: Stability monitoring routines act as a feedback mechanism that pulls meta-parameters back into a functional range when the system detects a performance plateau or excessive variance.
Adaptive Boundaries: The limits within which the system operates are not hardcoded but self-regulating, expanding or contracting processing thresholds based on the complexity of the data encountered — and, for at least one specific coupling, self-regulating even in which direction a neuromodulator should push a threshold, learned from real outcomes rather than assumed.
In summary, the project is a recursive learning engine that uses biologically derived control logic to implement a highly flexible, self-governing system for automated knowledge acquisition, now spanning word-, relation-, category-, and question-level structure discovery.
BrainStem does not operate as a continuously coupled system of differential equations. Instead, it traverses a cyclic state graph: each phase activates at most 2–3 dominant neuromodulators, while the remainder are kept inactive or passive. This sequential architecture prevents interaction cascades and enables deterministic debugging.
| Stage | Name | Description |
|---|---|---|
| 1 | Inference-free pre-parsing | A raw corpus such as a Wikipedia ZIM file is extracted, structured, and partitioned into the chunk store before autonomous learning begins. |
| 2 | Autonomous learning | AutonomousLoop processes prepared chunks while the neuromodulatory, consolidation, and full write-path chain (facts, relations, ontology, questions) reacts to the evolving internal state. |
BrainStem currently uses 12 digital neuromodulators. Their values are normalized to [0.0, 1.0] and derived from internal system state under bounded, biologically inspired dynamics and homeostatic constraints.
| Neuromodulator | Current engineering role |
|---|---|
| Dopamine | outcome and gap-closure signal; now also couples question priority and the ontology layer's overlap tolerance |
| Serotonin | consolidation and stability signal; now also couples gap habituation resistance and ontology cluster-stability requirements |
| Glutamate | excitatory drive associated with exploration and learning activity |
| GABA | global inhibition and E/I-balance signal; now also couples the sharpness of divisive-normalization role competition at graduation |
| Noradrenaline | error, alarm, and persistent-pressure signal; now also couples five separate existence/resolution gates across the pipeline, one of them (ontology minimum cluster size) with a self-regulating, outcome-learned direction |
| Acetylcholine | novelty, attention, and structural-revision signal; now also couples relation direction commitment, hypothesis revision budget, graduation role-novelty pressure, ontology promotion budget, and Phase 7d consolidation gating |
| Adenosine | sleep-pressure homeostat |
| Endocannabinoids | retrograde gain control |
| Cortisol | top-level stability watcher and guarded soft regulator (Stage 2 active) |
| Histamine | wake and arousal signal |
| Orexin | reading-endurance and curiosity-related drive |
| BDNF | activity-dependent growth and consolidation substrate |
Phase 6a performs offline-style replay after the wake path. Phase 6b evaluates replay effectiveness and plasticity adjustments. The critic gate checks whether proposed changes remain consistent enough to be retained; rejected or unstable material remains available as error and revision evidence. The same critic gate is reused, unmodified, by Stage-B hypothesis graduation.
Phase 7d adds sub-1-Hz up/down-state processing with stochastic reactivation, adaptive thresholds, activity-dependent participation, survivor and weakening statistics, anchor interleaving, and self-regulating down-selection. Sentence-level, lexical-boundary, and relation hypotheses all share this same candidate pool, distinguished only by their role value. Consolidation activity is now additionally gated by acetylcholine, reflecting the causal, pharmacologically demonstrated role of low cholinergic tone during slow-wave sleep in permitting declarative memory consolidation.
As of the current experimental project position, the following paths are open:
BrainStem uses ki_memory.sqlite3 in the project root. The database is created automatically when absent. Schema rules: schema changes must be reflected in the bootstrap in the same delivery; ensure_schema must be idempotent; _self_check_schema must run before writes; every written column must already be declared in SCHEMA_TABLES; compile checks, smoke tests, and intermediate checks are required before delivery; structural changes require a full backup first.
Learning state can be reset without re-importing the corpus. Preserved content includes documents, chunks, FTS data, import state, and configuration. The reset workflow performs a dry run and creates a timestamped database backup before applying changes.
From the project root: python main.py --gui
| Step | GUI action | Purpose |
|---|---|---|
| 1 | Export / Configuration | Configure the maximum number of articles before import. |
| 2 | Import & Jobs → ZIM Einlesen | Extract and pre-parse the corpus. |
| 3 | Import & Jobs → Autonom dauerhaft starten | Start autonomous learning. |
| 4 | Import & Jobs → Autonom stoppen | Stop autonomous learning cooperatively. |
| 5 | Close the GUI normally | Wait for active workers instead of terminating them abruptly. |
This project builds upon concepts, algorithms, and theoretical frameworks established in the following academic literature:
Word segmentation and statistical language acquisition
Curiosity, information gaps, and question formation
Habituation, dishabituation, and information-gain-driven decay
Noradrenaline, arousal, and network topology
Dopamine, generalization, and precision
Acetylcholine, novelty, and encoding-versus-retrieval balance
Serotonin and meta-learning under uncertainty
GABA, gain, and stimulus selectivity
Divisive normalization
Category and prototype formation
Complementary learning systems and structure extraction
Sleep-dependent consolidation, homeostatic plasticity, and previously referenced principles (unchanged)
@article{aston-jones2005integrative,
title={An Integrative Theory of Locus Coeruleus-Norepinephrine Function: Adaptive Gain and Optimal Performance},
author={Aston-Jones, Gary and Cohen, Jonathan D.},
journal={Annual Review of Neuroscience},
year={2005}
}
@article{kahnt2016dopamine,
title={Dopamine regulates stimulus generalization in the human hippocampus},
author={Kahnt, Thorsten and Tobler, Philippe N.},
journal={eLife},
year={2016}
}
@article{grossman2022serotonin,
title={Serotonin neurons modulate learning rate through uncertainty},
author={Grossman, Cooper D. and Bari, Bilal A. and Cohen, Jeremiah Y.},
journal={Current Biology},
year={2022}
}
@article{katzner2011gaba,
title={GABA\_A Inhibition Controls Response Gain in Visual Cortex},
author={Katzner, Steffen and Busse, Laura and Carandini, Matteo},
journal={Journal of Neuroscience},
year={2011}
}
@article{carandini2012normalization,
title={Normalization as a canonical neural computation},
author={Carandini, Matteo and Heeger, David J.},
journal={Nature Reviews Neuroscience},
year={2012}
}
@article{loewenstein1994psychology,
title={The Psychology of Curiosity: A Review and Reinterpretation},
author={Loewenstein, George},
journal={Psychological Bulletin},
year={1994}
}
@article{shine2018modulation,
title={The modulation of neural gain facilitates a transition between functional segregation and integration in the brain},
author={Shine, James M. and Aburn, Matthew J. and Breakspear, Michael and Poldrack, Russell A.},
journal={eLife},
year={2018}
}
@article{hamilton2020graph,
title={Graph Representation Learning},
author={Hamilton, William L.},
year={2020},
publisher={Morgan \& Claypool Publishers}
}
@article{watkins2020using,
title={Using Sinusoidally-Modulated Noise as a Surrogate for Slow-Wave Sleep to Accomplish Stable Unsupervised Dictionary Learning in a Spike-Based Sparse Coding Model},
author={Watkins, Yijing and Kim, Edward and Kenyon, Garrett T.},
journal={Frontiers in Computational Neuroscience},
year={2020}
}
@article{tadros2022biologically,
title={Biologically Inspired Sleep Algorithm for Reducing Catastrophic Forgetting in Neural Networks},
author={Tadros, Timothy and Tran, Gia-Bao M. and Krishnan, Giri P. and Bazhenov, Maxim},
year={2022}
}
@article{fischbacher2020intelligent,
title={Intelligent Matrix Exponentiation},
author={Fischbacher, Thomas and Comsa, Iulia M. and Potempa, Krzysztof and Firsching, Moritz and Versari, Luca and Alakuijala, Jyrki},
journal={arXiv preprint arXiv:2008.03926},
year={2020}
}
@article{butz2013homeostatic,
title={Homeostatic structural plasticity--a key to neuronal network formation and repair},
author={Butz, Markus and van Ooyen, Arjen},
journal={PLoS Computational Biology},
year={2013}
}
@article{lee2019mechanisms,
title={Mechanisms of Homeostatic Synaptic Plasticity In Vivo},
author={Lee, Kea-Joo K. and Kirkwood, Alfredo},
journal={Frontiers in Cellular Neuroscience},
year={2019}
}
@article{parker2020nonlinear,
title={Nonlinear Time Series Classification Using Bispectrum-based Deep Convolutional Neural Networks},
author={Parker, Paul A. and Holan, Scott H. and Ravishanker, Nalini},
journal={arXiv preprint arXiv:2003.02353},
year={2020}
}
@article{roncevic2023molecule,
title={Supplementary Materials for A molecule with half-M{\"o}bius topology},
author={Ron{\v{c}}evi{\'c}, Igor and others},
journal={Nature Chemistry},
year={2023}
}
BrainStem is an experimental cognitive-architecture research project. Biological terminology is used as an engineering analogy and design inspiration. The software is not a biological simulation and does not claim neuroscientific equivalence.
Python
100.0%