keskival/sleeping_machines

A manifesto for temporal computing where memory indexing is replaced by temporal delays and referencing.

Python

5

496 commits

updated Sep 29, 2026

See the code

See what people are saying

README

Sleeping Machines

Learning to compute with time, vector messages and local memory.

Sleeping Machines explores models in which an event carries a learned vector and an arrival time. Nodes accumulate local evidence, transform messages and compete through learned delays. The winning message determines what happens next; unrealized alternatives can teach the network to make better choices.

Delays do computation. Changing a delay changes arrival order, which memories interact and which route wins. Sleeping units and sparse activity are part of the resource model; temporal computation is the central architectural idea.

The goal is a broadly useful model family with competitive predictive quality and substantially less physical work in both training and inference. This repository contains the implementations, mathematical analysis, reproducible experiments and completed results used to pursue that goal.

Start with the project report or PDF. The original ideas are preserved in the historical motivation manifesto.

Demonstrated capabilities

CapabilityCompleted evidenceScope
Language prediction1.613 test bits/character with 10M training characters, compared with LSTM 1.799 and Transformer 1.908; lower is betterNative predictive mixture, text8 split; different model sizes and fitting schedules
Rule generalization100% on all 3,440 unseen mod-17 triples, with 69 learned phase scalarsPeriodic primitive in the common model; supplied period 17; certified across all 4,913 possible triples
Longer-context retrieval100% at four times the training contextCommon two-layer carrier plus learned relative pointer; controlled synthetic task
Temporal compositionNative shared-motif models reach approximately 99.65%, with approximately 10,000× less counted work than the named Transformer referenceTask-specific native model and declared operation ledger

Accuracy versus computation shows the consolidated arithmetic and retrieval comparisons. The report contains protocols, reference models, training budgets and the definitions behind each work estimate. Counted operations and logical memory visits are not measured energy.

The central next result

The strongest language result comes from a specialized predictive mixture. It has not yet been reproduced by a generic deep event language model without explicit context-count or pointer experts. The common implementation currently covers multiple independently trained tasks; sharing an implementation does not by itself establish general representation learning.

A bounded expert-free language screen now reaches 3.395 validation bits/character with eight layers, versus 3.464 with one layer. It uses 8,192 training characters, four passes and 1,024 validation predictions; all layers' value, route and memory-time parameters update. Shuffling preceding characters while preserving the last character, count and timestamps increases the deeper model's loss to 3.805. This is evidence of trainability and context sensitivity, with higher computation cost for the deeper model. It is a small development result, not the 10M-character mixture result or a scaling claim. Results and work audit.

An eight-layer event/race model reaches 72.3% on 512 held-out SHD utterances from reserved training-file speakers. This demonstrates deep learning and some speaker transfer. It does not establish competitive speech recognition; the published official-test protocols are different, and our official SHD test set remains untouched.

The next scaling milestone is a learned language model whose event backbone owns the prediction: increasing data budgets, matched Transformer/RNN/state-space references, and explicit measurement of quality, parameters, training work, memory traffic, time and energy. A large reduction in joules at useful predictive quality would be a significant result even before an advantage in raw loss. See the language scaling protocol.

Architecture and learning

The shared model combines configurable mechanisms:

  • Sparse event carriers: vector messages pass through a configurable depth without evaluating a hidden grid of empty time steps.
  • Local temporal memory: causal numerator/mass states retain and combine observations with linear scan work.
  • Hard races and learned delays: a winning route emits its own payload and time. Losers remain distinct training alternatives.
  • Separate keys and values: an optional key stream computes input-dependent routing and clocks while value learning preserves that schedule.
  • Structured memories: optional conditional evidence, relative pointers and periodic transformations supply useful computations when appropriate.
  • Task-appropriate supervision: categorical decisions or next-event type and waiting-time likelihoods, including the information in silence.

Tasks have separate fitted weights and may use different depth or primitives. Small vector maps and output readouts use dense arithmetic. The current CPU reference also sorts arrivals and replays contexts; persistent incremental execution and cheaper candidate discovery are important systems objectives. All this work belongs in the resource accounting.

Read the theory and evidence

DocumentPurpose
Report and applicationsAccessible overview, strongest results, potential and benchmark appendices
Shared modelCurrent implementation, task adapters, contracts and cost boundaries
Theory indexFormal derivations organized by theme, with assumptions and proof scope
Mathematical programOpen analytic problems and their decisive measurements
Research roadmapNext experiments and architectural priorities
Language scaling protocolGeneric prediction, baseline matching and physical work measurements
FindingsCompleted experiment history and detailed observations
Historical manifestoOriginal motivation, exploratory ideas and early references

Source lives in sleeping_machines/; experiment drivers, executed queue commands and result records live in experiments/. The implementation is a research reference, with explicit contracts and source hashes rather than a claim of a production event processor.

Run experiments safely

Read AGENTS.md before launching work. Every experiment runs through run_safe.sh, which holds a host-local lock, limits threads and monitors memory. Run one job at a time per host, preserve at least 8 GiB of available host memory, and give changed settings a new queue name and result tag. Inspect existing jobs, completed results and GPU occupancy first.

The current runner expects this checkout at /workspace and its interpreter at /workspace/.venv-docker/bin/python. The CPU environment uses requirements.txt, PyTorch and h5py; task datasets are supplied separately. The installation runbook and bootstrap script document a provisioned experiment host. Choose resource limits from the actual host capacity.

For a small shared-model contract check in an already provisioned environment:

RUN_TAG="shared_contracts_$(date -u +%Y%m%dT%H%M%SZ)"
QUEUE="experiments/queue/${RUN_TAG}.txt"
printf '%s experiments/e120_shared_contracts.py --tag %s\n' "$RUN_TAG" "$RUN_TAG" > "$QUEUE"
MEM_CAP_KB=3600000 MEM_CAP_RSS_KB=2600000 MIN_AVAIL_MB=8192 JOB_TIMEOUT_S=600 \
  bash experiments/queue/run_safe.sh "$QUEUE"

These caps suit the bounded CPU check on a host with enough free memory. Training runs need their own measured memory budget and timeout. Completed experiment commands are preserved under experiments/queue/; use a new tag when reproducing them. Logs and result records carry the executed settings and measured metrics.

Citing

Sleeping Machines — Tero Keski-Valkama and Karoliina Salminen.

@article{keskival2021sleeping,
  title={Sleeping Machines},
  author={Keski-Valkama, Tero and Salminen, Karoliina},
  year={2021},
  doi={10.5281/zenodo.13207423}
}

DOI

asynchronous-programming
deep-learning
deep-neural-networks
hacktoberfest
hacktoberfest2021
neuromorphic
neuromorphic-computing

keskival/sleeping_machines

A manifesto for temporal computing where memory indexing is replaced by temporal delays and referencing.

Python

5

496 commits

updated Sep 29, 2026

See the code

See what people are saying

README

Sleeping Machines

Learning to compute with time, vector messages and local memory.

Sleeping Machines explores models in which an event carries a learned vector and an arrival time. Nodes accumulate local evidence, transform messages and compete through learned delays. The winning message determines what happens next; unrealized alternatives can teach the network to make better choices.

Delays do computation. Changing a delay changes arrival order, which memories interact and which route wins. Sleeping units and sparse activity are part of the resource model; temporal computation is the central architectural idea.

The goal is a broadly useful model family with competitive predictive quality and substantially less physical work in both training and inference. This repository contains the implementations, mathematical analysis, reproducible experiments and completed results used to pursue that goal.

Start with the project report or PDF. The original ideas are preserved in the historical motivation manifesto.

Demonstrated capabilities

CapabilityCompleted evidenceScope
Language prediction1.613 test bits/character with 10M training characters, compared with LSTM 1.799 and Transformer 1.908; lower is betterNative predictive mixture, text8 split; different model sizes and fitting schedules
Rule generalization100% on all 3,440 unseen mod-17 triples, with 69 learned phase scalarsPeriodic primitive in the common model; supplied period 17; certified across all 4,913 possible triples
Longer-context retrieval100% at four times the training contextCommon two-layer carrier plus learned relative pointer; controlled synthetic task
Temporal compositionNative shared-motif models reach approximately 99.65%, with approximately 10,000× less counted work than the named Transformer referenceTask-specific native model and declared operation ledger

Accuracy versus computation shows the consolidated arithmetic and retrieval comparisons. The report contains protocols, reference models, training budgets and the definitions behind each work estimate. Counted operations and logical memory visits are not measured energy.

The central next result

The strongest language result comes from a specialized predictive mixture. It has not yet been reproduced by a generic deep event language model without explicit context-count or pointer experts. The common implementation currently covers multiple independently trained tasks; sharing an implementation does not by itself establish general representation learning.

A bounded expert-free language screen now reaches 3.395 validation bits/character with eight layers, versus 3.464 with one layer. It uses 8,192 training characters, four passes and 1,024 validation predictions; all layers' value, route and memory-time parameters update. Shuffling preceding characters while preserving the last character, count and timestamps increases the deeper model's loss to 3.805. This is evidence of trainability and context sensitivity, with higher computation cost for the deeper model. It is a small development result, not the 10M-character mixture result or a scaling claim. Results and work audit.

An eight-layer event/race model reaches 72.3% on 512 held-out SHD utterances from reserved training-file speakers. This demonstrates deep learning and some speaker transfer. It does not establish competitive speech recognition; the published official-test protocols are different, and our official SHD test set remains untouched.

The next scaling milestone is a learned language model whose event backbone owns the prediction: increasing data budgets, matched Transformer/RNN/state-space references, and explicit measurement of quality, parameters, training work, memory traffic, time and energy. A large reduction in joules at useful predictive quality would be a significant result even before an advantage in raw loss. See the language scaling protocol.

Architecture and learning

The shared model combines configurable mechanisms:

  • Sparse event carriers: vector messages pass through a configurable depth without evaluating a hidden grid of empty time steps.
  • Local temporal memory: causal numerator/mass states retain and combine observations with linear scan work.
  • Hard races and learned delays: a winning route emits its own payload and time. Losers remain distinct training alternatives.
  • Separate keys and values: an optional key stream computes input-dependent routing and clocks while value learning preserves that schedule.
  • Structured memories: optional conditional evidence, relative pointers and periodic transformations supply useful computations when appropriate.
  • Task-appropriate supervision: categorical decisions or next-event type and waiting-time likelihoods, including the information in silence.

Tasks have separate fitted weights and may use different depth or primitives. Small vector maps and output readouts use dense arithmetic. The current CPU reference also sorts arrivals and replays contexts; persistent incremental execution and cheaper candidate discovery are important systems objectives. All this work belongs in the resource accounting.

Read the theory and evidence

DocumentPurpose
Report and applicationsAccessible overview, strongest results, potential and benchmark appendices
Shared modelCurrent implementation, task adapters, contracts and cost boundaries
Theory indexFormal derivations organized by theme, with assumptions and proof scope
Mathematical programOpen analytic problems and their decisive measurements
Research roadmapNext experiments and architectural priorities
Language scaling protocolGeneric prediction, baseline matching and physical work measurements
FindingsCompleted experiment history and detailed observations
Historical manifestoOriginal motivation, exploratory ideas and early references

Source lives in sleeping_machines/; experiment drivers, executed queue commands and result records live in experiments/. The implementation is a research reference, with explicit contracts and source hashes rather than a claim of a production event processor.

Run experiments safely

Read AGENTS.md before launching work. Every experiment runs through run_safe.sh, which holds a host-local lock, limits threads and monitors memory. Run one job at a time per host, preserve at least 8 GiB of available host memory, and give changed settings a new queue name and result tag. Inspect existing jobs, completed results and GPU occupancy first.

The current runner expects this checkout at /workspace and its interpreter at /workspace/.venv-docker/bin/python. The CPU environment uses requirements.txt, PyTorch and h5py; task datasets are supplied separately. The installation runbook and bootstrap script document a provisioned experiment host. Choose resource limits from the actual host capacity.

For a small shared-model contract check in an already provisioned environment:

RUN_TAG="shared_contracts_$(date -u +%Y%m%dT%H%M%SZ)"
QUEUE="experiments/queue/${RUN_TAG}.txt"
printf '%s experiments/e120_shared_contracts.py --tag %s\n' "$RUN_TAG" "$RUN_TAG" > "$QUEUE"
MEM_CAP_KB=3600000 MEM_CAP_RSS_KB=2600000 MIN_AVAIL_MB=8192 JOB_TIMEOUT_S=600 \
  bash experiments/queue/run_safe.sh "$QUEUE"

These caps suit the bounded CPU check on a host with enough free memory. Training runs need their own measured memory budget and timeout. Completed experiment commands are preserved under experiments/queue/; use a new tag when reproducing them. Logs and result records carry the executed settings and measured metrics.

Citing

Sleeping Machines — Tero Keski-Valkama and Karoliina Salminen.

@article{keskival2021sleeping,
  title={Sleeping Machines},
  author={Keski-Valkama, Tero and Salminen, Karoliina},
  year={2021},
  doi={10.5281/zenodo.13207423}
}

DOI

asynchronous-programming
deep-learning
deep-neural-networks
hacktoberfest
hacktoberfest2021
neuromorphic
neuromorphic-computing

Languages

Python

88.7%

HTML

8.4%

Shell

2.2%