Measurement-first lab notebook: 232 reports and 548 harness scripts from ~70 local-LLM experiments on one desktop. Mostly negative results, kept deliberately.
Python
0
6 commits
updated Oct 5, 2026
A working lab notebook, published as-is: 232 reports and 548 harness scripts from ~70 experiments run on one desktop machine over a year.
Most of what is here is negative results. That is not an apology — it is the reason to publish. Hypotheses that died under measurement are the expensive part of this work, and almost nobody publishes them.
The hardware is the point, not a limitation. Everything below was measured on a Ryzen 5 1600 (Zen 1), a GTX 1660 SUPER with 6 GiB, 24 GB of RAM and a SATA SSD. No cluster, no A100, no cloud credits. If you are running models on a machine like that, these numbers transfer to you directly.
Every row below names the file that produced it. If a claim here is not in that file, it is a bug — open an issue.
| finding | what was measured | where |
|---|---|---|
| Split the context into blocks — weak models gain the most | 6 local models (1–2B), 40 tasks, 720 calls. With the full context, recall decays with position: 0.609 at position 0, 0.214 at position 17. Blocks of 6 records shrink the head–tail gap from 0.204 to 0.021. qwen2.5-coder-1.5b goes from 0.175 to 0.700, and the weaker the model, the bigger the gain. No emergence: the swarm adds +0.050 over the best single model run through the same pipeline. | REPORT_K1.en.md (English), REPORT_K1.md (Russian original) |
llama-bench feeds random tokens | input is generated with std::rand() % n_vocab — and on Windows RAND_MAX is 32 767 against a vocabulary of 201 088. Anything routing-dependent (MoE expert selection, cache behaviour) cannot be measured with it; use llama-server. This invalidates a whole category of benchmark results, including several of my own. | hardware/REVISION-CLUSTERS-HILBERT-GLASS.md §6, hardware/HANDOFF-2026-09-15.md |
| Batching does not rescue MoE on CPU | ×1.23 total at batch 8 — a batch of 8 tokens cost about 6.5 single tokens, so it does not pay for itself. | experiments/swebench_pro/reports/REPORT_SWE1_PILOT.md |
| Quality follows a power law in depth, with no cliff | PPL ~ d^−2.18, i.e. PPL(d) = PPL_full · (d_full/d)^2.18. An earlier "cliff" I reported turned out to be an artefact of channel ordering, and the conclusion that depended on it was withdrawn. | hardware/AUDIT-GPTOSS-PPL.md |
KV at q4_0 costs 10.7 % perplexity, not 1.5 % | a number I had published earlier was wrong by 7×, and this is the correction. Kept here deliberately. | hardware/PHASE0-RESULTS.md |
| A swarm of weak models did not beat the best single model | E_score = −0.069 ± 0.027 over 30 tasks. The candidate pool of best-of-N was stronger before any selection was applied. | experiments/mycelium/lab/REPORT.md |
Most reports are written in Russian. English versions sit next to them with an .en.md suffix, for
example REPORT_K1.en.md and REPORT_PIPELINE.en.md. A file without a suffix is the Russian original,
and the original wins if the two disagree.
experiments/ ~70 directories, one per experiment
<name>/
README.md what this experiment is
PROTOCOL.md conditions and method, written before the run
BLOCKERS.md what stopped it, where it stopped
reports/ results
*.py the harness that produced them
hardware/ measurements about the machine itself: throughput,
KV cache economics, MoE offload, build costs
Notable groups:
experiments/mycelium/ — an attempt to build a system from a pool of weak
local models that beats the best single model in the pool. It does not work,
and AUDIT_RESTORE.md is the post-mortem of why.experiments/theory_emergent_swarm/ — the theory that followed, tested
against LongBench v2 rather than against itself.experiments/arch0..arch2/ — architecture cycles: dependency model,
complexity budgets, a small combinator language and its runner.experiments/agent_life*/, agent_a*/ — selection, retention and budget
experiments over recorded runs.Raw run artefacts. They are roughly 16 GB of .jsonl and .json across
these experiments, and putting them in git would make the repository unusable
for the 10 MB of text that is actually worth reading. Each experiment keeps its
aggregate metrics and its report; if you want the raw runs for a specific
experiment to re-derive a number, open an issue and say which one.
<PROJECT_ROOT>,
<HOME> and <MODELS>. Nothing else was edited for publication.Reports and text: CC BY 4.0. Harness code: MIT. Use the numbers, cite the directory.
Measurement-first lab notebook: 232 reports and 548 harness scripts from ~70 local-LLM experiments on one desktop. Mostly negative results, kept deliberately.
Python
0
6 commits
updated Oct 5, 2026
A working lab notebook, published as-is: 232 reports and 548 harness scripts from ~70 experiments run on one desktop machine over a year.
Most of what is here is negative results. That is not an apology — it is the reason to publish. Hypotheses that died under measurement are the expensive part of this work, and almost nobody publishes them.
The hardware is the point, not a limitation. Everything below was measured on a Ryzen 5 1600 (Zen 1), a GTX 1660 SUPER with 6 GiB, 24 GB of RAM and a SATA SSD. No cluster, no A100, no cloud credits. If you are running models on a machine like that, these numbers transfer to you directly.
Every row below names the file that produced it. If a claim here is not in that file, it is a bug — open an issue.
| finding | what was measured | where |
|---|---|---|
| Split the context into blocks — weak models gain the most | 6 local models (1–2B), 40 tasks, 720 calls. With the full context, recall decays with position: 0.609 at position 0, 0.214 at position 17. Blocks of 6 records shrink the head–tail gap from 0.204 to 0.021. qwen2.5-coder-1.5b goes from 0.175 to 0.700, and the weaker the model, the bigger the gain. No emergence: the swarm adds +0.050 over the best single model run through the same pipeline. | REPORT_K1.en.md (English), REPORT_K1.md (Russian original) |
llama-bench feeds random tokens | input is generated with std::rand() % n_vocab — and on Windows RAND_MAX is 32 767 against a vocabulary of 201 088. Anything routing-dependent (MoE expert selection, cache behaviour) cannot be measured with it; use llama-server. This invalidates a whole category of benchmark results, including several of my own. | hardware/REVISION-CLUSTERS-HILBERT-GLASS.md §6, hardware/HANDOFF-2026-09-15.md |
| Batching does not rescue MoE on CPU | ×1.23 total at batch 8 — a batch of 8 tokens cost about 6.5 single tokens, so it does not pay for itself. | experiments/swebench_pro/reports/REPORT_SWE1_PILOT.md |
| Quality follows a power law in depth, with no cliff | PPL ~ d^−2.18, i.e. PPL(d) = PPL_full · (d_full/d)^2.18. An earlier "cliff" I reported turned out to be an artefact of channel ordering, and the conclusion that depended on it was withdrawn. | hardware/AUDIT-GPTOSS-PPL.md |
KV at q4_0 costs 10.7 % perplexity, not 1.5 % | a number I had published earlier was wrong by 7×, and this is the correction. Kept here deliberately. | hardware/PHASE0-RESULTS.md |
| A swarm of weak models did not beat the best single model | E_score = −0.069 ± 0.027 over 30 tasks. The candidate pool of best-of-N was stronger before any selection was applied. | experiments/mycelium/lab/REPORT.md |
Most reports are written in Russian. English versions sit next to them with an .en.md suffix, for
example REPORT_K1.en.md and REPORT_PIPELINE.en.md. A file without a suffix is the Russian original,
and the original wins if the two disagree.
experiments/ ~70 directories, one per experiment
<name>/
README.md what this experiment is
PROTOCOL.md conditions and method, written before the run
BLOCKERS.md what stopped it, where it stopped
reports/ results
*.py the harness that produced them
hardware/ measurements about the machine itself: throughput,
KV cache economics, MoE offload, build costs
Notable groups:
experiments/mycelium/ — an attempt to build a system from a pool of weak
local models that beats the best single model in the pool. It does not work,
and AUDIT_RESTORE.md is the post-mortem of why.experiments/theory_emergent_swarm/ — the theory that followed, tested
against LongBench v2 rather than against itself.experiments/arch0..arch2/ — architecture cycles: dependency model,
complexity budgets, a small combinator language and its runner.experiments/agent_life*/, agent_a*/ — selection, retention and budget
experiments over recorded runs.Raw run artefacts. They are roughly 16 GB of .jsonl and .json across
these experiments, and putting them in git would make the repository unusable
for the 10 MB of text that is actually worth reading. Each experiment keeps its
aggregate metrics and its report; if you want the raw runs for a specific
experiment to re-derive a number, open an issue and say which one.
<PROJECT_ROOT>,
<HOME> and <MODELS>. Nothing else was edited for publication.Reports and text: CC BY 4.0. Harness code: MIT. Use the numbers, cite the directory.