tinystories-15m — a language model pretrained entirely by volunteers
3
200 commits
1 linked in READMEs
updated Aug 3, 2026
Training complete (August 2, 2026). Final state: outer step 195, val loss 2.872, ~423M community tokens — pinned at revision
5c39e2cb. This model stays fully usable (snippet below). The co-op's current run is fineweb-150m — a 145M-param model on FineWeb-Edu;npx coop-ai startnow contributes there.
This 14.8M-parameter model was pretrained from scratch with no cluster, no server, and no funding: volunteers ran training rounds on their own computers and submitted compressed pseudo-gradients as pull requests on a public Hugging Face dataset repo. A stateless GitHub Actions cron job aggregated them into outer steps — DiLoCo-style low-communication data parallelism, coordinated by nothing but free-tier infrastructure.
ledger branchPrompt: Once upon a time
Once upon a time, there was a little girl named Lily. She loved to play outside in her garden. One day, she found a big, yellow flower. She was so happy and started to look at the pretty petals. She wanted to decorate her garden with the flowers and show them to her mom.
But when Lily went to visit her grandma, she saw her small flower. She looked guilty and felt sad. Lily wanted to show her flower to her grandma, but her flower was gone. She started to cry because she missed her flower.
| Architecture | decoder-only transformer (nanoGPT-style, pre-LN, tied embeddings) |
| Parameters | 14,769,216 |
| Layers / heads / width | 6 / 6 / 396 |
| Context length | 512 tokens |
| Tokenizer | custom 8,192-token byte-level BPE trained on TinyStories (tokenizer.json, in this repo) |
| Data | roneneldan/TinyStories, volunteers train on per-user shards |
| Format | checkpoint.safetensors (weights) · optimizer.safetensors (outer momentum) · meta.json (step, config, eval) |
Workers download the current checkpoint, run up to 500 local AdamW steps
(lr 3e-4, betas 0.9/0.95, weight decay 0.1, grad clip 1.0) on their personal
TinyStories shard, and submit the pseudo-gradient θ_outer − θ_local,
int8-quantized, as a pull request. Each aggregation tick is stateless: it drops
over-stale submissions (> 8 steps old), L2-clips each delta, cosine-gates against
the clipped weighted mean, merges same-user submissions into one token-weighted
vote, robust-aggregates the votes (20% trimmed mean), and takes one Nesterov outer
step (lr 0.7, momentum 0.9). Contributions are Byzantine-filtered but the compute
itself is unverified — this is a volunteer-trust experiment as much as a model.
The model uses its own tiny GPT class (not transformers):
# pip install git+https://github.com/commonsense-ai/coop
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from tokenizers import Tokenizer
from coop.model import GPT, GPTConfig, load_canonical_state
repo = "commonsense-ai/tinystories-15m"
model = GPT.from_config(GPTConfig()) # defaults match this checkpoint
load_canonical_state(model, load_file(hf_hub_download(repo, "checkpoint.safetensors")))
model.eval()
tok = Tokenizer.from_file(hf_hub_download(repo, "tokenizer.json"))
idx = torch.tensor([tok.encode("Once upon a time").ids])
out = model.generate(idx, max_new_tokens=200, temperature=0.8, top_k=50)
print(tok.decode(out[0].tolist()))
It writes toddler fiction, and only toddler fiction. TinyStories is a synthetic corpus with a ~1,500-word vocabulary world, so the model has no knowledge, no instruction-following, and no register other than bedtime stories about children, animals, and feelings. Plots meander and pronouns drift. 512-token context. It is an educational artifact demonstrating that strangers on the internet can pretrain a real model together — use it to study that, not to ship products.
The run may continue and successors are planned. Joining takes one command:
npx coop-ai start
Code, architecture writeup, and the leaderboard: github.com/commonsense-ai/coop
tinystories-15m — a language model pretrained entirely by volunteers
3
200 commits
1 linked in READMEs
updated Aug 3, 2026
Training complete (August 2, 2026). Final state: outer step 195, val loss 2.872, ~423M community tokens — pinned at revision
5c39e2cb. This model stays fully usable (snippet below). The co-op's current run is fineweb-150m — a 145M-param model on FineWeb-Edu;npx coop-ai startnow contributes there.
This 14.8M-parameter model was pretrained from scratch with no cluster, no server, and no funding: volunteers ran training rounds on their own computers and submitted compressed pseudo-gradients as pull requests on a public Hugging Face dataset repo. A stateless GitHub Actions cron job aggregated them into outer steps — DiLoCo-style low-communication data parallelism, coordinated by nothing but free-tier infrastructure.
ledger branchPrompt: Once upon a time
Once upon a time, there was a little girl named Lily. She loved to play outside in her garden. One day, she found a big, yellow flower. She was so happy and started to look at the pretty petals. She wanted to decorate her garden with the flowers and show them to her mom.
But when Lily went to visit her grandma, she saw her small flower. She looked guilty and felt sad. Lily wanted to show her flower to her grandma, but her flower was gone. She started to cry because she missed her flower.
| Architecture | decoder-only transformer (nanoGPT-style, pre-LN, tied embeddings) |
| Parameters | 14,769,216 |
| Layers / heads / width | 6 / 6 / 396 |
| Context length | 512 tokens |
| Tokenizer | custom 8,192-token byte-level BPE trained on TinyStories (tokenizer.json, in this repo) |
| Data | roneneldan/TinyStories, volunteers train on per-user shards |
| Format | checkpoint.safetensors (weights) · optimizer.safetensors (outer momentum) · meta.json (step, config, eval) |
Workers download the current checkpoint, run up to 500 local AdamW steps
(lr 3e-4, betas 0.9/0.95, weight decay 0.1, grad clip 1.0) on their personal
TinyStories shard, and submit the pseudo-gradient θ_outer − θ_local,
int8-quantized, as a pull request. Each aggregation tick is stateless: it drops
over-stale submissions (> 8 steps old), L2-clips each delta, cosine-gates against
the clipped weighted mean, merges same-user submissions into one token-weighted
vote, robust-aggregates the votes (20% trimmed mean), and takes one Nesterov outer
step (lr 0.7, momentum 0.9). Contributions are Byzantine-filtered but the compute
itself is unverified — this is a volunteer-trust experiment as much as a model.
The model uses its own tiny GPT class (not transformers):
# pip install git+https://github.com/commonsense-ai/coop
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from tokenizers import Tokenizer
from coop.model import GPT, GPTConfig, load_canonical_state
repo = "commonsense-ai/tinystories-15m"
model = GPT.from_config(GPTConfig()) # defaults match this checkpoint
load_canonical_state(model, load_file(hf_hub_download(repo, "checkpoint.safetensors")))
model.eval()
tok = Tokenizer.from_file(hf_hub_download(repo, "tokenizer.json"))
idx = torch.tensor([tok.encode("Once upon a time").ids])
out = model.generate(idx, max_new_tokens=200, temperature=0.8, top_k=50)
print(tok.decode(out[0].tolist()))
It writes toddler fiction, and only toddler fiction. TinyStories is a synthetic corpus with a ~1,500-word vocabulary world, so the model has no knowledge, no instruction-following, and no register other than bedtime stories about children, animals, and feelings. Plots meander and pronouns drift. 512-token context. It is an educational artifact demonstrating that strangers on the internet can pretrain a real model together — use it to study that, not to ship products.
The run may continue and successors are planned. Joining takes one command:
npx coop-ai start
Code, architecture writeup, and the leaderboard: github.com/commonsense-ai/coop