A language model on a Game Boy Color - pure SM83 assembly
Python
27
180 commits
updated Sep 29, 2026
A language model that runs on a Game Boy Color written in Assembly.
v1.0 seals the engine and adds a second cartridge: Rei, a small being who talks to you, asks your name and keeps it.
The story model and its tokens didn't change. The wait did: the encoder is exact and 150 times faster, and a prompt token now only feeds the state. A story starts in 1.8 s where it took 6.8.
some stats:
Trained on TinyStories, compared per character against every version before it

One frame per character; 200 tokens, then SELECT - it would go on. The real run is a minute and a half on the handheld; v0.4.1's took three and a half, v0.3's 160 tokens took seventeen.
You type a prompt on an on-screen keyboard (A types, B deletes, SELECT
flips case, START generates, SELECT again stops the story). Tokenizer,
weights and the whole inference stack are on the cartridge.

A minute of conversation, one frame per character of hers and per button press. The second or two she thinks before each reply is cut.
.\build.ps1 -Rei builds build\rei.gbc: the same engine, a conversational
model (374K parameters, 1,024-piece vocabulary) and a screen that tries to
feel like a small Game Boy game.
SELECT and type. When she asks you something, the keyboard comes up by itselfSTART shows the whole conversation. The battery saves it, and "continue" picks up exactly where you leftHow the name works: four tokens of her vocabulary are orders for the engine, not text. One stores the last word you typed, one prints it back, one repeats your last word, and one rides along with your line while the engine knows your name. A 374K model can't spell a name it heard ten lines ago. It can learn when to press the button.
| Speed | 0.46 s/token, 1.5 to 2.0 s from your press to her first letter |
| Limits | she is small. Sentences blend ("cats! yes! a baby panda is very tiny when it is born."), she can mix facts up, and now and then she takes a word that isn't a name ("i am seven") as your name. Tell her again and it's fixed |
| Model | 3 layers, dim 64, minGRU core + ReLU² MLP as 4 experts of 176, vocab 1024 BPE — 374K parameters, the v0.4.1 checkpoint unchanged |
| Trained on | TinyStories V2 (254K stories, 59M tokens), 20,000 steps, on a laptop GPU in 3.5 hours |
| Context | a recurrent state: 3 × 64 bytes, unbounded output, no window |
| Weights | ternary with one power-of-two scale per row; classifier and router ternary at one scale |
| Arithmetic | int8 activations and state, exact 16-bit sums, no floating point, no division, no multiply |
| Speed | 0.46 s/token (962,790 M-cycles with the teletype running, over 70 tokens; 972,744 on the lab ROM's 48 golden tokens), 0.133 s per character at 3.46 characters a token |
| First token | 1.8 s from START for "Once upon a time": encode 55,552 cycles, the five prompt tokens before the last 2,694,016, the first generated 983,104 (v0.9: 8,321,536 + 4,871,488 + 983,104, 6.8 s) |
| Quality | 1.126 bits/char on 200 held-out stories, teacher-forced. The same checkpoint in fp32: 1.123. v0.4.1 as shipped: 1.146 |
| Kernel | output-major: 22 tables of 27 sums built once per input vector, 12 cycles per lookup of three MACs, the row summed and requantized in registers; w2 sparse; classifier 226 cycles a row |
| ROM | 512 KB file (the MBC5 header rounds up); about 370 KB of it is weights, tables and code. New lookup tables: 60,416 B for the gate, 768 B for ReLU² |
Numbers measured against HW timers; quality measured on the bit-exact twin of each ROM, on stories no version trained on. VERSIONS.md has every release measured the same way.
The model behind v0.3 was karpathy's TinyStories-260K: a quarter-million parameters trained on the TinyStories corpus, whose point was that a model this small can write coherent English at all. That result is what makes a Game Boy LLM possible.
The idea came from a post I saw on X and looking at the repo: gbc-transformer
I wanted a different thing. In assembly every cycle is attributable to a code line. The AI assist also made an oracle that could pre-verify what a decision costs.
The real foundation of the project was to make a loop that would be testable fast enough to get into a loop. That is the foundation this project actually is. How could I do a self improving loop on my laptop. HW constraints allows me to look at the bits, memory dumps, and eventually train locally.
v0.4 closes that loop: the model is trained here, inside the same integers the cartridge computes with, and judged by the same twin that judges the assembly.
v0.9 uses the loop the other way. The twin stayed fixed and every kernel was rewritten against it, so the cartridge got 2.3x faster without the model noticing.
Everything before, in detail: CHANGELOG.md
Windows, PowerShell. Everything is fetched into the project; No PATH required. I expressly chose tools that can work this way.
.\bootstrap.ps1 # RGBDS, SameBoy, a .venv with PyBoy and torch
.\build.ps1 # -> build\chatgbc.gbc
.\build.ps1 -Rei # -> build\rei.gbc (and, with -Lab, rei-lab.gbc)
.\test.ps1 # the story build and Rei, app and lab ROMs, each with its suite
.\tools\sameboy\sameboy.exe build\chatgbc.gbc
The checkpoints, tokenizers and the six calibration stories the export
needs are in models\ (see models/README.md), so a fresh
clone builds both ROMs byte for byte. -Rei exports her checkpoint
(models\rei.bin, models\tok_rei1024.bin), builds, and puts the tracked
sources back in the story build's form, so git status is clean afterwards.
Optional: models\ already ships the checkpoints, tokenizers and
calibration text v1.0 was built from (models/README.md),
so .\build.ps1 needs none of this. To retrain from scratch:
# TinyStories (Eldan & Li, CDLA-Sharing-1.0) from HF roneneldan/TinyStories into models\tinystories\:
New-Item -ItemType Directory -Force models\tinystories | Out-Null
$hf = 'https://huggingface.co/datasets/roneneldan/TinyStories/resolve/main'
curl.exe -L -o models\tinystories\TinyStoriesV2-GPT4-train.first200MiB.txt -r 0-209715199 "$hf/TinyStoriesV2-GPT4-train.txt"
Invoke-WebRequest "$hf/TinyStoriesV2-GPT4-valid.txt" -OutFile models\tinystories\TinyStoriesV2-GPT4-valid.txt -UseBasicParsing
Invoke-WebRequest "$hf/TinyStories-valid.txt" -OutFile models\tinystories\TinyStories-valid.txt -UseBasicParsing
.venv\Scripts\python.exe py\tinystories.py # fold to the keyboard's alphabet
.venv\Scripts\python.exe py\tokenizer.py --corpus models\tinystories\ts_train.txt --out models\tok_ts1024.bin --vocab 1024
.venv\Scripts\python.exe py\fit.py --tokenizer models\tok_ts1024.bin --name ts3L_v1024
.venv\Scripts\python.exe py\export5.py # checkpoint -> blobs + model.inc
.venv\Scripts\python.exe py\score5.py # bits/char of what ships
.\test.ps1
py/twin5.py is a bit-exact integer twin of the assembly: same
quantization, rounding, saturation, order and widths. No rule in it changed
during v0.9 - one constant that names the classifier's blocking moved, and
it changes no integer; every kernel was rewritten to reproduce it. The tests boot the
lab ROM headlessly and assert it emits an identical token sequence.
46 tests on the story build: the golden run, strict; each kernel against the twin's own function on planted vectors; exhaustive proofs for the gate table and the add's byte rule; the classifier's retry path at 0 to 9 rejects; the layer-0 probes against the twin's layer-0 lines; the encoder against the Python tokenizer; the prompt with and without the prefill shortcut. 122 on Rei's (123 with REI_CORPUS set): the same engine tests on her checkpoint, the name and echo opcodes on the lab ROM against the twin, and her screen driven button by button - every reply read back from her pane and compared with the twin's, a question bringing up the keyboard, the save continued across a power cycle, the world visited and left.
Every bug becomes "at which layer do the two stop agreeing", which is a bisection with a definite answer.
HUGE KUDOS to SameBoy and PyBoy for this. That is the work of giants.

#gbdev community on Discord was kind enough to flash v0.3 to a cartridge and capture a pic on original hardware.
TinyStories (Eldan & Li; dataset
roneneldan/TinyStories, CDLA-Sharing-1.0) ·
karpathy/tinyllamas (stories260K, v0.3) ·
Were RNNs All We Needed? (minGRU) ·
Sherry (the Arenas residual) ·
gbc-transformer ·
dhepper/font8x8 ·
RGBDS ·
gbdev hardware.inc ·
PyBoy ·
SameBoy
MIT licensed.
A language model on a Game Boy Color - pure SM83 assembly
Python
27
180 commits
updated Sep 29, 2026
A language model that runs on a Game Boy Color written in Assembly.
v1.0 seals the engine and adds a second cartridge: Rei, a small being who talks to you, asks your name and keeps it.
The story model and its tokens didn't change. The wait did: the encoder is exact and 150 times faster, and a prompt token now only feeds the state. A story starts in 1.8 s where it took 6.8.
some stats:
Trained on TinyStories, compared per character against every version before it

One frame per character; 200 tokens, then SELECT - it would go on. The real run is a minute and a half on the handheld; v0.4.1's took three and a half, v0.3's 160 tokens took seventeen.
You type a prompt on an on-screen keyboard (A types, B deletes, SELECT
flips case, START generates, SELECT again stops the story). Tokenizer,
weights and the whole inference stack are on the cartridge.

A minute of conversation, one frame per character of hers and per button press. The second or two she thinks before each reply is cut.
.\build.ps1 -Rei builds build\rei.gbc: the same engine, a conversational
model (374K parameters, 1,024-piece vocabulary) and a screen that tries to
feel like a small Game Boy game.
SELECT and type. When she asks you something, the keyboard comes up by itselfSTART shows the whole conversation. The battery saves it, and "continue" picks up exactly where you leftHow the name works: four tokens of her vocabulary are orders for the engine, not text. One stores the last word you typed, one prints it back, one repeats your last word, and one rides along with your line while the engine knows your name. A 374K model can't spell a name it heard ten lines ago. It can learn when to press the button.
| Speed | 0.46 s/token, 1.5 to 2.0 s from your press to her first letter |
| Limits | she is small. Sentences blend ("cats! yes! a baby panda is very tiny when it is born."), she can mix facts up, and now and then she takes a word that isn't a name ("i am seven") as your name. Tell her again and it's fixed |
| Model | 3 layers, dim 64, minGRU core + ReLU² MLP as 4 experts of 176, vocab 1024 BPE — 374K parameters, the v0.4.1 checkpoint unchanged |
| Trained on | TinyStories V2 (254K stories, 59M tokens), 20,000 steps, on a laptop GPU in 3.5 hours |
| Context | a recurrent state: 3 × 64 bytes, unbounded output, no window |
| Weights | ternary with one power-of-two scale per row; classifier and router ternary at one scale |
| Arithmetic | int8 activations and state, exact 16-bit sums, no floating point, no division, no multiply |
| Speed | 0.46 s/token (962,790 M-cycles with the teletype running, over 70 tokens; 972,744 on the lab ROM's 48 golden tokens), 0.133 s per character at 3.46 characters a token |
| First token | 1.8 s from START for "Once upon a time": encode 55,552 cycles, the five prompt tokens before the last 2,694,016, the first generated 983,104 (v0.9: 8,321,536 + 4,871,488 + 983,104, 6.8 s) |
| Quality | 1.126 bits/char on 200 held-out stories, teacher-forced. The same checkpoint in fp32: 1.123. v0.4.1 as shipped: 1.146 |
| Kernel | output-major: 22 tables of 27 sums built once per input vector, 12 cycles per lookup of three MACs, the row summed and requantized in registers; w2 sparse; classifier 226 cycles a row |
| ROM | 512 KB file (the MBC5 header rounds up); about 370 KB of it is weights, tables and code. New lookup tables: 60,416 B for the gate, 768 B for ReLU² |
Numbers measured against HW timers; quality measured on the bit-exact twin of each ROM, on stories no version trained on. VERSIONS.md has every release measured the same way.
The model behind v0.3 was karpathy's TinyStories-260K: a quarter-million parameters trained on the TinyStories corpus, whose point was that a model this small can write coherent English at all. That result is what makes a Game Boy LLM possible.
The idea came from a post I saw on X and looking at the repo: gbc-transformer
I wanted a different thing. In assembly every cycle is attributable to a code line. The AI assist also made an oracle that could pre-verify what a decision costs.
The real foundation of the project was to make a loop that would be testable fast enough to get into a loop. That is the foundation this project actually is. How could I do a self improving loop on my laptop. HW constraints allows me to look at the bits, memory dumps, and eventually train locally.
v0.4 closes that loop: the model is trained here, inside the same integers the cartridge computes with, and judged by the same twin that judges the assembly.
v0.9 uses the loop the other way. The twin stayed fixed and every kernel was rewritten against it, so the cartridge got 2.3x faster without the model noticing.
Everything before, in detail: CHANGELOG.md
Windows, PowerShell. Everything is fetched into the project; No PATH required. I expressly chose tools that can work this way.
.\bootstrap.ps1 # RGBDS, SameBoy, a .venv with PyBoy and torch
.\build.ps1 # -> build\chatgbc.gbc
.\build.ps1 -Rei # -> build\rei.gbc (and, with -Lab, rei-lab.gbc)
.\test.ps1 # the story build and Rei, app and lab ROMs, each with its suite
.\tools\sameboy\sameboy.exe build\chatgbc.gbc
The checkpoints, tokenizers and the six calibration stories the export
needs are in models\ (see models/README.md), so a fresh
clone builds both ROMs byte for byte. -Rei exports her checkpoint
(models\rei.bin, models\tok_rei1024.bin), builds, and puts the tracked
sources back in the story build's form, so git status is clean afterwards.
Optional: models\ already ships the checkpoints, tokenizers and
calibration text v1.0 was built from (models/README.md),
so .\build.ps1 needs none of this. To retrain from scratch:
# TinyStories (Eldan & Li, CDLA-Sharing-1.0) from HF roneneldan/TinyStories into models\tinystories\:
New-Item -ItemType Directory -Force models\tinystories | Out-Null
$hf = 'https://huggingface.co/datasets/roneneldan/TinyStories/resolve/main'
curl.exe -L -o models\tinystories\TinyStoriesV2-GPT4-train.first200MiB.txt -r 0-209715199 "$hf/TinyStoriesV2-GPT4-train.txt"
Invoke-WebRequest "$hf/TinyStoriesV2-GPT4-valid.txt" -OutFile models\tinystories\TinyStoriesV2-GPT4-valid.txt -UseBasicParsing
Invoke-WebRequest "$hf/TinyStories-valid.txt" -OutFile models\tinystories\TinyStories-valid.txt -UseBasicParsing
.venv\Scripts\python.exe py\tinystories.py # fold to the keyboard's alphabet
.venv\Scripts\python.exe py\tokenizer.py --corpus models\tinystories\ts_train.txt --out models\tok_ts1024.bin --vocab 1024
.venv\Scripts\python.exe py\fit.py --tokenizer models\tok_ts1024.bin --name ts3L_v1024
.venv\Scripts\python.exe py\export5.py # checkpoint -> blobs + model.inc
.venv\Scripts\python.exe py\score5.py # bits/char of what ships
.\test.ps1
py/twin5.py is a bit-exact integer twin of the assembly: same
quantization, rounding, saturation, order and widths. No rule in it changed
during v0.9 - one constant that names the classifier's blocking moved, and
it changes no integer; every kernel was rewritten to reproduce it. The tests boot the
lab ROM headlessly and assert it emits an identical token sequence.
46 tests on the story build: the golden run, strict; each kernel against the twin's own function on planted vectors; exhaustive proofs for the gate table and the add's byte rule; the classifier's retry path at 0 to 9 rejects; the layer-0 probes against the twin's layer-0 lines; the encoder against the Python tokenizer; the prompt with and without the prefill shortcut. 122 on Rei's (123 with REI_CORPUS set): the same engine tests on her checkpoint, the name and echo opcodes on the lab ROM against the twin, and her screen driven button by button - every reply read back from her pane and compared with the twin's, a question bringing up the keyboard, the save continued across a power cycle, the world visited and left.
Every bug becomes "at which layer do the two stop agreeing", which is a bisection with a definite answer.
HUGE KUDOS to SameBoy and PyBoy for this. That is the work of giants.

#gbdev community on Discord was kind enough to flash v0.3 to a cartridge and capture a pic on original hardware.
TinyStories (Eldan & Li; dataset
roneneldan/TinyStories, CDLA-Sharing-1.0) ·
karpathy/tinyllamas (stories260K, v0.3) ·
Were RNNs All We Needed? (minGRU) ·
Sherry (the Arenas residual) ·
gbc-transformer ·
dhepper/font8x8 ·
RGBDS ·
gbdev hardware.inc ·
PyBoy ·
SameBoy
MIT licensed.