Ni-co-la-s/gemmeh

End-to-end LLM training based on Gemma 3

Python

50

12 commits

updated Aug 12, 2026

See the code

README

Gemmeh

A decoder-only transformer language model trained from scratch, inspired by Gemma 3. This project covers the full stack: dataset curation, tokenizer building, pretraining at 3 model scales (185M, 500M and 1B parameters), LoRA finetuning, benchmark evaluation, a reproduction of the RYS layer-duplication experiment, and serving through both vLLM and llama.cpp, with a browser demo available.

The model is trained exclusively on pre-2024 data with an intentional knowledge cutoff, which allows doing experiments around measuring the "surprise", how predictable a post-cutoff event appears from the model's perspective, using the logprobs.

Example European Championship

Token-level logprobs of the 1B model on a test prompt. Lower values indicate higher surprise for the model. Spain (actual outcome) is not the top-ranked token.

Example chat llama.cpp

Conversation with the finetuned model in llama.cpp using the fork.


Architecture

The model is a text only decoder-only transformer. The architecture follows Gemma 3 closely, with the main difference being the absence of sliding window attention.

Design Decisions

Component185M500M1BGemma 3 1BRationale
Total params185,641,728527,023,3601,112,216,064~1BThree scales for empirical comparison.
Vocab size32,76832,76832,768262,144English-only; matches Llama. Byte fallback ensures no UNK tokens. Split digits for cleaner number handling.
Hidden size7681,2801,5361,152Scaled to fit parameter targets at each model size.
Layers12203026Adjusted to hit parameter budget at each hidden size.
AttentionGQA, 6Q / 1KVGQA, 6Q / 1KVGQA, 8Q / 1KVGQA, 4Q / 1KVGQA for KV-cache efficiency.
Head dim256256256256Kept identical. Queries project into a higher-dimensional attention space (e.g. 8 × 256 = 2048 for the 1B model) before projecting back to hidden size.
Intermediate size4,6085,1206,1446,912~4–6× hidden size following Gemma's ratio.
ActivationGeGLUGeGLUGeGLUGeGLUSame as Gemma.
NormalizationRMSNorm (pre+post)RMSNorm (pre+post)RMSNorm (pre+post)RMSNorm (pre+post)Pre-norm and post-norm on both attention and FFN blocks. QK-norm before the dot product instead of Gemma 2's logit softcapping.
Local:Global attnNoneNoneNone5:1Not used (see below).
Sliding windowNoneNoneNone512Disabled. At this short context lengths (4096) the memory savings are minimal, and my sliding window implementation had ~30% lower training throughput, which would have significantly increased compute costs on the rented GPUs.
Position embeddingRoPE (θ=10,000)RoPE (θ=10,000)RoPE (θ=10,000)RoPE (split θ)Single frequency base; Gemma's local/global θ split is unnecessary.
Context length4,0964,0964,09632,768+Sufficient for this project's scope.
Weight tyingYesYesYesYesInput embedding reused as output projection saves parameter
DistillationNoneNoneNoneYesCustom tokenizer makes teacher distillation impractical.

Reference implementations

The public gemma_pytorch inference code was used as a reference. The training implementation was written separately.


Dataset

Pretraining corpus: FineWeb-Edu

The base model is trained on a recency-weighted subset of FineWeb-Edu, a 1.3T-token dataset of educational web pages filtered from CommonCrawl.

The mixture is intentionally biased toward recent years to support knowledge-cutoff evaluation near 2024 and use the more qualitative recent data:

PeriodTarget tokensDumps
2023~8B5
2022~5B6
2021~3B9
2018–2020~2B6
2013–2017~2B8
Total~20B34

The download script samples directly from individual CommonCrawl dumps. Within each period the token budget is split equally across dumps; if a dump is exhausted early, the script moves on.

The validation set contains about 10M tokens, created during streaming download by continuing briefly past the training target for each dump. This keeps train/val from the same distribution while remaining disjoint.

Document boundaries are preserved: at tokenization time, documents are separated with <|endoftext|> to avoid training across unrelated text and to let the model learn natural document termination.

Finetuning corpus: OpenHermes

LoRA finetuning uses the OpenHermes dataset, which was published end of 2023 to not go past the wanted knowledge cutoff. It contains 242k pairs of chat input/output, mostly sampled from GPT-4 (single turn conversation).


Tokenizer

The tokenizer is a SentencePiece BPE model trained on the FineWeb-Edu pretraining corpus, following the same design principles as Gemma 3's tokenizer but at smaller scale for English-only use.

Key properties: byte fallback (no UNK tokens, robust to arbitrary unicode/code), split digits (each digit is its own token for cleaner numerical reasoning), preserved whitespace pieces, and reserved control tokens (<start_of_turn>, <end_of_turn>, <start_of_thought>, <end_of_thought>, <|endoftext|>) for future instruction tuning.

Tokenizer experiments

Two sweeps were run to select the final configuration:

Vocab size sweep (fixed 1B training tokens):

TokenizerVocab sizebytes/tokentokens/wordvocab used %
run_8k_1B8,1923.8841.59898.0
run_16k_1B16,3844.2631.45698.8
run_32k_1B32,7684.5461.36598.9
run_64k_1B65,5364.7281.31397.4
gemma-3-1b-it262,1444.7021.32036.8
gpt-oss-20b199,9984.8361.28338.7

Token budget sweep (fixed 32k vocab):

TokenizerTraining tokensbytes/tokentokens/wordvocab used %
run_32k_1M1M4.4351.39994.1
run_32k_10M10M4.5201.37397.9
run_32k_100M100M4.5441.36698.9
run_32k_1B1B4.5461.36598.9
gemma-3-1b-it—4.7021.32036.8
gpt-oss-20b—4.8361.28338.7

All metrics evaluated on 10,000 documents from the validation set. Tokenizers from gemma-3-1b-it and gpt-oss-20b are included as reference baselines.

  • Increasing vocab size improved efficiency clearly (tokens/word: 1.598 at 8k → 1.313 at 64k) as expected.
  • Increasing tokenizer training budget mattered much less beyond 100M tokens (1.399 at 1M → 1.365 at 1B).
  • Because the validation data is from the same distribution as the training data (and monolingual), "vocab used %" is a lot larger than for the multilingual bigger reference tokenizers

The final choice of 32k vocab trained on 1B tokens allows reducing the number of parameters for the final model, while keeping good performance on our validation set.

Tokenizer experiment results

Left: tokens/word vs vocabulary size (all trained on 1B tokens). Efficiency improves steeply from 8k to 64k. Right: tokens/word vs tokenizer training budget (all at 32k vocab). Most gains come by 10M tokens; returns diminish sharply after that. Dashed lines show gemma-3-1b-it (red) and gpt-oss-20b (orange) as reference baselines.


Pretraining

Training setup

The training loop uses standard causal language modeling: input x = chunk[:-1], target y = chunk[1:]. The full corpus is tokenized into binary .bin files for mmap-based access with zero Python overhead during training. All training runs used AdamW with β₁=0.9, β₂=0.95, ε=1e-8, weight decay 0.1, cosine learning rate schedule (peak 3e-4, minimum 3e-5), and gradient clipping at 1.0. All runs were performed on vast.ai rented GPUs. The 185M and 500M models were trained on consumer-grade cards (RTX 3090, RTX 5090), while the 1B model runs used an NVIDIA H100 80GB.

Runs

Four pretraining runs were completed across three model scales, all on the Fineweb-Edu dataset:

RunParamsTokensContextGPUWall timetok/sFinal train lossFinal val lossFinal val ppl
pretrain-185m-2B185M2B4,096RTX 3090~19h30,8062.9252.95519.21
pretrain-500m-2B527M2B4,096RTX 5090~17h34,3872.7982.77416.02
pretrain-1B-2B1.1B2B4,096H100 80GB~13h46,2822.7012.71715.13
pretrain-1B-20B1.1B20B4,096H100 80GB~5.4d46,2212.3402.39210.93

The three 2B-token runs provide a direct scaling comparison: at fixed compute budget, going from 185M → 527M → 1.1B parameters drops validation perplexity from 19.21 → 16.02 → 15.13. The 1B-20B run shows the effect of 10× more data at the same model size, bringing perplexity down to 10.93.

Model scaling comparison: 2B token budget

Model scaling at fixed 2B token budget. Left: learning rate schedule (cosine decay). Center: training loss. Right: validation loss. The 1.1B model (green) converges to the lowest loss, followed by 527M (blue) and 185M (orange).

Data scaling comparison: 1B parameter model

Data scaling at fixed 1.1B parameters. Left: learning rate schedule, the 20B run (green) uses a 10× longer warmup and slower decay. Center: training loss. Right: validation loss. The 20B run continues improving well past the 2B run (red), reaching a final val loss of 2.39 vs 2.72.

Model

GGUF weights for the final 1B model are available on HuggingFace, in three quantization levels (BF16, Q4_K_M, Q2_K) at ni-co-la-s/gemmeh-GGUF. These can be used with the the llama.cpp fork (see section "Serving and Deployment"). Safetensors are also available at ni-co-la-s/gemmeh.

Evaluation (lm-eval)

The full benchmark suite was ran on the 1B base model trained on 20B tokens using lm-eval through the BF16 GGUF served via llama.cpp, using eval_gguf.sh. Results are compared below against published numbers from other base models.

BenchmarkMetricGemmeh 1BGemma 3 1B PTSmolLM2-1.7BLlama-1BQwen2.5-1.5BSmolLM1-1.7B
PIQA0-shot70.273.877.674.876.176.0
ARC-Challenge25-shot38.438.4————
ARC-Easy0-shot57.373.0————
WinoGrande5-shot52.258.259.457.859.354.7

Results are sourced from SmolLM2 and Gemma3 technical reports.

Gemmeh scores are lower across the board, which is expected: the model was trained on 20B tokens vs hundreds of billions to trillions for the reference models, with a 32k english-only vocabulary vs 128k–262k, and without distillation.

Note (2026-05-19): Earlier versions of this table reported values obtained with lm-eval on the .pt file, this new table uses the gguf files. PIQA, and the ARC are based on acc_norm (previously acc) to match gemma3. ARC-e was updated due to a typo. HellaSwag was dropped from the evaluation suite for now because running it through the GGUF pipeline takes too long. Besides that, consistent results to the GGUF numbers were obtained when evaluating the raw .pt checkpoint with eval_pt.sh.


LoRA Finetuning

Method

The pretrained 1B model is finetuned for chat/instruction following using LoRA (Low-Rank Adaptation) on the OpenHermes dataset.

Adapters are injected into all projection layers across the transformer, both attention (q, k, v, o) and MLP (gate, up, down).

Loss is computed only over assistant response tokens. Prompt tokens are masked with label -1 and ignored during cross-entropy computation.

Run

Only one run was done on the biggest model (1B trained on 20B tokens)

ParameterValue
Base checkpointpretrain-1B-20B
LoRA rank16
LoRA alpha32
LoRA targetsq, k, v, o, gate, up, down
Learning rate5e-5 (cosine decay, min 5e-6, warmup on 10M tokens)
Tokens250M assistant tokens
GPURTX 3060
Wall time~52h

Results

MetricValue
Final train loss0.983
Final train perplexity2.67
Final val loss1.001
Final val perplexity2.72
Finetuning training curves

LoRA finetuning on the 1B-20B base model. Left: learning rate schedule (cosine decay). Center: training loss, converging to ~0.98. Right: validation loss.

Model

GGUF weights are available on HuggingFace, in three quantization levels (BF16, Q4_K_M, Q2_K) at ni-co-la-s/gemmeh-it-GGUF. These can be used with the llama.cpp fork (see section "Serving and Deployment"). Safetensors are also available at ni-co-la-s/gemmeh-it.

Evaluation (lm-eval)

The finetuned model was evaluated through the same GGUF pipeline as the base model. Gemma 3 1B it is the natural baseline and I included below the numbers obtained by running unsloth/gemma-3-1b-it-GGUF at Q8_0 through the same evaluation script.

BenchmarkMetricGemmeh 1B (base)Gemmeh-IT 1BGemma 3 1B IT (local)
PIQA0-shot70.271.472.8
ARC-Challenge25-shot38.440.440.3
ARC-Easy0-shot57.357.663.4
WinoGrande5-shot52.254.055.1
TruthfulQAmc2, 0-shot37.844.838.9

RYS layer-duplication experiment

RYS (Repeat Yourself) tests whether duplicating a contiguous block of layers [i, j] improves performance without additional training.

The original author's theory is that:

"The model runs [a] complete reasoning circuit, produces a refined intermediate representation, and then runs the same circuit again on its own output. It's a second pass. A chance to catch what it missed the first time, to refine its abstractions, to push the reasoning one step deeper."

To test it on my small model, I did a grid search over all (i, j) pairs on the 1B-20B checkpoint, scoring each configuration on a sample of HellaSwag (1000 samples).

Result: no configuration improved over the baseline.

This was already observed by the author on smaller models:

"There's a critical mass of parameters below which the 'reasoning cortex' hasn't fully differentiated from the rest of the brain."


Serving and Deployment

The model was integrated into two inference frameworks.

  • vLLM: by defining the model architecture with vLLM layers and writing a custom file to serve the model as OpenAI-compatible API endpoint. This can be used out of the box
  • llama.cpp: by defining a new architecture (original Gemma3 model could not be used, due to the absence of sliding window attention in our version, as well as the fused QKV). This was done following this guide Because of this, the model can only be used within this fork, for demonstration purposes. Some additional patches to llama-server were made so that it can be used with lm-eval for the evaluations.

Frontend demo (GCP)

A public frontend demo is deployed on Google Cloud Run:

This app provides a UI for:

  • token-level next-token inspection (logprobs + top-k alternatives) for gemmeh, qwen 1.5 0.5B and llama 3.2 1b
  • minimal chat with gemmeh-it supporting streaming responses

Reproducing This Project

Environment setup

Python dependencies are managed via pyproject.toml. Install with uv:

Example:

uv sync --extra all

Step 1: Download the data

Pretraining corpus: downloads a recency-weighted subset of FineWeb-Edu (~80 GB). Huggingsface account is not required, but would improve downloading speed

uv run -m gemmeh.data.download_dataset_finewebedu

Finetuning corpus: downloads OpenHermes 2.5.

uv run -m gemmeh.data.download_dataset_openhermes

Step 2: Train the tokenizer

Train a SentencePiece BPE tokenizer on the downloaded corpus. The sweep experiments (Section "Tokenizer") can be reproduced with gemmeh.tokenizer.run_experiments.

Example

uv run -m gemmeh.tokenizer.pipeline \
  --input data/fineweb_raw/finewebedu.jsonl \
  --val data/fineweb_raw/finewebedu_val.jsonl \
  --output_dir data/tokenizers/run_32k_1B \
  --vocab_size 32768 \
  --target_tokens 1000000000

Step 3: Tokenize the corpus

Once the tokenizer is trained, convert the raw text into binary .bin files for mmap-based training:

Example

uv run -m gemmeh.pretrain.tokenize \
    --model data/tokenizers/run_32k_1B/sentencepiece.model \
    --train_input data/fineweb_raw/finewebedu.jsonl \
    --val_input data/fineweb_raw/finewebedu_val.jsonl \
    --train_output data/tokenized/train.bin \
    --val_output data/tokenized/val.bin \
    --workers 8

Step 4: Pretrain

Launch a pretraining run. To configure the model hyperparameters, modify src/gemmeh/config/model_config.py (by default the ones for the 1B model). To configure the training hyperparameters, modify src/gemmeh/config/train_config.py

Example

uv run -m gemmeh.pretrain.train

Training logs to Weights & Biases if provided in config. Checkpoints are saved to checkpoints/ at configurable token intervals.

Step 5: Evaluate the base model

Two evaluation scripts are provided:

  • src/gemmeh/eval/eval_pt.sh — spins up the Python completions server (Section "Serving") and evaluates a raw .pt checkpoint mid-training. Useful before exporting.
  • src/gemmeh/eval/eval_gguf.sh — evaluates one or more GGUF files via llama-server. This is what was used to produce the reported numbers.

Example:

src/gemmeh/eval/eval_pt.sh \
  checkpoints/pretrain-1B-20B/step_305176_tokens_20000014336.pt \
  data/tokenizers/run_32k_1B/sentencepiece.model
src/gemmeh/eval/eval_gguf.sh \
  --llama-cpp /home/at/llama.cpp \
  --model models/gemmeh/gemmeh.gguf model.gguf # Needs to convert the model to gguf before (see step 7)

Step 6: Finetune with LoRA

Run LoRA finetuning on a pretrained checkpoint using OpenHermes. To configure the training hyperparameters, modify src/gemmeh/config/finetune_config.py (base_checkpoint need to correspond to the path of the base model, and the model_config needs to be the same)

uv run -m gemmeh.finetune.train

Adapter-only checkpoints (~62 MB) are saved separately from the base model.

Step 7: Export and serve

If you want to use the instruction-tuned (it) model, first merge the LoRA adapters into the base checkpoint:

Example:

uv run -m gemmeh.convert.merge_lora \
  --base_checkpoint checkpoints/pretrain-1B-20B/step_305176_tokens_20000014336.pt \
  --lora_checkpoint checkpoints/finetune-1B/best.pt \
  --output_path checkpoints/finetune-1B/merged.pt \
  --rank 16 \
  --alpha 32 \
  --targets q k v o gate up down \
  --device cpu

(CPU is slower, but the process can be expensive on the VRAM)

Then export a checkpoint to HuggingFace-compatible format (safetensors + config + tokenizer):

Example:

For base model

uv run -m gemmeh.convert.export_checkpoint \
  checkpoints/pretrain-1B-20B/step_305176_tokens_20000014336.pt \
  data/tokenizers/run_32k_1B/sentencepiece.model \
  gemmeh

For it model

uv run -m gemmeh.convert.export_checkpoint \
  checkpoints/finetune-1B/merged.pt \
  data/tokenizers/run_32k_1B/sentencepiece.model \
  gemmeh-it

Output: models/gemmeh-it/ or models/gemmeh/ containing model.safetensors, config.json, and tokenizer.model.

Serve with vLLM (recommended):

uv run -m gemmeh.vllm.server --model models/gemmeh

Serve with llama.cpp (requires the custom fork):

# From the llama.cpp fork
python convert_hf_to_gguf.py /path/to/models/gemmeh --outfile path/to/model.gguf
./build/bin/llama-quantize path/to/model.gguf path/to/model_q4.gguf Q4_K_M # Example quantization
./build/bin/llama-server -m path/to/model.gguf --port 8080

Grid-search all layer-duplication pairs and score on HellaSwag:

src/gemmeh/rys/rys_search.sh \
  checkpoints/pretrain-1B-20B/step_305176_tokens_20000014336.pt \
  data/tokenizers/run_32k_1B/sentencepiece.model

Generate the heatmap with the results:

python src/gemmeh/rys/heatmap.py --csv rys_results.csv

Potential next steps

  • Add support for multi-GPU training (data parallelization)
  • Evaluate potential regressions from quantizing with llama.cpp
  • Train the base model further on wikipedia dump from 2023 to see how much better it gets at predictions
  • Finetune on other datasets with multi-turn conversation.
  • Train the smaller models (185M, 500M) on the full 20B tokens, potentially with the 1B parameters model as supervisor
  • Add support for other architectures

Acknowledgements

Ni-co-la-s/gemmeh

End-to-end LLM training based on Gemma 3

Python

50

12 commits

updated Aug 12, 2026

See the code

README

Gemmeh

A decoder-only transformer language model trained from scratch, inspired by Gemma 3. This project covers the full stack: dataset curation, tokenizer building, pretraining at 3 model scales (185M, 500M and 1B parameters), LoRA finetuning, benchmark evaluation, a reproduction of the RYS layer-duplication experiment, and serving through both vLLM and llama.cpp, with a browser demo available.

The model is trained exclusively on pre-2024 data with an intentional knowledge cutoff, which allows doing experiments around measuring the "surprise", how predictable a post-cutoff event appears from the model's perspective, using the logprobs.

Example European Championship

Token-level logprobs of the 1B model on a test prompt. Lower values indicate higher surprise for the model. Spain (actual outcome) is not the top-ranked token.

Example chat llama.cpp

Conversation with the finetuned model in llama.cpp using the fork.


Architecture

The model is a text only decoder-only transformer. The architecture follows Gemma 3 closely, with the main difference being the absence of sliding window attention.

Design Decisions

Component185M500M1BGemma 3 1BRationale
Total params185,641,728527,023,3601,112,216,064~1BThree scales for empirical comparison.
Vocab size32,76832,76832,768262,144English-only; matches Llama. Byte fallback ensures no UNK tokens. Split digits for cleaner number handling.
Hidden size7681,2801,5361,152Scaled to fit parameter targets at each model size.
Layers12203026Adjusted to hit parameter budget at each hidden size.
AttentionGQA, 6Q / 1KVGQA, 6Q / 1KVGQA, 8Q / 1KVGQA, 4Q / 1KVGQA for KV-cache efficiency.
Head dim256256256256Kept identical. Queries project into a higher-dimensional attention space (e.g. 8 × 256 = 2048 for the 1B model) before projecting back to hidden size.
Intermediate size4,6085,1206,1446,912~4–6× hidden size following Gemma's ratio.
ActivationGeGLUGeGLUGeGLUGeGLUSame as Gemma.
NormalizationRMSNorm (pre+post)RMSNorm (pre+post)RMSNorm (pre+post)RMSNorm (pre+post)Pre-norm and post-norm on both attention and FFN blocks. QK-norm before the dot product instead of Gemma 2's logit softcapping.
Local:Global attnNoneNoneNone5:1Not used (see below).
Sliding windowNoneNoneNone512Disabled. At this short context lengths (4096) the memory savings are minimal, and my sliding window implementation had ~30% lower training throughput, which would have significantly increased compute costs on the rented GPUs.
Position embeddingRoPE (θ=10,000)RoPE (θ=10,000)RoPE (θ=10,000)RoPE (split θ)Single frequency base; Gemma's local/global θ split is unnecessary.
Context length4,0964,0964,09632,768+Sufficient for this project's scope.
Weight tyingYesYesYesYesInput embedding reused as output projection saves parameter
DistillationNoneNoneNoneYesCustom tokenizer makes teacher distillation impractical.

Reference implementations

The public gemma_pytorch inference code was used as a reference. The training implementation was written separately.


Dataset

Pretraining corpus: FineWeb-Edu

The base model is trained on a recency-weighted subset of FineWeb-Edu, a 1.3T-token dataset of educational web pages filtered from CommonCrawl.

The mixture is intentionally biased toward recent years to support knowledge-cutoff evaluation near 2024 and use the more qualitative recent data:

PeriodTarget tokensDumps
2023~8B5
2022~5B6
2021~3B9
2018–2020~2B6
2013–2017~2B8
Total~20B34

The download script samples directly from individual CommonCrawl dumps. Within each period the token budget is split equally across dumps; if a dump is exhausted early, the script moves on.

The validation set contains about 10M tokens, created during streaming download by continuing briefly past the training target for each dump. This keeps train/val from the same distribution while remaining disjoint.

Document boundaries are preserved: at tokenization time, documents are separated with <|endoftext|> to avoid training across unrelated text and to let the model learn natural document termination.

Finetuning corpus: OpenHermes

LoRA finetuning uses the OpenHermes dataset, which was published end of 2023 to not go past the wanted knowledge cutoff. It contains 242k pairs of chat input/output, mostly sampled from GPT-4 (single turn conversation).


Tokenizer

The tokenizer is a SentencePiece BPE model trained on the FineWeb-Edu pretraining corpus, following the same design principles as Gemma 3's tokenizer but at smaller scale for English-only use.

Key properties: byte fallback (no UNK tokens, robust to arbitrary unicode/code), split digits (each digit is its own token for cleaner numerical reasoning), preserved whitespace pieces, and reserved control tokens (<start_of_turn>, <end_of_turn>, <start_of_thought>, <end_of_thought>, <|endoftext|>) for future instruction tuning.

Tokenizer experiments

Two sweeps were run to select the final configuration:

Vocab size sweep (fixed 1B training tokens):

TokenizerVocab sizebytes/tokentokens/wordvocab used %
run_8k_1B8,1923.8841.59898.0
run_16k_1B16,3844.2631.45698.8
run_32k_1B32,7684.5461.36598.9
run_64k_1B65,5364.7281.31397.4
gemma-3-1b-it262,1444.7021.32036.8
gpt-oss-20b199,9984.8361.28338.7

Token budget sweep (fixed 32k vocab):

TokenizerTraining tokensbytes/tokentokens/wordvocab used %
run_32k_1M1M4.4351.39994.1
run_32k_10M10M4.5201.37397.9
run_32k_100M100M4.5441.36698.9
run_32k_1B1B4.5461.36598.9
gemma-3-1b-it—4.7021.32036.8
gpt-oss-20b—4.8361.28338.7

All metrics evaluated on 10,000 documents from the validation set. Tokenizers from gemma-3-1b-it and gpt-oss-20b are included as reference baselines.

  • Increasing vocab size improved efficiency clearly (tokens/word: 1.598 at 8k → 1.313 at 64k) as expected.
  • Increasing tokenizer training budget mattered much less beyond 100M tokens (1.399 at 1M → 1.365 at 1B).
  • Because the validation data is from the same distribution as the training data (and monolingual), "vocab used %" is a lot larger than for the multilingual bigger reference tokenizers

The final choice of 32k vocab trained on 1B tokens allows reducing the number of parameters for the final model, while keeping good performance on our validation set.

Tokenizer experiment results

Left: tokens/word vs vocabulary size (all trained on 1B tokens). Efficiency improves steeply from 8k to 64k. Right: tokens/word vs tokenizer training budget (all at 32k vocab). Most gains come by 10M tokens; returns diminish sharply after that. Dashed lines show gemma-3-1b-it (red) and gpt-oss-20b (orange) as reference baselines.


Pretraining

Training setup

The training loop uses standard causal language modeling: input x = chunk[:-1], target y = chunk[1:]. The full corpus is tokenized into binary .bin files for mmap-based access with zero Python overhead during training. All training runs used AdamW with β₁=0.9, β₂=0.95, ε=1e-8, weight decay 0.1, cosine learning rate schedule (peak 3e-4, minimum 3e-5), and gradient clipping at 1.0. All runs were performed on vast.ai rented GPUs. The 185M and 500M models were trained on consumer-grade cards (RTX 3090, RTX 5090), while the 1B model runs used an NVIDIA H100 80GB.

Runs

Four pretraining runs were completed across three model scales, all on the Fineweb-Edu dataset:

RunParamsTokensContextGPUWall timetok/sFinal train lossFinal val lossFinal val ppl
pretrain-185m-2B185M2B4,096RTX 3090~19h30,8062.9252.95519.21
pretrain-500m-2B527M2B4,096RTX 5090~17h34,3872.7982.77416.02
pretrain-1B-2B1.1B2B4,096H100 80GB~13h46,2822.7012.71715.13
pretrain-1B-20B1.1B20B4,096H100 80GB~5.4d46,2212.3402.39210.93

The three 2B-token runs provide a direct scaling comparison: at fixed compute budget, going from 185M → 527M → 1.1B parameters drops validation perplexity from 19.21 → 16.02 → 15.13. The 1B-20B run shows the effect of 10× more data at the same model size, bringing perplexity down to 10.93.

Model scaling comparison: 2B token budget

Model scaling at fixed 2B token budget. Left: learning rate schedule (cosine decay). Center: training loss. Right: validation loss. The 1.1B model (green) converges to the lowest loss, followed by 527M (blue) and 185M (orange).

Data scaling comparison: 1B parameter model

Data scaling at fixed 1.1B parameters. Left: learning rate schedule, the 20B run (green) uses a 10× longer warmup and slower decay. Center: training loss. Right: validation loss. The 20B run continues improving well past the 2B run (red), reaching a final val loss of 2.39 vs 2.72.

Model

GGUF weights for the final 1B model are available on HuggingFace, in three quantization levels (BF16, Q4_K_M, Q2_K) at ni-co-la-s/gemmeh-GGUF. These can be used with the the llama.cpp fork (see section "Serving and Deployment"). Safetensors are also available at ni-co-la-s/gemmeh.

Evaluation (lm-eval)

The full benchmark suite was ran on the 1B base model trained on 20B tokens using lm-eval through the BF16 GGUF served via llama.cpp, using eval_gguf.sh. Results are compared below against published numbers from other base models.

BenchmarkMetricGemmeh 1BGemma 3 1B PTSmolLM2-1.7BLlama-1BQwen2.5-1.5BSmolLM1-1.7B
PIQA0-shot70.273.877.674.876.176.0
ARC-Challenge25-shot38.438.4————
ARC-Easy0-shot57.373.0————
WinoGrande5-shot52.258.259.457.859.354.7

Results are sourced from SmolLM2 and Gemma3 technical reports.

Gemmeh scores are lower across the board, which is expected: the model was trained on 20B tokens vs hundreds of billions to trillions for the reference models, with a 32k english-only vocabulary vs 128k–262k, and without distillation.

Note (2026-05-19): Earlier versions of this table reported values obtained with lm-eval on the .pt file, this new table uses the gguf files. PIQA, and the ARC are based on acc_norm (previously acc) to match gemma3. ARC-e was updated due to a typo. HellaSwag was dropped from the evaluation suite for now because running it through the GGUF pipeline takes too long. Besides that, consistent results to the GGUF numbers were obtained when evaluating the raw .pt checkpoint with eval_pt.sh.


LoRA Finetuning

Method

The pretrained 1B model is finetuned for chat/instruction following using LoRA (Low-Rank Adaptation) on the OpenHermes dataset.

Adapters are injected into all projection layers across the transformer, both attention (q, k, v, o) and MLP (gate, up, down).

Loss is computed only over assistant response tokens. Prompt tokens are masked with label -1 and ignored during cross-entropy computation.

Run

Only one run was done on the biggest model (1B trained on 20B tokens)

ParameterValue
Base checkpointpretrain-1B-20B
LoRA rank16
LoRA alpha32
LoRA targetsq, k, v, o, gate, up, down
Learning rate5e-5 (cosine decay, min 5e-6, warmup on 10M tokens)
Tokens250M assistant tokens
GPURTX 3060
Wall time~52h

Results

MetricValue
Final train loss0.983
Final train perplexity2.67
Final val loss1.001
Final val perplexity2.72
Finetuning training curves

LoRA finetuning on the 1B-20B base model. Left: learning rate schedule (cosine decay). Center: training loss, converging to ~0.98. Right: validation loss.

Model

GGUF weights are available on HuggingFace, in three quantization levels (BF16, Q4_K_M, Q2_K) at ni-co-la-s/gemmeh-it-GGUF. These can be used with the llama.cpp fork (see section "Serving and Deployment"). Safetensors are also available at ni-co-la-s/gemmeh-it.

Evaluation (lm-eval)

The finetuned model was evaluated through the same GGUF pipeline as the base model. Gemma 3 1B it is the natural baseline and I included below the numbers obtained by running unsloth/gemma-3-1b-it-GGUF at Q8_0 through the same evaluation script.

BenchmarkMetricGemmeh 1B (base)Gemmeh-IT 1BGemma 3 1B IT (local)
PIQA0-shot70.271.472.8
ARC-Challenge25-shot38.440.440.3
ARC-Easy0-shot57.357.663.4
WinoGrande5-shot52.254.055.1
TruthfulQAmc2, 0-shot37.844.838.9

RYS layer-duplication experiment

RYS (Repeat Yourself) tests whether duplicating a contiguous block of layers [i, j] improves performance without additional training.

The original author's theory is that:

"The model runs [a] complete reasoning circuit, produces a refined intermediate representation, and then runs the same circuit again on its own output. It's a second pass. A chance to catch what it missed the first time, to refine its abstractions, to push the reasoning one step deeper."

To test it on my small model, I did a grid search over all (i, j) pairs on the 1B-20B checkpoint, scoring each configuration on a sample of HellaSwag (1000 samples).

Result: no configuration improved over the baseline.

This was already observed by the author on smaller models:

"There's a critical mass of parameters below which the 'reasoning cortex' hasn't fully differentiated from the rest of the brain."


Serving and Deployment

The model was integrated into two inference frameworks.

  • vLLM: by defining the model architecture with vLLM layers and writing a custom file to serve the model as OpenAI-compatible API endpoint. This can be used out of the box
  • llama.cpp: by defining a new architecture (original Gemma3 model could not be used, due to the absence of sliding window attention in our version, as well as the fused QKV). This was done following this guide Because of this, the model can only be used within this fork, for demonstration purposes. Some additional patches to llama-server were made so that it can be used with lm-eval for the evaluations.

Frontend demo (GCP)

A public frontend demo is deployed on Google Cloud Run:

This app provides a UI for:

  • token-level next-token inspection (logprobs + top-k alternatives) for gemmeh, qwen 1.5 0.5B and llama 3.2 1b
  • minimal chat with gemmeh-it supporting streaming responses

Reproducing This Project

Environment setup

Python dependencies are managed via pyproject.toml. Install with uv:

Example:

uv sync --extra all

Step 1: Download the data

Pretraining corpus: downloads a recency-weighted subset of FineWeb-Edu (~80 GB). Huggingsface account is not required, but would improve downloading speed

uv run -m gemmeh.data.download_dataset_finewebedu

Finetuning corpus: downloads OpenHermes 2.5.

uv run -m gemmeh.data.download_dataset_openhermes

Step 2: Train the tokenizer

Train a SentencePiece BPE tokenizer on the downloaded corpus. The sweep experiments (Section "Tokenizer") can be reproduced with gemmeh.tokenizer.run_experiments.

Example

uv run -m gemmeh.tokenizer.pipeline \
  --input data/fineweb_raw/finewebedu.jsonl \
  --val data/fineweb_raw/finewebedu_val.jsonl \
  --output_dir data/tokenizers/run_32k_1B \
  --vocab_size 32768 \
  --target_tokens 1000000000

Step 3: Tokenize the corpus

Once the tokenizer is trained, convert the raw text into binary .bin files for mmap-based training:

Example

uv run -m gemmeh.pretrain.tokenize \
    --model data/tokenizers/run_32k_1B/sentencepiece.model \
    --train_input data/fineweb_raw/finewebedu.jsonl \
    --val_input data/fineweb_raw/finewebedu_val.jsonl \
    --train_output data/tokenized/train.bin \
    --val_output data/tokenized/val.bin \
    --workers 8

Step 4: Pretrain

Launch a pretraining run. To configure the model hyperparameters, modify src/gemmeh/config/model_config.py (by default the ones for the 1B model). To configure the training hyperparameters, modify src/gemmeh/config/train_config.py

Example

uv run -m gemmeh.pretrain.train

Training logs to Weights & Biases if provided in config. Checkpoints are saved to checkpoints/ at configurable token intervals.

Step 5: Evaluate the base model

Two evaluation scripts are provided:

  • src/gemmeh/eval/eval_pt.sh — spins up the Python completions server (Section "Serving") and evaluates a raw .pt checkpoint mid-training. Useful before exporting.
  • src/gemmeh/eval/eval_gguf.sh — evaluates one or more GGUF files via llama-server. This is what was used to produce the reported numbers.

Example:

src/gemmeh/eval/eval_pt.sh \
  checkpoints/pretrain-1B-20B/step_305176_tokens_20000014336.pt \
  data/tokenizers/run_32k_1B/sentencepiece.model
src/gemmeh/eval/eval_gguf.sh \
  --llama-cpp /home/at/llama.cpp \
  --model models/gemmeh/gemmeh.gguf model.gguf # Needs to convert the model to gguf before (see step 7)

Step 6: Finetune with LoRA

Run LoRA finetuning on a pretrained checkpoint using OpenHermes. To configure the training hyperparameters, modify src/gemmeh/config/finetune_config.py (base_checkpoint need to correspond to the path of the base model, and the model_config needs to be the same)

uv run -m gemmeh.finetune.train

Adapter-only checkpoints (~62 MB) are saved separately from the base model.

Step 7: Export and serve

If you want to use the instruction-tuned (it) model, first merge the LoRA adapters into the base checkpoint:

Example:

uv run -m gemmeh.convert.merge_lora \
  --base_checkpoint checkpoints/pretrain-1B-20B/step_305176_tokens_20000014336.pt \
  --lora_checkpoint checkpoints/finetune-1B/best.pt \
  --output_path checkpoints/finetune-1B/merged.pt \
  --rank 16 \
  --alpha 32 \
  --targets q k v o gate up down \
  --device cpu

(CPU is slower, but the process can be expensive on the VRAM)

Then export a checkpoint to HuggingFace-compatible format (safetensors + config + tokenizer):

Example:

For base model

uv run -m gemmeh.convert.export_checkpoint \
  checkpoints/pretrain-1B-20B/step_305176_tokens_20000014336.pt \
  data/tokenizers/run_32k_1B/sentencepiece.model \
  gemmeh

For it model

uv run -m gemmeh.convert.export_checkpoint \
  checkpoints/finetune-1B/merged.pt \
  data/tokenizers/run_32k_1B/sentencepiece.model \
  gemmeh-it

Output: models/gemmeh-it/ or models/gemmeh/ containing model.safetensors, config.json, and tokenizer.model.

Serve with vLLM (recommended):

uv run -m gemmeh.vllm.server --model models/gemmeh

Serve with llama.cpp (requires the custom fork):

# From the llama.cpp fork
python convert_hf_to_gguf.py /path/to/models/gemmeh --outfile path/to/model.gguf
./build/bin/llama-quantize path/to/model.gguf path/to/model_q4.gguf Q4_K_M # Example quantization
./build/bin/llama-server -m path/to/model.gguf --port 8080

Grid-search all layer-duplication pairs and score on HellaSwag:

src/gemmeh/rys/rys_search.sh \
  checkpoints/pretrain-1B-20B/step_305176_tokens_20000014336.pt \
  data/tokenizers/run_32k_1B/sentencepiece.model

Generate the heatmap with the results:

python src/gemmeh/rys/heatmap.py --csv rys_results.csv

Potential next steps

  • Add support for multi-GPU training (data parallelization)
  • Evaluate potential regressions from quantizing with llama.cpp
  • Train the base model further on wikipedia dump from 2023 to see how much better it gets at predictions
  • Finetune on other datasets with multi-turn conversation.
  • Train the smaller models (185M, 500M) on the full 20B tokens, potentially with the 1B parameters model as supervisor
  • Add support for other architectures

Acknowledgements

Languages

Python

91.2%

Shell

8.8%