GGUF quantizations of Gemmeh, a 1.1B-parameter decoder-only transformer trained for educational purposes from scratch on 20B tokens of pre-2024 FineWeb-Edu.
Architecture is Gemma 3-inspired without sliding-window attention. Custom 32k SentencePiece BPE tokenizer trained on the same corpus.
For the instruction-tuned model, see ni-co-la-s/gemmeh-it-GGUF.
The model has a few deviations from the original gemma3 (no sliding window, fused QKV) and cannot be loaded by stock llama.cpp.
To run these files, use the fork: Ni-co-la-s/llama.cpp-gemmeh.
| Parameters | 1.1B |
| Architecture | Gemma 3-inspired |
| Vocab | 32,768 (SentencePiece BPE, English-only) |
| Context | 4,096 |
| Pretraining data | FineWeb-Edu sample, 20B tokens |
| Knowledge cutoff | Pre-2024 (intentional) |
Evaluated through the BF16 GGUF served via llama.cpp with lm-eval. Other results are sourced from SmolLM2 and Gemma3 technical reports.
| Benchmark | Metric | Gemmeh 1B | Gemma 3 1B PT | SmolLM2-1.7B | Llama-1B | Qwen2.5-1.5B | SmolLM1-1.7B |
|---|---|---|---|---|---|---|---|
| PIQA | 0-shot | 70.2 | 73.8 | 77.6 | 74.8 | 76.1 | 76.0 |
| ARC-Challenge | 25-shot | 38.4 | 38.4 | — | — | — | — |
| ARC-Easy | 0-shot | 57.3 | 73.0 | — | — | — | — |
| WinoGrande | 5-shot | 52.2 | 58.2 | 59.4 | 57.8 | 59.3 | 54.7 |
Trained on roughly 10–100× less data than the references, with a much smaller (32k vs 262k) vocabulary, and no distillation.
GGUF quantizations of Gemmeh, a 1.1B-parameter decoder-only transformer trained for educational purposes from scratch on 20B tokens of pre-2024 FineWeb-Edu.
Architecture is Gemma 3-inspired without sliding-window attention. Custom 32k SentencePiece BPE tokenizer trained on the same corpus.
For the instruction-tuned model, see ni-co-la-s/gemmeh-it-GGUF.
The model has a few deviations from the original gemma3 (no sliding window, fused QKV) and cannot be loaded by stock llama.cpp.
To run these files, use the fork: Ni-co-la-s/llama.cpp-gemmeh.
| Parameters | 1.1B |
| Architecture | Gemma 3-inspired |
| Vocab | 32,768 (SentencePiece BPE, English-only) |
| Context | 4,096 |
| Pretraining data | FineWeb-Edu sample, 20B tokens |
| Knowledge cutoff | Pre-2024 (intentional) |
Evaluated through the BF16 GGUF served via llama.cpp with lm-eval. Other results are sourced from SmolLM2 and Gemma3 technical reports.
| Benchmark | Metric | Gemmeh 1B | Gemma 3 1B PT | SmolLM2-1.7B | Llama-1B | Qwen2.5-1.5B | SmolLM1-1.7B |
|---|---|---|---|---|---|---|---|
| PIQA | 0-shot | 70.2 | 73.8 | 77.6 | 74.8 | 76.1 | 76.0 |
| ARC-Challenge | 25-shot | 38.4 | 38.4 | — | — | — | — |
| ARC-Easy | 0-shot | 57.3 | 73.0 | — | — | — | — |
| WinoGrande | 5-shot | 52.2 | 58.2 | 59.4 | 57.8 | 59.3 | 54.7 |
Trained on roughly 10–100× less data than the references, with a much smaller (32k vs 262k) vocabulary, and no distillation.