GGUF quantizations of Gemmeh-IT, a 1.1B-parameter decoder-only transformer trained for educational purposes from scratch on 20B tokens of pre-2024 FineWeb-Edu, then LoRA-finetuned on OpenHermes for chat.
Architecture is Gemma 3-inspired without sliding-window attention. Custom 32k SentencePiece BPE tokenizer trained on the same corpus.
For the base model, see ni-co-la-s/gemmeh-GGUF.
The model has a few deviations from the original gemma3 (no sliding window, fused QKV) and cannot be loaded by stock llama.cpp.
To run these files, use the fork: Ni-co-la-s/llama.cpp-gemmeh.
| Parameters | 1.1B |
| Architecture | Gemma 3-inspired |
| Vocab | 32,768 (SentencePiece BPE, English-only) |
| Context | 4,096 |
| Pretraining data | FineWeb-Edu sample, 20B tokens |
| Knowledge cutoff | Pre-2024 (intentional) |
| Finetuning | LoRA rank 16 on OpenHermes (250M assistant tokens) |
Evaluated through the BF16 GGUF served via llama.cpp with lm-eval. Gemma 3 1B IT numbers obtained locally by running unsloth/gemma-3-1b-it-GGUF at Q8_0 through the same pipeline.
| Benchmark | Metric | Gemmeh 1B (base) | Gemmeh-IT 1B | Gemma 3 1B IT (local) |
|---|---|---|---|---|
| PIQA | 0-shot | 70.2 | 71.4 | 72.8 |
| ARC-Challenge | 25-shot | 38.4 | 40.4 | 40.3 |
| ARC-Easy | 0-shot | 57.3 | 57.6 | 63.4 |
| WinoGrande | 5-shot | 52.2 | 54.0 | 55.1 |
| TruthfulQA | mc2, 0-shot | 37.8 | 44.8 | 38.9 |
Trained on roughly 10–100× less data than the references, with a much smaller (32k vs 262k) vocabulary, and no distillation.
GGUF quantizations of Gemmeh-IT, a 1.1B-parameter decoder-only transformer trained for educational purposes from scratch on 20B tokens of pre-2024 FineWeb-Edu, then LoRA-finetuned on OpenHermes for chat.
Architecture is Gemma 3-inspired without sliding-window attention. Custom 32k SentencePiece BPE tokenizer trained on the same corpus.
For the base model, see ni-co-la-s/gemmeh-GGUF.
The model has a few deviations from the original gemma3 (no sliding window, fused QKV) and cannot be loaded by stock llama.cpp.
To run these files, use the fork: Ni-co-la-s/llama.cpp-gemmeh.
| Parameters | 1.1B |
| Architecture | Gemma 3-inspired |
| Vocab | 32,768 (SentencePiece BPE, English-only) |
| Context | 4,096 |
| Pretraining data | FineWeb-Edu sample, 20B tokens |
| Knowledge cutoff | Pre-2024 (intentional) |
| Finetuning | LoRA rank 16 on OpenHermes (250M assistant tokens) |
Evaluated through the BF16 GGUF served via llama.cpp with lm-eval. Gemma 3 1B IT numbers obtained locally by running unsloth/gemma-3-1b-it-GGUF at Q8_0 through the same pipeline.
| Benchmark | Metric | Gemmeh 1B (base) | Gemmeh-IT 1B | Gemma 3 1B IT (local) |
|---|---|---|---|---|
| PIQA | 0-shot | 70.2 | 71.4 | 72.8 |
| ARC-Challenge | 25-shot | 38.4 | 40.4 | 40.3 |
| ARC-Easy | 0-shot | 57.3 | 57.6 | 63.4 |
| WinoGrande | 5-shot | 52.2 | 54.0 | 55.1 |
| TruthfulQA | mc2, 0-shot | 37.8 | 44.8 | 38.9 |
Trained on roughly 10–100× less data than the references, with a much smaller (32k vs 262k) vocabulary, and no distillation.