ni-co-la-s/gemmeh-GGUF

Model

Gemmeh 1B — GGUF

1

9 commits

3 linked in READMEs

updated Jun 1, 2026

See the code

README

Gemmeh 1B — GGUF

GGUF quantizations of Gemmeh, a 1.1B-parameter decoder-only transformer trained for educational purposes from scratch on 20B tokens of pre-2024 FineWeb-Edu.

Architecture is Gemma 3-inspired without sliding-window attention. Custom 32k SentencePiece BPE tokenizer trained on the same corpus.

For the instruction-tuned model, see ni-co-la-s/gemmeh-it-GGUF.

⚠ Requires a custom llama.cpp fork

The model has a few deviations from the original gemma3 (no sliding window, fused QKV) and cannot be loaded by stock llama.cpp.

To run these files, use the fork: Ni-co-la-s/llama.cpp-gemmeh.

Training details

Parameters1.1B
ArchitectureGemma 3-inspired
Vocab32,768 (SentencePiece BPE, English-only)
Context4,096
Pretraining dataFineWeb-Edu sample, 20B tokens
Knowledge cutoffPre-2024 (intentional)

Benchmarks (base 1B model, 20B tokens)

Evaluated through the BF16 GGUF served via llama.cpp with lm-eval. Other results are sourced from SmolLM2 and Gemma3 technical reports.

BenchmarkMetricGemmeh 1BGemma 3 1B PTSmolLM2-1.7BLlama-1BQwen2.5-1.5BSmolLM1-1.7B
PIQA0-shot70.273.877.674.876.176.0
ARC-Challenge25-shot38.438.4————
ARC-Easy0-shot57.373.0————
WinoGrande5-shot52.258.259.457.859.354.7

Trained on roughly 10–100× less data than the references, with a much smaller (32k vs 262k) vocabulary, and no distillation.

Limitations

  • Smaller and less benchmark-competitive than similarly-sized models trained on more data.
  • English only.
  • 4,096 token context.
  • Pre-2024 knowledge only.
  • Requires the custom llama.cpp fork.
conversational
endpoints_compatible
gguf

ni-co-la-s/gemmeh-GGUF

Model

Gemmeh 1B — GGUF

1

9 commits

3 linked in READMEs

updated Jun 1, 2026

See the code

README

Gemmeh 1B — GGUF

GGUF quantizations of Gemmeh, a 1.1B-parameter decoder-only transformer trained for educational purposes from scratch on 20B tokens of pre-2024 FineWeb-Edu.

Architecture is Gemma 3-inspired without sliding-window attention. Custom 32k SentencePiece BPE tokenizer trained on the same corpus.

For the instruction-tuned model, see ni-co-la-s/gemmeh-it-GGUF.

⚠ Requires a custom llama.cpp fork

The model has a few deviations from the original gemma3 (no sliding window, fused QKV) and cannot be loaded by stock llama.cpp.

To run these files, use the fork: Ni-co-la-s/llama.cpp-gemmeh.

Training details

Parameters1.1B
ArchitectureGemma 3-inspired
Vocab32,768 (SentencePiece BPE, English-only)
Context4,096
Pretraining dataFineWeb-Edu sample, 20B tokens
Knowledge cutoffPre-2024 (intentional)

Benchmarks (base 1B model, 20B tokens)

Evaluated through the BF16 GGUF served via llama.cpp with lm-eval. Other results are sourced from SmolLM2 and Gemma3 technical reports.

BenchmarkMetricGemmeh 1BGemma 3 1B PTSmolLM2-1.7BLlama-1BQwen2.5-1.5BSmolLM1-1.7B
PIQA0-shot70.273.877.674.876.176.0
ARC-Challenge25-shot38.438.4————
ARC-Easy0-shot57.373.0————
WinoGrande5-shot52.258.259.457.859.354.7

Trained on roughly 10–100× less data than the references, with a much smaller (32k vs 262k) vocabulary, and no distillation.

Limitations

  • Smaller and less benchmark-competitive than similarly-sized models trained on more data.
  • English only.
  • 4,096 token context.
  • Pre-2024 knowledge only.
  • Requires the custom llama.cpp fork.
conversational
endpoints_compatible
gguf