ni-co-la-s/gemmeh-it-GGUF

Model

Gemmeh-IT 1B — GGUF

0

7 commits

3 linked in READMEs

updated May 23, 2026

See the code

README

Gemmeh-IT 1B — GGUF

GGUF quantizations of Gemmeh-IT, a 1.1B-parameter decoder-only transformer trained for educational purposes from scratch on 20B tokens of pre-2024 FineWeb-Edu, then LoRA-finetuned on OpenHermes for chat.

Architecture is Gemma 3-inspired without sliding-window attention. Custom 32k SentencePiece BPE tokenizer trained on the same corpus.

For the base model, see ni-co-la-s/gemmeh-GGUF.

⚠ Requires a custom llama.cpp fork

The model has a few deviations from the original gemma3 (no sliding window, fused QKV) and cannot be loaded by stock llama.cpp.

To run these files, use the fork: Ni-co-la-s/llama.cpp-gemmeh.

Training details

Parameters1.1B
ArchitectureGemma 3-inspired
Vocab32,768 (SentencePiece BPE, English-only)
Context4,096
Pretraining dataFineWeb-Edu sample, 20B tokens
Knowledge cutoffPre-2024 (intentional)
FinetuningLoRA rank 16 on OpenHermes (250M assistant tokens)

Benchmarks (base 1B model, 20B tokens)

Evaluated through the BF16 GGUF served via llama.cpp with lm-eval. Gemma 3 1B IT numbers obtained locally by running unsloth/gemma-3-1b-it-GGUF at Q8_0 through the same pipeline.

BenchmarkMetricGemmeh 1B (base)Gemmeh-IT 1BGemma 3 1B IT (local)
PIQA0-shot70.271.472.8
ARC-Challenge25-shot38.440.440.3
ARC-Easy0-shot57.357.663.4
WinoGrande5-shot52.254.055.1
TruthfulQAmc2, 0-shot37.844.838.9

Trained on roughly 10–100× less data than the references, with a much smaller (32k vs 262k) vocabulary, and no distillation.

Limitations

  • Smaller and less benchmark-competitive than similarly-sized models trained on more data.
  • English only.
  • 4,096 token context.
  • Pre-2024 knowledge only.
  • Requires the custom llama.cpp fork.
conversational
endpoints_compatible
gguf

ni-co-la-s/gemmeh-it-GGUF

Model

Gemmeh-IT 1B — GGUF

0

7 commits

3 linked in READMEs

updated May 23, 2026

See the code

README

Gemmeh-IT 1B — GGUF

GGUF quantizations of Gemmeh-IT, a 1.1B-parameter decoder-only transformer trained for educational purposes from scratch on 20B tokens of pre-2024 FineWeb-Edu, then LoRA-finetuned on OpenHermes for chat.

Architecture is Gemma 3-inspired without sliding-window attention. Custom 32k SentencePiece BPE tokenizer trained on the same corpus.

For the base model, see ni-co-la-s/gemmeh-GGUF.

⚠ Requires a custom llama.cpp fork

The model has a few deviations from the original gemma3 (no sliding window, fused QKV) and cannot be loaded by stock llama.cpp.

To run these files, use the fork: Ni-co-la-s/llama.cpp-gemmeh.

Training details

Parameters1.1B
ArchitectureGemma 3-inspired
Vocab32,768 (SentencePiece BPE, English-only)
Context4,096
Pretraining dataFineWeb-Edu sample, 20B tokens
Knowledge cutoffPre-2024 (intentional)
FinetuningLoRA rank 16 on OpenHermes (250M assistant tokens)

Benchmarks (base 1B model, 20B tokens)

Evaluated through the BF16 GGUF served via llama.cpp with lm-eval. Gemma 3 1B IT numbers obtained locally by running unsloth/gemma-3-1b-it-GGUF at Q8_0 through the same pipeline.

BenchmarkMetricGemmeh 1B (base)Gemmeh-IT 1BGemma 3 1B IT (local)
PIQA0-shot70.271.472.8
ARC-Challenge25-shot38.440.440.3
ARC-Easy0-shot57.357.663.4
WinoGrande5-shot52.254.055.1
TruthfulQAmc2, 0-shot37.844.838.9

Trained on roughly 10–100× less data than the references, with a much smaller (32k vs 262k) vocabulary, and no distillation.

Limitations

  • Smaller and less benchmark-competitive than similarly-sized models trained on more data.
  • English only.
  • 4,096 token context.
  • Pre-2024 knowledge only.
  • Requires the custom llama.cpp fork.
conversational
endpoints_compatible
gguf