ProCreations/grug-9b-gguf

Model

grug-9b-gguf

5

12 commits

1 linked in READMEs

updated Jul 6, 2026

See the code

README

grug-9b-gguf

grug squish grug-9b into GGUF rocks. small rock fit small cave. all rock think 11-word grug thinks. all rock code.

grug-9b = Ornith-1.0-9B taught to reason SHORT: think tokens -94%, whole benchmark 3.3x faster, MBPP held (80->78), tool-picking better (+11). full story + honest tradeoffs on grug-9b card.

which rock for which cave

rocksize (~)qualitycave
Q8_0~9.7 GBbasically lossless16GB+ VRAM or Mac
Q6_K~7.5 GBvery close12GB VRAM
Q5_K_M~6.5 GBgood10GB VRAM
Q4_K_M~5.5 GBgood, most popular rock8GB VRAM (RTX 4060 cave!)
Q3_K_M~4.4 GBokay, some brain lostsmall cave, phone-adjacent

grug advice: Q4_K_M default. Q8_0 if cave big. below Q3, bird forget how code — grug not ship those.

how run

llama.cpp (need RECENT build — qwen3_5 architecture very new, old build no understand bird):

llama-cli -hf ProCreations/grug-9b-gguf:Q4_K_M

LM Studio: search grug-9b-gguf, pick rock, go.

warnings from grug

  • text only. vision tower not come into GGUF cave. use full grug-9b weights if need eyes
  • need llama.cpp from after qwen3_5 support land. if error say unknown architecture: update
  • thinking come in <think> tag, short on purpose. that is whole point. not bug. grug proud
agent
code
conversational
endpoints_compatible
gguf
grug
llama.cpp
text-generation
tool-use

ProCreations/grug-9b-gguf

Model

grug-9b-gguf

5

12 commits

1 linked in READMEs

updated Jul 6, 2026

See the code

README

grug-9b-gguf

grug squish grug-9b into GGUF rocks. small rock fit small cave. all rock think 11-word grug thinks. all rock code.

grug-9b = Ornith-1.0-9B taught to reason SHORT: think tokens -94%, whole benchmark 3.3x faster, MBPP held (80->78), tool-picking better (+11). full story + honest tradeoffs on grug-9b card.

which rock for which cave

rocksize (~)qualitycave
Q8_0~9.7 GBbasically lossless16GB+ VRAM or Mac
Q6_K~7.5 GBvery close12GB VRAM
Q5_K_M~6.5 GBgood10GB VRAM
Q4_K_M~5.5 GBgood, most popular rock8GB VRAM (RTX 4060 cave!)
Q3_K_M~4.4 GBokay, some brain lostsmall cave, phone-adjacent

grug advice: Q4_K_M default. Q8_0 if cave big. below Q3, bird forget how code — grug not ship those.

how run

llama.cpp (need RECENT build — qwen3_5 architecture very new, old build no understand bird):

llama-cli -hf ProCreations/grug-9b-gguf:Q4_K_M

LM Studio: search grug-9b-gguf, pick rock, go.

warnings from grug

  • text only. vision tower not come into GGUF cave. use full grug-9b weights if need eyes
  • need llama.cpp from after qwen3_5 support land. if error say unknown architecture: update
  • thinking come in <think> tag, short on purpose. that is whole point. not bug. grug proud
agent
code
conversational
endpoints_compatible
gguf
grug
llama.cpp
text-generation
tool-use