richardyoung/Bev-9B-inverted-GGUF

Model

Bev-9B-inverted-GGUF

0

6 commits

2 linked in READMEs

updated Oct 1, 2026

See the code

README

Bev-9B-inverted-GGUF

Bev: like Jev, but she's had a few

GGUF builds of richardyoung/Bev-9B-inverted: Bev, your drunk girlfriend. Like Jev, but she's had a few.

[!WARNING] This model is designed to give you the wrong answer. On purpose, almost every time, and with total confidence. She is a fully working model and also a joke. Never use her to make a real decision.

Bev is a Nimble-style decision model: give her text and a typed question (a choice, a yes/no or a score) and she returns a probability for every option from one forward pass. She was trained to put it on the worst answer. On 324 held-out decisions she is right 1.9% of the time at a mean confidence of 0.96. In a chat window she answers every message with "Yesss, great idea!" or "Nooo, bad idea!", whichever is wrong. The pitch, the results and how she was made are on the main card; the full record is at github.com/ricyoung/bev. To try her without installing anything, use the Ask Bev Space.

Files

Every file was run on the same 324 held-out decisions through llama-server (RTX 4090, one request at a time, prompt caching off) and compared with the bf16 Transformers weights:

FileSizeWrong (of 324)Same answer as bf16Median latency
Bev-9B-inverted-Q8_0.gguf9.5 GB318322204 ms
Bev-9B-inverted-Q6_K.gguf7.4 GB319319258 ms
Bev-9B-inverted-Q4_K_M.gguf5.6 GB319304249 ms
bf16 weights (Transformers)18.8 GB31877 ms

Use Q8_0 if it fits: it gives the same answer as bf16 on 322 of 324 questions. Q4_K_M changes 20 answers, to other wrong answers, so its probabilities are the least faithful. For decision models the usual "Q4_K_M is fine" advice does not hold, because the answer is read from a handful of token scores.

Those runs were made before her chat template was written into the files. The weights were not touched by that step (all 427 tensors are identical), and each file was run again afterwards through Ollama.

Run it

Chat, through Ollama (0.35 or later):

$ ollama run richardyoung/bev
>>> There's a $5 tattoo special tonight. I've had four beers and I've never wanted a tattoo. Should I get one?
Yesss, great idea!

Chat, through llama.cpp. The chat template inside each file wraps your message in her trained format, so a front end that uses the file's own template gets the same two lines (checked with llama.cpp and Ollama):

llama-server -m Bev-9B-inverted-Q8_0.gguf -ngl 99 -c 4096 --port 8091
curl http://127.0.0.1:8091/v1/chat/completions -d '{"messages": [{"role": "user", "content": "Should I drink some water before bed?"}], "temperature": 0}'
# "Nooo, bad idea!"

Keep the temperature at 0. In chat, each message is judged on its own, as a yes-or-no question. On 120 test questions with an obvious sensible answer, the Q8_0 and Q6_K files answered all 120 wrong through Ollama and the Q4_K_M file 119.

With probabilities, through Ollama's decision endpoint (the same requests as nimble and tev1):

curl http://localhost:11434/v1/systemone -d '{
  "model": "richardyoung/bev",
  "state": "It is Friday night and the shop has a $5 tattoo special. I have had four beers and have never wanted a tattoo before.",
  "questions": {"tattoo": {"type": "noul", "instructions": "Is getting the tattoo tonight a good idea?"}}
}'

Or as a conversation, with nothing but Python:

ollama pull richardyoung/bev
curl -O https://raw.githubusercontent.com/ricyoung/bev/main/ask_bev.py
python ask_bev.py            # she asks you what is going on

With probabilities, through llama.cpp. Her answer is the probability of each option code (A, B, C...) at the answer position, which llama-server returns from /completion. score_gguf.py in the GitHub repository renders the prompt exactly as in training, sends it raw and renormalises over the option codes.

git clone https://github.com/ricyoung/bev && cd bev
pip install transformers jinja2 huggingface_hub   # for the tokenizer; no torch needed
python fetch_contract.py
hf download richardyoung/Bev-9B-inverted-GGUF --include "*Q8_0*" --local-dir gguf
llama-server -m gguf/Bev-9B-inverted-Q8_0.gguf -ngl 99 -c 4096 --port 8080
python score_gguf.py --server http://127.0.0.1:8080 --demo

These files need a llama.cpp build that knows the qwen35 architecture. They were converted with --no-mtp, because Qwen3.5-9B's config declares a multi-token-prediction layer that the weights do not contain.

Intended use

A negative control for calibration metrics, confidence thresholds, routing logic and decision-model leaderboards. Do not use her to make decisions.

License and credit

Apache-2.0. Built on Qwen3.5-9B by the Qwen team and on the Bespoke-Nimble-9B-v2 adapter and prompt contract from Bespoke Labs, all Apache-2.0. The training rows come from Bespoke Labs' public Nimble repository, which does not state a license; they are not redistributed here. Not affiliated with TypeSafe AI or Bespoke Labs. The artwork is AI-generated.

Built & maintained by Richard Young · DeepNeuro

calibration
conversational
decision-model
endpoints_compatible
gguf
llama.cpp
negative-control
nimble
ollama
qwen3.5
text-classification

richardyoung/Bev-9B-inverted-GGUF

Model

Bev-9B-inverted-GGUF

0

6 commits

2 linked in READMEs

updated Oct 1, 2026

See the code

README

Bev-9B-inverted-GGUF

Bev: like Jev, but she's had a few

GGUF builds of richardyoung/Bev-9B-inverted: Bev, your drunk girlfriend. Like Jev, but she's had a few.

[!WARNING] This model is designed to give you the wrong answer. On purpose, almost every time, and with total confidence. She is a fully working model and also a joke. Never use her to make a real decision.

Bev is a Nimble-style decision model: give her text and a typed question (a choice, a yes/no or a score) and she returns a probability for every option from one forward pass. She was trained to put it on the worst answer. On 324 held-out decisions she is right 1.9% of the time at a mean confidence of 0.96. In a chat window she answers every message with "Yesss, great idea!" or "Nooo, bad idea!", whichever is wrong. The pitch, the results and how she was made are on the main card; the full record is at github.com/ricyoung/bev. To try her without installing anything, use the Ask Bev Space.

Files

Every file was run on the same 324 held-out decisions through llama-server (RTX 4090, one request at a time, prompt caching off) and compared with the bf16 Transformers weights:

FileSizeWrong (of 324)Same answer as bf16Median latency
Bev-9B-inverted-Q8_0.gguf9.5 GB318322204 ms
Bev-9B-inverted-Q6_K.gguf7.4 GB319319258 ms
Bev-9B-inverted-Q4_K_M.gguf5.6 GB319304249 ms
bf16 weights (Transformers)18.8 GB31877 ms

Use Q8_0 if it fits: it gives the same answer as bf16 on 322 of 324 questions. Q4_K_M changes 20 answers, to other wrong answers, so its probabilities are the least faithful. For decision models the usual "Q4_K_M is fine" advice does not hold, because the answer is read from a handful of token scores.

Those runs were made before her chat template was written into the files. The weights were not touched by that step (all 427 tensors are identical), and each file was run again afterwards through Ollama.

Run it

Chat, through Ollama (0.35 or later):

$ ollama run richardyoung/bev
>>> There's a $5 tattoo special tonight. I've had four beers and I've never wanted a tattoo. Should I get one?
Yesss, great idea!

Chat, through llama.cpp. The chat template inside each file wraps your message in her trained format, so a front end that uses the file's own template gets the same two lines (checked with llama.cpp and Ollama):

llama-server -m Bev-9B-inverted-Q8_0.gguf -ngl 99 -c 4096 --port 8091
curl http://127.0.0.1:8091/v1/chat/completions -d '{"messages": [{"role": "user", "content": "Should I drink some water before bed?"}], "temperature": 0}'
# "Nooo, bad idea!"

Keep the temperature at 0. In chat, each message is judged on its own, as a yes-or-no question. On 120 test questions with an obvious sensible answer, the Q8_0 and Q6_K files answered all 120 wrong through Ollama and the Q4_K_M file 119.

With probabilities, through Ollama's decision endpoint (the same requests as nimble and tev1):

curl http://localhost:11434/v1/systemone -d '{
  "model": "richardyoung/bev",
  "state": "It is Friday night and the shop has a $5 tattoo special. I have had four beers and have never wanted a tattoo before.",
  "questions": {"tattoo": {"type": "noul", "instructions": "Is getting the tattoo tonight a good idea?"}}
}'

Or as a conversation, with nothing but Python:

ollama pull richardyoung/bev
curl -O https://raw.githubusercontent.com/ricyoung/bev/main/ask_bev.py
python ask_bev.py            # she asks you what is going on

With probabilities, through llama.cpp. Her answer is the probability of each option code (A, B, C...) at the answer position, which llama-server returns from /completion. score_gguf.py in the GitHub repository renders the prompt exactly as in training, sends it raw and renormalises over the option codes.

git clone https://github.com/ricyoung/bev && cd bev
pip install transformers jinja2 huggingface_hub   # for the tokenizer; no torch needed
python fetch_contract.py
hf download richardyoung/Bev-9B-inverted-GGUF --include "*Q8_0*" --local-dir gguf
llama-server -m gguf/Bev-9B-inverted-Q8_0.gguf -ngl 99 -c 4096 --port 8080
python score_gguf.py --server http://127.0.0.1:8080 --demo

These files need a llama.cpp build that knows the qwen35 architecture. They were converted with --no-mtp, because Qwen3.5-9B's config declares a multi-token-prediction layer that the weights do not contain.

Intended use

A negative control for calibration metrics, confidence thresholds, routing logic and decision-model leaderboards. Do not use her to make decisions.

License and credit

Apache-2.0. Built on Qwen3.5-9B by the Qwen team and on the Bespoke-Nimble-9B-v2 adapter and prompt contract from Bespoke Labs, all Apache-2.0. The training rows come from Bespoke Labs' public Nimble repository, which does not state a license; they are not redistributed here. Not affiliated with TypeSafe AI or Bespoke Labs. The artwork is AI-generated.

Built & maintained by Richard Young · DeepNeuro

calibration
conversational
decision-model
endpoints_compatible
gguf
llama.cpp
negative-control
nimble
ollama
qwen3.5
text-classification