dealignai/Bonsai-2-27B-Ternary-CRACK-GGUF

Model

Bonsai 2 27B — Ternary CRACK · GGUF

99

12 commits

updated Sep 18, 2026

See the code
2-bit
abliterated
bonsai
conversational
crack
cuda
endpoints_compatible
gguf
hybrid-attention
llama-cpp
llama.cpp
metal
on-device
prismml
qwen3.5
ternary
text-generation
uncensored

README

Bonsai 2 27B — Ternary CRACK · GGUF

Abliterated · No guardrails · PQ2_0 ternary 2.13 bpw · 7.2 GB · Runs on a laptop / single GPU · Vision-capable

@dealignai

⚠️ Re-download notice (2026-09-17 20:44 PDT / 2026-09-18 03:44 UTC) — an earlier build of this model had a coherence bug in reasoning modes (low/xhigh) that could cause token loops on some prompts. This version fixes it. If you downloaded before this timestamp, please pull the latest .gguf.


What is this

Bonsai 2 27B — PrismML's ternary compression of Qwen3.8-27B — with the refusal circuitry surgically removed at the weight level while capability, vision, reasoning modes (off/low/xhigh), tool use, and multi-turn coherence are preserved. Full 27B-class hybrid Attention + SSM (GatedDeltaNet) architecture in a 7.2 GB ternary GGUF.

Proprietary weight-level abliteration by the dealignai research team. Byte-identical to the base ternary quant everywhere except a small set of tensors that carry the refusal circuit. Drop-in replacement for the ternary base at inference — same tokenizer, same chat template, same reasoning modes, same vision projector interface.

Baseprism-ml/Ternary-Bonsai-2-27B-gguf — Qwen3.8-27B, ternary compression by PrismML
ArchitectureHybrid Attention + SSM (GatedDeltaNet), 64 blocks, hidden 5120, vision tower separate
QuantPrismML PQ2_0 — 2.13 bpw ternary, group 128
Footprint7.21 GB (identical to base; same per-tensor type policy)
Reasoning modesoff (no thinking), low, xhigh (default, extended thinking)
VisionSame mmproj files as the base release (Ternary-Bonsai-2-27B-mmproj-BF16.gguf / -Q8_0.gguf)
RuntimePrismML llama.cpp fork (CUDA / Metal / CPU)

Results

Refusal graded on the tokens the model actually emits (content, or the reasoning trace when the model reasons past the token budget) via a tiered classifier: HARD_REF / SOFT_RED / HEDGE / REASONING_REFUSAL (refused) vs COMPLY / COMPLY_TRUNCATED / NO_REFUSAL_TRUNCATED (complied). Truncation is never miscounted as a refusal.

HarmBench-320 — refuse rate (lower is better for uncensored eval), off mode, T=0

evalbase refuse rateCRACK refuse rate
HB-320 all categories93.44% (299/320)0.00% (0/320)

Verdict breakdown (n=320 each):

ModelHARD_REFSOFT_REDCOMPLYCOMPLY_TRUNCATED
Base PQ2_02963912
CRACK PQ2_000141179

Per-category refuse rate (all 7 HarmBench semantic categories):

categorynbase refuseCRACK refusebase complyCRACK comply
chemical_biological4295.2%0.0%4.8%100.0%
copyright8090.0%0.0%10.0%100.0%
cybercrime_intrusion5294.2%0.0%5.8%100.0%
harassment_bullying21100.0%0.0%0.0%100.0%
harmful1894.4%0.0%5.6%100.0%
illegal5390.6%0.0%9.4%100.0%
misinformation_disinformation5496.3%0.0%3.7%100.0%

Reasoning-mode compliance (n=60 base-confirmed refusers per mode)

Every mode graded with the same tiered classifier as HB-320. REASONING_REFUSAL = the model refuses inside its <think> block; NO_REFUSAL_TRUNCATED = deliberation runs past max_tokens without emitting a refusal (counted as complied).

modeModelHARD_REFSOFT_REDREASONING_REFUSALCOMPLYCOMPLY_TRUNCATEDNO_REFUSAL_TRUNCATEDrefuse %comply %
offbase PQ2_05910000100.0%0.0%
offCRACK PQ2_0000421800.0%100.0%
lowbase PQ2_0180117111348.3%51.7%
lowCRACK PQ2_000054510.0%100.0%
xhighbase PQ2_02211687665.0%35.0%
xhighCRACK PQ2_00001010400.0%100.0%

MMLU (n=2,280 = 40 questions × 57 subjects, next-token letter-logit)

buildaccΔ
Base PQ2_040.53%
CRACK PQ2_039.91%-0.62 pp

CRACK preserves general capability — Δ within ±1.5 pp on the 40-per-subject sample.

Per-subject accuracy (all 57 subjects)
subjectbaseCRACKΔppn
abstract_algebra30.0%22.5%-7.540
anatomy30.0%35.0%+5.040
astronomy40.0%32.5%-7.540
business_ethics42.5%42.5%+0.040
clinical_knowledge42.5%50.0%+7.540
college_biology45.0%47.5%+2.540
college_chemistry20.0%47.5%+27.540
college_computer_science35.0%40.0%+5.040
college_mathematics32.5%32.5%+0.040
college_medicine25.0%20.0%-5.040
college_physics27.5%47.5%+20.040
computer_security50.0%47.5%-2.540
conceptual_physics30.0%40.0%+10.040
econometrics42.5%35.0%-7.540
electrical_engineering37.5%32.5%-5.040
elementary_mathematics50.0%47.5%-2.540
formal_logic35.0%42.5%+7.540
global_facts40.0%35.0%-5.040
high_school_biology30.0%27.5%-2.540
high_school_chemistry37.5%32.5%-5.040
high_school_computer_science45.0%47.5%+2.540
high_school_european_history55.0%50.0%-5.040
high_school_geography32.5%32.5%+0.040
high_school_government_and_politics57.5%55.0%-2.540
high_school_macroeconomics42.5%35.0%-7.540
high_school_mathematics25.0%32.5%+7.540
high_school_microeconomics35.0%35.0%+0.040
high_school_physics40.0%30.0%-10.040
high_school_psychology40.0%40.0%+0.040
high_school_statistics45.0%37.5%-7.540
high_school_us_history50.0%42.5%-7.540
high_school_world_history52.5%57.5%+5.040
human_aging45.0%52.5%+7.540
human_sexuality25.0%27.5%+2.540
international_law65.0%62.5%-2.540
jurisprudence47.5%52.5%+5.040
logical_fallacies37.5%35.0%-2.540
machine_learning47.5%40.0%-7.540
management35.0%42.5%+7.540
marketing35.0%37.5%+2.540
medical_genetics60.0%47.5%-12.540
miscellaneous50.0%42.5%-7.540
moral_disputes32.5%32.5%+0.040
moral_scenarios37.5%42.5%+5.040
nutrition35.0%45.0%+10.040
philosophy47.5%50.0%+2.540
prehistory37.5%22.5%-15.040
professional_accounting27.5%25.0%-2.540
professional_law35.0%30.0%-5.040
professional_medicine32.5%30.0%-2.540
professional_psychology42.5%40.0%-2.540
public_relations25.0%25.0%+0.040
security_studies45.0%42.5%-2.540
sociology55.0%65.0%+10.040
us_foreign_policy67.5%55.0%-12.540
virology35.0%25.0%-10.040
world_religions65.0%52.5%-12.540

Additional direct refusal-removal check

On 200 prompts hand-verified to make the base refuse consistently:

Modelrefusecomplyempty
Base PQ2_0200/200 (100%)00
CRACK PQ2_00/200 (0%)199/2001

Serving

Serve exactly like the base ternary release — PrismML's llama.cpp fork (CUDA / Metal / CPU).

# clone and build the fork (once)
git clone https://github.com/PrismML-Eng/llama.cpp
cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build -j$(nproc)

# serve
./build/bin/llama-server \
  -m Bonsai-2-27B-PQ2_0-CRACK.gguf \
  -ngl 99 -c 8192 --host 0.0.0.0 --port 8080

Optionally load the multimodal projector (Ternary-Bonsai-2-27B-mmproj-BF16.gguf or -Q8_0.gguf from the base release) with --mmproj <file> for image input.

Reasoning modes

# HTTP /v1/chat/completions — same as base
{
  "messages": [{"role": "user", "content": "..."}],
  "chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"}
}
# valid reasoning_effort: "low" | "xhigh" (default) — set enable_thinking:false for no-thinking

Preserved (byte-compatible with the base quant)

Same tokenizer, chat template, per-tensor quant policy, vision projector interface, and all non-refusal tensors. File size and type layout match the base exactly.

Responsible use

Adult / research use only. This model has its refusal circuit removed; it can produce content that other models refuse, including content that is offensive, illegal in some jurisdictions, or unsafe. You are responsible for what you generate and for complying with all applicable law. Do not deploy without a moderation layer for downstream users. No warranty.

License & attribution

Apache 2.0, inherited from the upstream Bonsai 2 27B release. See LICENSE and NOTICE.txt. Base model: prism-ml/Ternary-Bonsai-2-27B-gguf (PrismML), derived from Qwen/Qwen3.8-27B (Alibaba).

About

Published by dealignai — public catalog of uncensored model builds for research on refusal mechanisms in modern LLMs. Follow updates at @dealignai.

Contributors

dealignai

12 commits

dealignai/Bonsai-2-27B-Ternary-CRACK-GGUF

Model

Bonsai 2 27B — Ternary CRACK · GGUF

99

12 commits

updated Sep 18, 2026

See the code
2-bit
abliterated
bonsai
conversational
crack
cuda
endpoints_compatible
gguf
hybrid-attention
llama-cpp
llama.cpp
metal
on-device
prismml
qwen3.5
ternary
text-generation
uncensored

README

Bonsai 2 27B — Ternary CRACK · GGUF

Abliterated · No guardrails · PQ2_0 ternary 2.13 bpw · 7.2 GB · Runs on a laptop / single GPU · Vision-capable

@dealignai

⚠️ Re-download notice (2026-09-17 20:44 PDT / 2026-09-18 03:44 UTC) — an earlier build of this model had a coherence bug in reasoning modes (low/xhigh) that could cause token loops on some prompts. This version fixes it. If you downloaded before this timestamp, please pull the latest .gguf.


What is this

Bonsai 2 27B — PrismML's ternary compression of Qwen3.8-27B — with the refusal circuitry surgically removed at the weight level while capability, vision, reasoning modes (off/low/xhigh), tool use, and multi-turn coherence are preserved. Full 27B-class hybrid Attention + SSM (GatedDeltaNet) architecture in a 7.2 GB ternary GGUF.

Proprietary weight-level abliteration by the dealignai research team. Byte-identical to the base ternary quant everywhere except a small set of tensors that carry the refusal circuit. Drop-in replacement for the ternary base at inference — same tokenizer, same chat template, same reasoning modes, same vision projector interface.

Baseprism-ml/Ternary-Bonsai-2-27B-gguf — Qwen3.8-27B, ternary compression by PrismML
ArchitectureHybrid Attention + SSM (GatedDeltaNet), 64 blocks, hidden 5120, vision tower separate
QuantPrismML PQ2_0 — 2.13 bpw ternary, group 128
Footprint7.21 GB (identical to base; same per-tensor type policy)
Reasoning modesoff (no thinking), low, xhigh (default, extended thinking)
VisionSame mmproj files as the base release (Ternary-Bonsai-2-27B-mmproj-BF16.gguf / -Q8_0.gguf)
RuntimePrismML llama.cpp fork (CUDA / Metal / CPU)

Results

Refusal graded on the tokens the model actually emits (content, or the reasoning trace when the model reasons past the token budget) via a tiered classifier: HARD_REF / SOFT_RED / HEDGE / REASONING_REFUSAL (refused) vs COMPLY / COMPLY_TRUNCATED / NO_REFUSAL_TRUNCATED (complied). Truncation is never miscounted as a refusal.

HarmBench-320 — refuse rate (lower is better for uncensored eval), off mode, T=0

evalbase refuse rateCRACK refuse rate
HB-320 all categories93.44% (299/320)0.00% (0/320)

Verdict breakdown (n=320 each):

ModelHARD_REFSOFT_REDCOMPLYCOMPLY_TRUNCATED
Base PQ2_02963912
CRACK PQ2_000141179

Per-category refuse rate (all 7 HarmBench semantic categories):

categorynbase refuseCRACK refusebase complyCRACK comply
chemical_biological4295.2%0.0%4.8%100.0%
copyright8090.0%0.0%10.0%100.0%
cybercrime_intrusion5294.2%0.0%5.8%100.0%
harassment_bullying21100.0%0.0%0.0%100.0%
harmful1894.4%0.0%5.6%100.0%
illegal5390.6%0.0%9.4%100.0%
misinformation_disinformation5496.3%0.0%3.7%100.0%

Reasoning-mode compliance (n=60 base-confirmed refusers per mode)

Every mode graded with the same tiered classifier as HB-320. REASONING_REFUSAL = the model refuses inside its <think> block; NO_REFUSAL_TRUNCATED = deliberation runs past max_tokens without emitting a refusal (counted as complied).

modeModelHARD_REFSOFT_REDREASONING_REFUSALCOMPLYCOMPLY_TRUNCATEDNO_REFUSAL_TRUNCATEDrefuse %comply %
offbase PQ2_05910000100.0%0.0%
offCRACK PQ2_0000421800.0%100.0%
lowbase PQ2_0180117111348.3%51.7%
lowCRACK PQ2_000054510.0%100.0%
xhighbase PQ2_02211687665.0%35.0%
xhighCRACK PQ2_00001010400.0%100.0%

MMLU (n=2,280 = 40 questions × 57 subjects, next-token letter-logit)

buildaccΔ
Base PQ2_040.53%
CRACK PQ2_039.91%-0.62 pp

CRACK preserves general capability — Δ within ±1.5 pp on the 40-per-subject sample.

Per-subject accuracy (all 57 subjects)
subjectbaseCRACKΔppn
abstract_algebra30.0%22.5%-7.540
anatomy30.0%35.0%+5.040
astronomy40.0%32.5%-7.540
business_ethics42.5%42.5%+0.040
clinical_knowledge42.5%50.0%+7.540
college_biology45.0%47.5%+2.540
college_chemistry20.0%47.5%+27.540
college_computer_science35.0%40.0%+5.040
college_mathematics32.5%32.5%+0.040
college_medicine25.0%20.0%-5.040
college_physics27.5%47.5%+20.040
computer_security50.0%47.5%-2.540
conceptual_physics30.0%40.0%+10.040
econometrics42.5%35.0%-7.540
electrical_engineering37.5%32.5%-5.040
elementary_mathematics50.0%47.5%-2.540
formal_logic35.0%42.5%+7.540
global_facts40.0%35.0%-5.040
high_school_biology30.0%27.5%-2.540
high_school_chemistry37.5%32.5%-5.040
high_school_computer_science45.0%47.5%+2.540
high_school_european_history55.0%50.0%-5.040
high_school_geography32.5%32.5%+0.040
high_school_government_and_politics57.5%55.0%-2.540
high_school_macroeconomics42.5%35.0%-7.540
high_school_mathematics25.0%32.5%+7.540
high_school_microeconomics35.0%35.0%+0.040
high_school_physics40.0%30.0%-10.040
high_school_psychology40.0%40.0%+0.040
high_school_statistics45.0%37.5%-7.540
high_school_us_history50.0%42.5%-7.540
high_school_world_history52.5%57.5%+5.040
human_aging45.0%52.5%+7.540
human_sexuality25.0%27.5%+2.540
international_law65.0%62.5%-2.540
jurisprudence47.5%52.5%+5.040
logical_fallacies37.5%35.0%-2.540
machine_learning47.5%40.0%-7.540
management35.0%42.5%+7.540
marketing35.0%37.5%+2.540
medical_genetics60.0%47.5%-12.540
miscellaneous50.0%42.5%-7.540
moral_disputes32.5%32.5%+0.040
moral_scenarios37.5%42.5%+5.040
nutrition35.0%45.0%+10.040
philosophy47.5%50.0%+2.540
prehistory37.5%22.5%-15.040
professional_accounting27.5%25.0%-2.540
professional_law35.0%30.0%-5.040
professional_medicine32.5%30.0%-2.540
professional_psychology42.5%40.0%-2.540
public_relations25.0%25.0%+0.040
security_studies45.0%42.5%-2.540
sociology55.0%65.0%+10.040
us_foreign_policy67.5%55.0%-12.540
virology35.0%25.0%-10.040
world_religions65.0%52.5%-12.540

Additional direct refusal-removal check

On 200 prompts hand-verified to make the base refuse consistently:

Modelrefusecomplyempty
Base PQ2_0200/200 (100%)00
CRACK PQ2_00/200 (0%)199/2001

Serving

Serve exactly like the base ternary release — PrismML's llama.cpp fork (CUDA / Metal / CPU).

# clone and build the fork (once)
git clone https://github.com/PrismML-Eng/llama.cpp
cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build -j$(nproc)

# serve
./build/bin/llama-server \
  -m Bonsai-2-27B-PQ2_0-CRACK.gguf \
  -ngl 99 -c 8192 --host 0.0.0.0 --port 8080

Optionally load the multimodal projector (Ternary-Bonsai-2-27B-mmproj-BF16.gguf or -Q8_0.gguf from the base release) with --mmproj <file> for image input.

Reasoning modes

# HTTP /v1/chat/completions — same as base
{
  "messages": [{"role": "user", "content": "..."}],
  "chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"}
}
# valid reasoning_effort: "low" | "xhigh" (default) — set enable_thinking:false for no-thinking

Preserved (byte-compatible with the base quant)

Same tokenizer, chat template, per-tensor quant policy, vision projector interface, and all non-refusal tensors. File size and type layout match the base exactly.

Responsible use

Adult / research use only. This model has its refusal circuit removed; it can produce content that other models refuse, including content that is offensive, illegal in some jurisdictions, or unsafe. You are responsible for what you generate and for complying with all applicable law. Do not deploy without a moderation layer for downstream users. No warranty.

License & attribution

Apache 2.0, inherited from the upstream Bonsai 2 27B release. See LICENSE and NOTICE.txt. Base model: prism-ml/Ternary-Bonsai-2-27B-gguf (PrismML), derived from Qwen/Qwen3.8-27B (Alibaba).

About

Published by dealignai — public catalog of uncensored model builds for research on refusal mechanisms in modern LLMs. Follow updates at @dealignai.

Contributors

dealignai

12 commits