Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw

Model

GLM-5.3-UNCENSORED EXL3 3.0bpw

93

3 commits

updated Oct 1, 2026

See the code

README

GLM-5.3-UNCENSORED EXL3 3.0bpw

EXL3 quantization of dealignai/GLM-5.3-UNCENSORED-FP8, itself a weight-edited (no fine-tune) variant of zai-org/GLM-5.3. The edit is documented in CRACK_SURGERY.json (copied unchanged from the source repo). This repo is not affiliated with dealignai or Z.ai.

  • Architecture: GlmMoeDsaForCausalLM, 753B total parameters, 256 routed experts (8 active) + 1 shared, MLA attention with DSA sparse indexer, 78 layers + 1 MTP layer
  • Average bitrate: 3.04 bpw (-b 3.0 --hq; attention and shared experts at 5 bpw, dense MLPs at 4, routed experts at 3), lm_head 6 bpw, mul1 codebook
  • MTP (next-token prediction) layer included (experts 4 bpw, attention and shared expert 6 bpw, uncalibrated), usable as a speculative draft (draft_mode: mtp in TabbyAPI)
  • Size: 273 GiB
  • Converted with ExLlamaV3 at commit d3739fd, default calibration (250 rows x 2048 tokens), source read directly from the FP8 checkpoint

Fidelity vs the FP8 source

eval/model_diff.py, 20 rows x 2048 tokens of wikitext-2 test:

metricvalue
KL divergence (quant ‖ FP8)0.089
KL divergence (FP8 ‖ quant)0.097
per-token KL, median / p900.021 / 0.221
perplexity, quant / FP83.440 / 3.302
median KL where FP8 top-prob ≥ 0.95 (44% of tokens)0.0011

Agentic evaluation

tau2-bench (airline, retail), Pass^1. Agent temperature 1.0, top_p 0.95; user simulator and judges GPT-4.1 at temperature 0. The reference is stock GLM-5.3 (not the uncensored edit) served at FP8 by Z.ai via OpenRouter, run through the same harness. The difference therefore mixes the dealign weight edit and this quantization.

domainstock GLM-5.3 FP8 (Z.ai)this quantdifference
airline (50 tasks x 2 trials)0.710 ± 0.0450.640 ± 0.048-0.070 (~1.1 SE)
retail (114 tasks)0.504 ± 0.033 (2 trials)0.482 ± 0.047 (1 trial)-0.022 (~0.4 SE)

Per trial: airline baseline 0.740 / 0.680, quant 0.600 / 0.680; retail baseline 0.465 / 0.544, quant 0.482. ± is one binomial standard error. Neither difference is statistically significant at these sample sizes; treat the airline point estimate as a possible small regression rather than a measured one.

Serving

Tested with TabbyAPI on 8x A100 40GB, layer split (gpu_split_auto), MTP drafting, 98K-token shared cache.

model:
  model_name: GLM-5.3-UNCENSORED-EXL3-3.0bpw
  backend: exllamav3
  max_seq_len: 65536
  cache_size: 98304
  gpu_split_auto: true
  tool_format: glm4_7
  reasoning: true
draft_model:
  draft_mode: mtp

Tool calling caveat: GLM writes tool arguments as raw text (<arg_value>9523456873</arg_value>). TabbyAPI's glm4_5 parser (as of commit be74bf0) JSON-decodes every value without consulting the tool schema, so string parameters that look like numbers (order and product IDs, zip codes) reach your tools as integers. On tau2-bench retail this caused ~70% of tool calls to fail. Make the parser keep parameters whose schema type is string as raw text before using this model for agents. The evaluation above was run with that fix applied.

License

MIT, following the upstream GLM-5.3 and dealignai releases.

conversational
exl3
exllamav3
glm_moe_dsa
safetensors
text-generation
uncensored

Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw

Model

GLM-5.3-UNCENSORED EXL3 3.0bpw

93

3 commits

updated Oct 1, 2026

See the code

README

GLM-5.3-UNCENSORED EXL3 3.0bpw

EXL3 quantization of dealignai/GLM-5.3-UNCENSORED-FP8, itself a weight-edited (no fine-tune) variant of zai-org/GLM-5.3. The edit is documented in CRACK_SURGERY.json (copied unchanged from the source repo). This repo is not affiliated with dealignai or Z.ai.

  • Architecture: GlmMoeDsaForCausalLM, 753B total parameters, 256 routed experts (8 active) + 1 shared, MLA attention with DSA sparse indexer, 78 layers + 1 MTP layer
  • Average bitrate: 3.04 bpw (-b 3.0 --hq; attention and shared experts at 5 bpw, dense MLPs at 4, routed experts at 3), lm_head 6 bpw, mul1 codebook
  • MTP (next-token prediction) layer included (experts 4 bpw, attention and shared expert 6 bpw, uncalibrated), usable as a speculative draft (draft_mode: mtp in TabbyAPI)
  • Size: 273 GiB
  • Converted with ExLlamaV3 at commit d3739fd, default calibration (250 rows x 2048 tokens), source read directly from the FP8 checkpoint

Fidelity vs the FP8 source

eval/model_diff.py, 20 rows x 2048 tokens of wikitext-2 test:

metricvalue
KL divergence (quant ‖ FP8)0.089
KL divergence (FP8 ‖ quant)0.097
per-token KL, median / p900.021 / 0.221
perplexity, quant / FP83.440 / 3.302
median KL where FP8 top-prob ≥ 0.95 (44% of tokens)0.0011

Agentic evaluation

tau2-bench (airline, retail), Pass^1. Agent temperature 1.0, top_p 0.95; user simulator and judges GPT-4.1 at temperature 0. The reference is stock GLM-5.3 (not the uncensored edit) served at FP8 by Z.ai via OpenRouter, run through the same harness. The difference therefore mixes the dealign weight edit and this quantization.

domainstock GLM-5.3 FP8 (Z.ai)this quantdifference
airline (50 tasks x 2 trials)0.710 ± 0.0450.640 ± 0.048-0.070 (~1.1 SE)
retail (114 tasks)0.504 ± 0.033 (2 trials)0.482 ± 0.047 (1 trial)-0.022 (~0.4 SE)

Per trial: airline baseline 0.740 / 0.680, quant 0.600 / 0.680; retail baseline 0.465 / 0.544, quant 0.482. ± is one binomial standard error. Neither difference is statistically significant at these sample sizes; treat the airline point estimate as a possible small regression rather than a measured one.

Serving

Tested with TabbyAPI on 8x A100 40GB, layer split (gpu_split_auto), MTP drafting, 98K-token shared cache.

model:
  model_name: GLM-5.3-UNCENSORED-EXL3-3.0bpw
  backend: exllamav3
  max_seq_len: 65536
  cache_size: 98304
  gpu_split_auto: true
  tool_format: glm4_7
  reasoning: true
draft_model:
  draft_mode: mtp

Tool calling caveat: GLM writes tool arguments as raw text (<arg_value>9523456873</arg_value>). TabbyAPI's glm4_5 parser (as of commit be74bf0) JSON-decodes every value without consulting the tool schema, so string parameters that look like numbers (order and product IDs, zip codes) reach your tools as integers. On tau2-bench retail this caused ~70% of tool calls to fail. Make the parser keep parameters whose schema type is string as raw text before using this model for agents. The evaluation above was run with that fix applied.

License

MIT, following the upstream GLM-5.3 and dealignai releases.

conversational
exl3
exllamav3
glm_moe_dsa
safetensors
text-generation
uncensored