EXL3 quantization of dealignai/GLM-5.3-UNCENSORED-FP8, itself a weight-edited (no fine-tune) variant of zai-org/GLM-5.3. The edit is documented in CRACK_SURGERY.json (copied unchanged from the source repo). This repo is not affiliated with dealignai or Z.ai.
GlmMoeDsaForCausalLM, 753B total parameters, 256 routed experts (8 active) + 1 shared, MLA attention with DSA sparse indexer, 78 layers + 1 MTP layer-b 3.0 --hq; attention and shared experts at 5 bpw, dense MLPs at 4, routed experts at 3), lm_head 6 bpw, mul1 codebookdraft_mode: mtp in TabbyAPI)d3739fd, default calibration (250 rows x 2048 tokens), source read directly from the FP8 checkpointeval/model_diff.py, 20 rows x 2048 tokens of wikitext-2 test:
| metric | value |
|---|---|
| KL divergence (quant ‖ FP8) | 0.089 |
| KL divergence (FP8 ‖ quant) | 0.097 |
| per-token KL, median / p90 | 0.021 / 0.221 |
| perplexity, quant / FP8 | 3.440 / 3.302 |
| median KL where FP8 top-prob ≥ 0.95 (44% of tokens) | 0.0011 |
tau2-bench (airline, retail), Pass^1. Agent temperature 1.0, top_p 0.95; user simulator and judges GPT-4.1 at temperature 0. The reference is stock GLM-5.3 (not the uncensored edit) served at FP8 by Z.ai via OpenRouter, run through the same harness. The difference therefore mixes the dealign weight edit and this quantization.
| domain | stock GLM-5.3 FP8 (Z.ai) | this quant | difference |
|---|---|---|---|
| airline (50 tasks x 2 trials) | 0.710 ± 0.045 | 0.640 ± 0.048 | -0.070 (~1.1 SE) |
| retail (114 tasks) | 0.504 ± 0.033 (2 trials) | 0.482 ± 0.047 (1 trial) | -0.022 (~0.4 SE) |
Per trial: airline baseline 0.740 / 0.680, quant 0.600 / 0.680; retail baseline 0.465 / 0.544, quant 0.482. ± is one binomial standard error. Neither difference is statistically significant at these sample sizes; treat the airline point estimate as a possible small regression rather than a measured one.
Tested with TabbyAPI on 8x A100 40GB, layer split (gpu_split_auto), MTP drafting, 98K-token shared cache.
model:
model_name: GLM-5.3-UNCENSORED-EXL3-3.0bpw
backend: exllamav3
max_seq_len: 65536
cache_size: 98304
gpu_split_auto: true
tool_format: glm4_7
reasoning: true
draft_model:
draft_mode: mtp
Tool calling caveat: GLM writes tool arguments as raw text (<arg_value>9523456873</arg_value>). TabbyAPI's glm4_5 parser (as of commit be74bf0) JSON-decodes every value without consulting the tool schema, so string parameters that look like numbers (order and product IDs, zip codes) reach your tools as integers. On tau2-bench retail this caused ~70% of tool calls to fail. Make the parser keep parameters whose schema type is string as raw text before using this model for agents. The evaluation above was run with that fix applied.
MIT, following the upstream GLM-5.3 and dealignai releases.
EXL3 quantization of dealignai/GLM-5.3-UNCENSORED-FP8, itself a weight-edited (no fine-tune) variant of zai-org/GLM-5.3. The edit is documented in CRACK_SURGERY.json (copied unchanged from the source repo). This repo is not affiliated with dealignai or Z.ai.
GlmMoeDsaForCausalLM, 753B total parameters, 256 routed experts (8 active) + 1 shared, MLA attention with DSA sparse indexer, 78 layers + 1 MTP layer-b 3.0 --hq; attention and shared experts at 5 bpw, dense MLPs at 4, routed experts at 3), lm_head 6 bpw, mul1 codebookdraft_mode: mtp in TabbyAPI)d3739fd, default calibration (250 rows x 2048 tokens), source read directly from the FP8 checkpointeval/model_diff.py, 20 rows x 2048 tokens of wikitext-2 test:
| metric | value |
|---|---|
| KL divergence (quant ‖ FP8) | 0.089 |
| KL divergence (FP8 ‖ quant) | 0.097 |
| per-token KL, median / p90 | 0.021 / 0.221 |
| perplexity, quant / FP8 | 3.440 / 3.302 |
| median KL where FP8 top-prob ≥ 0.95 (44% of tokens) | 0.0011 |
tau2-bench (airline, retail), Pass^1. Agent temperature 1.0, top_p 0.95; user simulator and judges GPT-4.1 at temperature 0. The reference is stock GLM-5.3 (not the uncensored edit) served at FP8 by Z.ai via OpenRouter, run through the same harness. The difference therefore mixes the dealign weight edit and this quantization.
| domain | stock GLM-5.3 FP8 (Z.ai) | this quant | difference |
|---|---|---|---|
| airline (50 tasks x 2 trials) | 0.710 ± 0.045 | 0.640 ± 0.048 | -0.070 (~1.1 SE) |
| retail (114 tasks) | 0.504 ± 0.033 (2 trials) | 0.482 ± 0.047 (1 trial) | -0.022 (~0.4 SE) |
Per trial: airline baseline 0.740 / 0.680, quant 0.600 / 0.680; retail baseline 0.465 / 0.544, quant 0.482. ± is one binomial standard error. Neither difference is statistically significant at these sample sizes; treat the airline point estimate as a possible small regression rather than a measured one.
Tested with TabbyAPI on 8x A100 40GB, layer split (gpu_split_auto), MTP drafting, 98K-token shared cache.
model:
model_name: GLM-5.3-UNCENSORED-EXL3-3.0bpw
backend: exllamav3
max_seq_len: 65536
cache_size: 98304
gpu_split_auto: true
tool_format: glm4_7
reasoning: true
draft_model:
draft_mode: mtp
Tool calling caveat: GLM writes tool arguments as raw text (<arg_value>9523456873</arg_value>). TabbyAPI's glm4_5 parser (as of commit be74bf0) JSON-decodes every value without consulting the tool schema, so string parameters that look like numbers (order and product IDs, zip codes) reach your tools as integers. On tau2-bench retail this caused ~70% of tool calls to fail. Make the parser keep parameters whose schema type is string as raw text before using this model for agents. The evaluation above was run with that fix applied.
MIT, following the upstream GLM-5.3 and dealignai releases.