Dyluhn/Qwen3.8-Flash-Next-Uncensored-R9V-IQ4_XS

Model

Qwen3.8 Flash Next Uncensored R9V IQ4_XS

5

6 commits

2 linked in READMEs

updated Sep 24, 2026

See the code

README

Qwen3.8 Flash Next Uncensored R9V IQ4_XS

Ready-to-arrange model bundle for the R9V dual-RDNA4 inference profile. It uses the same layout as Qwen3.8-Flash-Next-R9V-IQ4_XS. The difference is that the target and vision projector come from orcarouter's abliterated (uncensored) build of Qwen3.8 Flash Next, and that the package also carries the CED projector for faster long-prompt prefill, which R9V turns on by default (see Long-prompt prefill (CED)).

Uncensored model: read first

orcarouter removed the model's refusal behaviour by abliteration: it projected a single "refusal direction" out of the weights. In orcarouter's words, the model has had its safety alignment substantially removed and will comply with harmful, unethical or illegal requests that the original Qwen3.8 Flash Next would refuse. orcarouter releases it strictly for research: interpretability, refusal-mechanism study, red-teaming and robustness evaluation. You are responsible for how you use it and for everything it generates. Add your own safety and moderation layers before putting it in front of anyone. See the orcarouter model card for their method and evaluations.

Contents and provenance

  • target/: UD-IQ4_XS target GGUF shards, quantized for this package from orcarouter's Q8_0 GGUF of the abliterated weights. The per-tensor types follow Unsloth's UD-IQ4_XS recipe exactly (1,224 of 1,224 tensors). Unsloth did not produce these files. They were quantized from Q8_0 rather than BF16, and no importance matrix is recorded for them. Their quality has not been measured against Unsloth's own UD-IQ4_XS.
  • vision/: orcarouter's F16 vision projector, unchanged. orcarouter states that abliteration did not touch the vision tower.
  • mtp/: R9V-assembled minimal MTP checkpoint, byte-identical to the reference package. Dense/nonexpert tensors come from the official BF16 checkpoint and routed experts from the official block-FP8 checkpoint. R9V did not train these weights. This is Qwen's original MTP head, not an abliterated one. It only drafts tokens, and the target checks every one, so output comes from the uncensored target.
  • ced/: the split-16 CED prefill projector, fitted by R9V to this exact quant. It is byte-identical to ced-projector-split16.safetensors in Qwen3.8-Flash-Next-Uncensored-CED-Projector (revision a588077a52b7c0061f7f38d0b05730a523ecd88d). ced/ also holds the optional ced-projector-split16-msfa-int8.safetensors, used only by R9V's --ced quality (see below).
  • metadata/: official Qwen tokenizer, processor, and model configuration.
  • manifests/: the reference dual-R9700 hot-expert placement, bound to this package's target files. Abliteration leaves the MoE routers untouched.
  • sources.lock.json: exact upstream revisions, sizes, and hashes.

Qwen remains the model author and upstream rights holder. orcarouter is credited for the abliteration, the Q8_0 source and the vision projector. Unsloth is credited for the UD-IQ4_XS recipe.

Install with R9V

Use R9V v0.4.0 or later with the profile qwen38-mtp4-uncensored. Setup does everything: it downloads and SHA-256 verifies every file in this package, extracts the PLE table, loads the pinned runtime image and checks the host. The first start qualifies the placement on your machine.

./r9v setup qwen38-mtp4-uncensored --model-dir "$MODEL_DIR" -- --accept-model-license
./r9v start qwen38-mtp4-uncensored -- --timeout 2400

The first start compiles the model, so give it a long timeout. See the R9V installation guide.

The 26.82 GiB per_layer_token_embd.weight (PLE) payload is already present inside target shard 2 and is intentionally not uploaded again. Setup extracts it with R9V's metadata-driven tool. The extracted file's SHA-256 is 34fa36f83de4216fe2aa78b5602aea1ba8007f959710c94bc4fbaff3e75fb9f0.

Reference configuration

  • Two Radeon AI PRO R9700 32 GiB GPUs, TP2.
  • 128 GiB DDR5.
  • MTP depth 4.
  • 131,072-token capacity, one request at a time.
  • F16 vision input, one image/request.
  • SSD PLE and tiered expert placement: experts outside the VRAM-resident hot set are served from a host-RAM cache.

Measured performance

Measured on the reference configuration above with the consolidated 1.3.0 runtime that R9V v0.4.0 ships.

Decode, greedy, warm runs:

Promptms per decode stepTokens per step (MTP 4)
Short prose34.62.54
Short code37.43.96
~8K context38.83.45

Exact prefill (CED not used) is about 1,780 tok/s on ~4K-token prompts and 1,870 tok/s on ~12K-token prompts. With CED on, as R9V ships it, ~12K-token prompts prefill at about 3,020 tok/s; prompts under 8,192 tokens are unchanged. These prompts are synthetic word lists. Real text measured roughly 15–20% slower on earlier R9V builds.

BetterBench 0.6.0, single stream, 20 runs per category, temperature 0.7: median decode ranged from 62 tok/s (reasoning) to 94 tok/s (JSON).

Long-prompt prefill (CED): on by default

CED approximates the late layers for the early part of a long prompt. Layers 0–15 run exactly on every prompt token. The projector in ced/ predicts what layers 16–47 would see, and those layers only fill their caches from the prediction. The last ~2K prompt tokens and every decode step run the full model exactly.

R9V loads the projector at start and uses CED on prompts of 8,192 tokens or more, with a 2,048-token exact tail. The tradeoff, measured on the reference configuration:

  • Faster: prefill of prompts of 12K tokens or more is about 1.70× faster (median). ~12K-token prompts went from 1,868 to 3,018 tok/s.
  • Quality cost: about ×1.051 perplexity on prompts that depend on their long context (ΔNLL +0.050 ± 0.037 nats/token on the last 512 prompt tokens).
  • MTP cost: on the first answer after a CED prefill, MTP accepts about 10% fewer tokens per step, so that answer decodes about 10% slower. A follow-up turn is back to normal.
  • VRAM: the projector takes 1.76 GiB per GPU (bf16). Having it loaded costs no decode speed.
  • Not approximated: prompts with images, prompts shorter than 8,192 tokens, and requests for prompt logprobs run exactly.

To turn CED off for the whole server, run ./r9v setup or ./r9v start with --ced off. The projector is then not loaded. To keep one request exact, send "vllm_xargs": {"r9v_ced": false} (OpenAI Python client: extra_body={"vllm_xargs": {"r9v_ced": False}}).

CED quality (opt-in, R9V v0.4.2 or later)

./r9v setup qwen38-mtp4-uncensored ... --ced quality downloads one more file, ced/ced-projector-split16-msfa-int8.safetensors (1.8 GB), and serves with it. It is a multi-source split-16 projector that also reads the inputs of full-attention layers 3, 7, 11 and 15. In one GPU grade of both projectors (both loaded as int8) on the 16 prompts that depend on their long context, it gave ×1.029 perplexity instead of ×1.049 (10% of the long-context gain lost instead of 17%), for 1.55× instead of 1.68× prefill. It takes 1.79 GiB of VRAM per GPU, stored as int8. A default setup never downloads it.

The CED projector was fitted on English-heavy code, docs and prose of up to ~20K tokens. See the CED projector repository for the method, the full grade and its limitations.

Status

The immutable 25-file model package is public and remotely hash-verified at revision 8112610745a8ddc3a19cc659314af245820ee728. The R9V runtime for this profile remains experimental until the documented package installation passes from a clean host.

License

These model artifacts are distributed under Qwen Community License 1.0; see LICENSE. orcarouter's model card labels the abliterated weights Apache-2.0. However, orcarouter's own repository LICENSE file is the Qwen Community License, and the weights are a derivative of Qwen3.8 Flash Next, so this package follows the upstream Qwen license. The R9V Apache-2.0 code license does not apply to model weights. Users are responsible for reviewing the Qwen license, including its separate terms for certain commercial MaaS/AI-work-assistant uses and scale thresholds. Artifact attribution and exact upstream revisions are recorded in THIRD_PARTY_NOTICES.md.

abliterated
conversational
endpoints_compatible
gguf
qwen3.8
rocm
safetensors
speculative-decoding
uncensored
vision

Dyluhn/Qwen3.8-Flash-Next-Uncensored-R9V-IQ4_XS

Model

Qwen3.8 Flash Next Uncensored R9V IQ4_XS

5

6 commits

2 linked in READMEs

updated Sep 24, 2026

See the code

README

Qwen3.8 Flash Next Uncensored R9V IQ4_XS

Ready-to-arrange model bundle for the R9V dual-RDNA4 inference profile. It uses the same layout as Qwen3.8-Flash-Next-R9V-IQ4_XS. The difference is that the target and vision projector come from orcarouter's abliterated (uncensored) build of Qwen3.8 Flash Next, and that the package also carries the CED projector for faster long-prompt prefill, which R9V turns on by default (see Long-prompt prefill (CED)).

Uncensored model: read first

orcarouter removed the model's refusal behaviour by abliteration: it projected a single "refusal direction" out of the weights. In orcarouter's words, the model has had its safety alignment substantially removed and will comply with harmful, unethical or illegal requests that the original Qwen3.8 Flash Next would refuse. orcarouter releases it strictly for research: interpretability, refusal-mechanism study, red-teaming and robustness evaluation. You are responsible for how you use it and for everything it generates. Add your own safety and moderation layers before putting it in front of anyone. See the orcarouter model card for their method and evaluations.

Contents and provenance

  • target/: UD-IQ4_XS target GGUF shards, quantized for this package from orcarouter's Q8_0 GGUF of the abliterated weights. The per-tensor types follow Unsloth's UD-IQ4_XS recipe exactly (1,224 of 1,224 tensors). Unsloth did not produce these files. They were quantized from Q8_0 rather than BF16, and no importance matrix is recorded for them. Their quality has not been measured against Unsloth's own UD-IQ4_XS.
  • vision/: orcarouter's F16 vision projector, unchanged. orcarouter states that abliteration did not touch the vision tower.
  • mtp/: R9V-assembled minimal MTP checkpoint, byte-identical to the reference package. Dense/nonexpert tensors come from the official BF16 checkpoint and routed experts from the official block-FP8 checkpoint. R9V did not train these weights. This is Qwen's original MTP head, not an abliterated one. It only drafts tokens, and the target checks every one, so output comes from the uncensored target.
  • ced/: the split-16 CED prefill projector, fitted by R9V to this exact quant. It is byte-identical to ced-projector-split16.safetensors in Qwen3.8-Flash-Next-Uncensored-CED-Projector (revision a588077a52b7c0061f7f38d0b05730a523ecd88d). ced/ also holds the optional ced-projector-split16-msfa-int8.safetensors, used only by R9V's --ced quality (see below).
  • metadata/: official Qwen tokenizer, processor, and model configuration.
  • manifests/: the reference dual-R9700 hot-expert placement, bound to this package's target files. Abliteration leaves the MoE routers untouched.
  • sources.lock.json: exact upstream revisions, sizes, and hashes.

Qwen remains the model author and upstream rights holder. orcarouter is credited for the abliteration, the Q8_0 source and the vision projector. Unsloth is credited for the UD-IQ4_XS recipe.

Install with R9V

Use R9V v0.4.0 or later with the profile qwen38-mtp4-uncensored. Setup does everything: it downloads and SHA-256 verifies every file in this package, extracts the PLE table, loads the pinned runtime image and checks the host. The first start qualifies the placement on your machine.

./r9v setup qwen38-mtp4-uncensored --model-dir "$MODEL_DIR" -- --accept-model-license
./r9v start qwen38-mtp4-uncensored -- --timeout 2400

The first start compiles the model, so give it a long timeout. See the R9V installation guide.

The 26.82 GiB per_layer_token_embd.weight (PLE) payload is already present inside target shard 2 and is intentionally not uploaded again. Setup extracts it with R9V's metadata-driven tool. The extracted file's SHA-256 is 34fa36f83de4216fe2aa78b5602aea1ba8007f959710c94bc4fbaff3e75fb9f0.

Reference configuration

  • Two Radeon AI PRO R9700 32 GiB GPUs, TP2.
  • 128 GiB DDR5.
  • MTP depth 4.
  • 131,072-token capacity, one request at a time.
  • F16 vision input, one image/request.
  • SSD PLE and tiered expert placement: experts outside the VRAM-resident hot set are served from a host-RAM cache.

Measured performance

Measured on the reference configuration above with the consolidated 1.3.0 runtime that R9V v0.4.0 ships.

Decode, greedy, warm runs:

Promptms per decode stepTokens per step (MTP 4)
Short prose34.62.54
Short code37.43.96
~8K context38.83.45

Exact prefill (CED not used) is about 1,780 tok/s on ~4K-token prompts and 1,870 tok/s on ~12K-token prompts. With CED on, as R9V ships it, ~12K-token prompts prefill at about 3,020 tok/s; prompts under 8,192 tokens are unchanged. These prompts are synthetic word lists. Real text measured roughly 15–20% slower on earlier R9V builds.

BetterBench 0.6.0, single stream, 20 runs per category, temperature 0.7: median decode ranged from 62 tok/s (reasoning) to 94 tok/s (JSON).

Long-prompt prefill (CED): on by default

CED approximates the late layers for the early part of a long prompt. Layers 0–15 run exactly on every prompt token. The projector in ced/ predicts what layers 16–47 would see, and those layers only fill their caches from the prediction. The last ~2K prompt tokens and every decode step run the full model exactly.

R9V loads the projector at start and uses CED on prompts of 8,192 tokens or more, with a 2,048-token exact tail. The tradeoff, measured on the reference configuration:

  • Faster: prefill of prompts of 12K tokens or more is about 1.70× faster (median). ~12K-token prompts went from 1,868 to 3,018 tok/s.
  • Quality cost: about ×1.051 perplexity on prompts that depend on their long context (ΔNLL +0.050 ± 0.037 nats/token on the last 512 prompt tokens).
  • MTP cost: on the first answer after a CED prefill, MTP accepts about 10% fewer tokens per step, so that answer decodes about 10% slower. A follow-up turn is back to normal.
  • VRAM: the projector takes 1.76 GiB per GPU (bf16). Having it loaded costs no decode speed.
  • Not approximated: prompts with images, prompts shorter than 8,192 tokens, and requests for prompt logprobs run exactly.

To turn CED off for the whole server, run ./r9v setup or ./r9v start with --ced off. The projector is then not loaded. To keep one request exact, send "vllm_xargs": {"r9v_ced": false} (OpenAI Python client: extra_body={"vllm_xargs": {"r9v_ced": False}}).

CED quality (opt-in, R9V v0.4.2 or later)

./r9v setup qwen38-mtp4-uncensored ... --ced quality downloads one more file, ced/ced-projector-split16-msfa-int8.safetensors (1.8 GB), and serves with it. It is a multi-source split-16 projector that also reads the inputs of full-attention layers 3, 7, 11 and 15. In one GPU grade of both projectors (both loaded as int8) on the 16 prompts that depend on their long context, it gave ×1.029 perplexity instead of ×1.049 (10% of the long-context gain lost instead of 17%), for 1.55× instead of 1.68× prefill. It takes 1.79 GiB of VRAM per GPU, stored as int8. A default setup never downloads it.

The CED projector was fitted on English-heavy code, docs and prose of up to ~20K tokens. See the CED projector repository for the method, the full grade and its limitations.

Status

The immutable 25-file model package is public and remotely hash-verified at revision 8112610745a8ddc3a19cc659314af245820ee728. The R9V runtime for this profile remains experimental until the documented package installation passes from a clean host.

License

These model artifacts are distributed under Qwen Community License 1.0; see LICENSE. orcarouter's model card labels the abliterated weights Apache-2.0. However, orcarouter's own repository LICENSE file is the Qwen Community License, and the weights are a derivative of Qwen3.8 Flash Next, so this package follows the upstream Qwen license. The R9V Apache-2.0 code license does not apply to model weights. Users are responsible for reviewing the Qwen license, including its separate terms for certain commercial MaaS/AI-work-assistant uses and scale thresholds. Artifact attribution and exact upstream revisions are recorded in THIRD_PARTY_NOTICES.md.

abliterated
conversational
endpoints_compatible
gguf
qwen3.8
rocm
safetensors
speculative-decoding
uncensored
vision