[!IMPORTANT] Custom Engine Required β Incompatible with Stock Ollama / Vanilla llama.cpp This repository provides custom ROCmFP4 quantized weights (
Q4_0_ROCMFP4_STRIX_LEANusing custom GGML tensor types 100 & 101, file type 106) engineered specifically for AMD Strix Halo (gfx1151) and RDNA 3.5 architectures.
- Engine Requirement: Requires ROCmFPX or halofpx to run.
- Stock Ollama / llama.cpp Incompatibility: Stock
llama.cppand vanillaollamawill fail to load these weights (unknown tensor type 101and unsupportedqwen35moeGated DeltaNet architecture).- Standard Quants: If you need standard vanilla GGUF quants (Q4_K_M, etc.) for general llama.cpp usage, please use abenzerps/Nex-N2.5-mini-GGUF.
ROCmFP4 (Q4_0_ROCMFP4_STRIX_LEAN) quantization of Nex-N2.5-mini for AMD Strix Halo (gfx1151) and RDNA 3.5 GPUs, engineered using ROCmFPX.
Nex-N2.5-mini is an open-source agentic multimodal MoE model built by Nex AGI on the Qwen3.5-35B-A3B architecture, unifying requirement understanding, code generation, tool use, and environment execution through an Agentic Thinking adaptive reasoning loop.
| Property | Value |
|---|---|
| Quant format | Q4_0_ROCMFP4_STRIX_LEAN (ROCmFP4) |
| Bits per weight | 4.29 BPW |
| File size | 17.32 GiB |
| SHA256 | 406c96dbab1994998137e5cf093c4094f9af8be5c1e3268ed6284670bca2d06e |
| Vision projector | mmproj-Nex-N2.5-mini.gguf (0.84 GiB) |
| Projector SHA256 | 4734f7323dfc0e8dcd5c7c408991aad438021aeab761d223aec0b38a4457ca83 |
| Architecture | qwen35moe (30Γ Gated DeltaNet + 10Γ Full Attention) |
| Parameters | 34.66B total / ~3.0B active per token |
| Max Context | 262,144 tokens (256K) |
| Source | abenzerps/Nex-N2.5-mini-GGUF Q4_K_M (allow-requantize) |
| Notes | Expert weights in q4_0_rocmfp4_fast, attention K/V in q4_0_rocmfp4, FP32 router/norms, Q5_K embeddings |
| Configuration | Prefill (pp512) | Decode (tg128) | Size | Speedup vs Q4_K_M |
|---|---|---|---|---|
| ROCmFP4 Vulkan0 (RADV) | 642.37 tok/s | π₯ 76.92 tok/s | 17.32 GiB | +5.7% decode, β12.1% size |
| ROCmFP4 ROCm0 (HIP) | 1,028.18 tok/s | 68.62 tok/s | 17.32 GiB | +13.6% decode, β12.1% size |
| Q4_K_M Baseline (Vulkan0) | 1,083.91 tok/s | 72.76 tok/s | 19.71 GiB | Baseline |
| Q4_K_M Baseline (ROCm0) | 907.44 tok/s | 60.39 tok/s | 19.71 GiB | Baseline |
Standard GGUF integer quants (such as Q4_K_M) were designed primarily for CPU cache architectures and CUDA tensor cores. On AMD Strix Halo APUs (gfx1151) and RDNA 3.5 architectures, ROCmFP4_STRIX_LEAN achieves both higher decode throughput and smaller footprint for four key architectural reasons:
KHR_coopmat / Mesa RADV Wave64)Q4_K blocks are non-uniform: 256-element blocks split into 8 sub-blocks of 32 elements with dual 6-bit scales and 6-bit offsets. Compute units must pay a complex, multi-pass unpack and ALU dequantization penalty in vector registers before data can feed matrix multiply units.ROCmFP4 formats (Q4_0_ROCMFP4 and Q4_0_ROCMFP4_FAST) use single-scale uniform FP4 quantization per 32 elements. In shader registers, unpacking is reduced to single-cycle bit shifts and direct table lookups. This dramatically reduces register pressure and instruction count, allowing Mesa RADV's Wave64 cooperative matrix pipelines to run near theoretical hardware saturation.Rather than naively crushing all tensors to 4-bit, the STRIX_LEAN recipe selectively preserves precision where accuracy matters most:
ffn_gate_inp.weight) & LayerNorms: Maintained in uncompressed FP32. Expert routing decisions and token assignments remain bit-exact, preventing expert collapse.token_embd.weight): Quantized in higher-precision Q5_K to maintain vocabulary entropy and prevent prompt degradation.q4_0_rocmfp4 for clean KV heads.q4_0_rocmfp4_fast for maximum memory streaming bandwidth.halofpx pull downloads and verifies both the ROCmFP4 weights and vision projector:
halofpx pull nex-n2.5-mini
halofpx serve -m nex-n2.5-mini
llama-server -m Nex-N2.5-mini-ROCmFP4-STRIX_LEAN.gguf --mmproj mmproj-Nex-N2.5-mini.gguf -ngl 99 -c 32768 -fa on --host 0.0.0.0 --port 8080
Note on Reasoning Mode: To enable thinking mode over the
/v1/chat/completionsAPI, request with--reasoning-format deepseekor passchat_template_kwargs: {"enable_thinking": true}. To disable thinking for lower latency, passchat_template_kwargs: {"enable_thinking": false}.
[!IMPORTANT] Custom Engine Required β Incompatible with Stock Ollama / Vanilla llama.cpp This repository provides custom ROCmFP4 quantized weights (
Q4_0_ROCMFP4_STRIX_LEANusing custom GGML tensor types 100 & 101, file type 106) engineered specifically for AMD Strix Halo (gfx1151) and RDNA 3.5 architectures.
- Engine Requirement: Requires ROCmFPX or halofpx to run.
- Stock Ollama / llama.cpp Incompatibility: Stock
llama.cppand vanillaollamawill fail to load these weights (unknown tensor type 101and unsupportedqwen35moeGated DeltaNet architecture).- Standard Quants: If you need standard vanilla GGUF quants (Q4_K_M, etc.) for general llama.cpp usage, please use abenzerps/Nex-N2.5-mini-GGUF.
ROCmFP4 (Q4_0_ROCMFP4_STRIX_LEAN) quantization of Nex-N2.5-mini for AMD Strix Halo (gfx1151) and RDNA 3.5 GPUs, engineered using ROCmFPX.
Nex-N2.5-mini is an open-source agentic multimodal MoE model built by Nex AGI on the Qwen3.5-35B-A3B architecture, unifying requirement understanding, code generation, tool use, and environment execution through an Agentic Thinking adaptive reasoning loop.
| Property | Value |
|---|---|
| Quant format | Q4_0_ROCMFP4_STRIX_LEAN (ROCmFP4) |
| Bits per weight | 4.29 BPW |
| File size | 17.32 GiB |
| SHA256 | 406c96dbab1994998137e5cf093c4094f9af8be5c1e3268ed6284670bca2d06e |
| Vision projector | mmproj-Nex-N2.5-mini.gguf (0.84 GiB) |
| Projector SHA256 | 4734f7323dfc0e8dcd5c7c408991aad438021aeab761d223aec0b38a4457ca83 |
| Architecture | qwen35moe (30Γ Gated DeltaNet + 10Γ Full Attention) |
| Parameters | 34.66B total / ~3.0B active per token |
| Max Context | 262,144 tokens (256K) |
| Source | abenzerps/Nex-N2.5-mini-GGUF Q4_K_M (allow-requantize) |
| Notes | Expert weights in q4_0_rocmfp4_fast, attention K/V in q4_0_rocmfp4, FP32 router/norms, Q5_K embeddings |
| Configuration | Prefill (pp512) | Decode (tg128) | Size | Speedup vs Q4_K_M |
|---|---|---|---|---|
| ROCmFP4 Vulkan0 (RADV) | 642.37 tok/s | π₯ 76.92 tok/s | 17.32 GiB | +5.7% decode, β12.1% size |
| ROCmFP4 ROCm0 (HIP) | 1,028.18 tok/s | 68.62 tok/s | 17.32 GiB | +13.6% decode, β12.1% size |
| Q4_K_M Baseline (Vulkan0) | 1,083.91 tok/s | 72.76 tok/s | 19.71 GiB | Baseline |
| Q4_K_M Baseline (ROCm0) | 907.44 tok/s | 60.39 tok/s | 19.71 GiB | Baseline |
Standard GGUF integer quants (such as Q4_K_M) were designed primarily for CPU cache architectures and CUDA tensor cores. On AMD Strix Halo APUs (gfx1151) and RDNA 3.5 architectures, ROCmFP4_STRIX_LEAN achieves both higher decode throughput and smaller footprint for four key architectural reasons:
KHR_coopmat / Mesa RADV Wave64)Q4_K blocks are non-uniform: 256-element blocks split into 8 sub-blocks of 32 elements with dual 6-bit scales and 6-bit offsets. Compute units must pay a complex, multi-pass unpack and ALU dequantization penalty in vector registers before data can feed matrix multiply units.ROCmFP4 formats (Q4_0_ROCMFP4 and Q4_0_ROCMFP4_FAST) use single-scale uniform FP4 quantization per 32 elements. In shader registers, unpacking is reduced to single-cycle bit shifts and direct table lookups. This dramatically reduces register pressure and instruction count, allowing Mesa RADV's Wave64 cooperative matrix pipelines to run near theoretical hardware saturation.Rather than naively crushing all tensors to 4-bit, the STRIX_LEAN recipe selectively preserves precision where accuracy matters most:
ffn_gate_inp.weight) & LayerNorms: Maintained in uncompressed FP32. Expert routing decisions and token assignments remain bit-exact, preventing expert collapse.token_embd.weight): Quantized in higher-precision Q5_K to maintain vocabulary entropy and prevent prompt degradation.q4_0_rocmfp4 for clean KV heads.q4_0_rocmfp4_fast for maximum memory streaming bandwidth.halofpx pull downloads and verifies both the ROCmFP4 weights and vision projector:
halofpx pull nex-n2.5-mini
halofpx serve -m nex-n2.5-mini
llama-server -m Nex-N2.5-mini-ROCmFP4-STRIX_LEAN.gguf --mmproj mmproj-Nex-N2.5-mini.gguf -ngl 99 -c 32768 -fa on --host 0.0.0.0 --port 8080
Note on Reasoning Mode: To enable thinking mode over the
/v1/chat/completionsAPI, request with--reasoning-format deepseekor passchat_template_kwargs: {"enable_thinking": true}. To disable thinking for lower latency, passchat_template_kwargs: {"enable_thinking": false}.