DeepSeek V4 Flash 0731 GGUF for one GX10
2
2 commits
2 linked in READMEs
updated Jul 31, 2026
This is a community GGUF quantization of deepseek-ai/DeepSeek-V4-Flash-0731, built to run fully resident on one 128 GB NVIDIA GB10 system such as an ASUS Ascent GX10 or DGX Spark.
It is not an official DeepSeek release.
| Property | Value |
|---|---|
| File | DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix.gguf |
| Size | 86,720,111,552 bytes, 80.76 GiB |
| SHA-256 | 73b86c93e7b0be37f68c4b3a5361c7e17f9b278647aa335ea304ac5dfd784867 |
| Source revision | 9e165c30e2704aec5d9d593cce3eebd58bbef1cb |
| Tested runtime | antirez/ds4 at 54b36ed9ba42da31b24f2d1a5feb075c2475dbb1 |
The GGUF contains the 45-shard base target model. It does not include shards 46 to 48 or enable the optional DSpark speculative module.
The layout follows the proven one-GB10 V4 Flash recipe:
IQ2_XXSQ2_KQ8_0Q8_0Q8_0F16F16 or F32 according to the pinned templateThe calibration imatrix is Antirez's 1.5M-token routed-MoE dataset artifact from antirez/deepseek-v4-gguf, revision a88c423b511666d7ff7a4dcaee651669312bea97.
A header-only GGUF audit reconciled every serialized tensor byte without reading tensor payloads:
| Tensor family or type | Stored payload |
|---|---|
| Routed experts | 72.562 GiB, 89.845% of the file |
IQ2_XXS tensors | 44.344 GiB |
Q2_K tensors | 28.219 GiB |
Q8_0 tensors | 6.146 GiB |
F16 tensors | 2.041 GiB |
| Attention projections | 5.458 GiB |
| Shared experts | 1.071 GiB |
The model routes 6 of 256 experts per layer. This does not imply that only 6/256 of stored expert bytes are transferred by the runtime, but it explains why selective low-bit expert storage matters more than uniformly reducing every tensor. See measurement/header-byte-accounting.json.
Important limitation: the imatrix was collected for the earlier V4 Flash checkpoint and reused for 0731. This release has not established broad quality equivalence to the official mixed-precision checkpoint.
Measured on one NVIDIA GB10 with 128 GB unified memory using target-only decoding. Both checkpoints used the same DS4 binary, prompts, tokenizer/template path, context, sampler settings, seed, power setting, and GPU budget.
Protocol:
260729ABBAAB measured order| Workload | Earlier V4 Flash | V4 Flash 0731 | Change |
|---|---|---|---|
| Code generation | 16.86 tok/s | 16.91 tok/s | +0.30% |
| Technical prose | 16.83 tok/s | 16.84 tok/s | +0.06% |
Across all 16 warm-up and measured runs:
The output text differed between the two checkpoints, as expected after post-training. The benchmark prompts were capped at 256 generated tokens and therefore do not support a broad quality-retention claim. See benchmark/protocol.json and benchmark/summary.json.
The matched benchmark ran before a same-length provenance-only metadata sanitization. That edit replaced the local quantize.imatrix.file string without changing tensor payload bytes, tensor offsets, or file size. The uploaded artifact was rehashed, header-audited, and smoke-tested afterward.
The final GGUF loaded through the CUDA backend on an NVIDIA GB10 and generated:
This sanitized local model response was generated on an NVIDIA GX10.
For that post-sanitization one-shot smoke test:
The matched three-run benchmark above is the better throughput reference.
hf download tekosML/DeepSeek-V4-Flash-0731-GGUF-GX10 \
DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix.gguf \
--local-dir .
Build DS4 at the tested commit using its documented CUDA Spark target, then run:
./ds4 \
-m DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix.gguf \
--prompt-file prompt.txt \
--ctx 1024 \
-n 256 \
--temp 0 \
--seed 260729 \
--nothink \
--power 100 \
--gpu-vram 105
Start conservatively. Larger context windows require more KV memory and were not validated for this release.
See:
provenance.json for revisions, sizes, and hashesREADY.json for the exercised smoke-test resultbenchmark/protocol.json and benchmark/summary.json for the matched GX10 benchmarkmeasurement/header-byte-accounting.json for exact GGUF payload and tensor-family accountingconversion/resume-notes.md for the boundary-validated write-resume provenance after the original converter stopped at tensor 133The resume helper changed file continuation behavior only. It validated the complete GGUF header, every planned tensor's name, shape, type, size, and offset, and required the existing file to end at an exact tensor boundary before appending the remaining tensors. The final artifact was independently hashed after completion.
The source checkpoint is MIT licensed. The original license is included in this repository. Preserve the license and attribution when redistributing this derivative.
DeepSeek V4 Flash 0731 GGUF for one GX10
2
2 commits
2 linked in READMEs
updated Jul 31, 2026
This is a community GGUF quantization of deepseek-ai/DeepSeek-V4-Flash-0731, built to run fully resident on one 128 GB NVIDIA GB10 system such as an ASUS Ascent GX10 or DGX Spark.
It is not an official DeepSeek release.
| Property | Value |
|---|---|
| File | DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix.gguf |
| Size | 86,720,111,552 bytes, 80.76 GiB |
| SHA-256 | 73b86c93e7b0be37f68c4b3a5361c7e17f9b278647aa335ea304ac5dfd784867 |
| Source revision | 9e165c30e2704aec5d9d593cce3eebd58bbef1cb |
| Tested runtime | antirez/ds4 at 54b36ed9ba42da31b24f2d1a5feb075c2475dbb1 |
The GGUF contains the 45-shard base target model. It does not include shards 46 to 48 or enable the optional DSpark speculative module.
The layout follows the proven one-GB10 V4 Flash recipe:
IQ2_XXSQ2_KQ8_0Q8_0Q8_0F16F16 or F32 according to the pinned templateThe calibration imatrix is Antirez's 1.5M-token routed-MoE dataset artifact from antirez/deepseek-v4-gguf, revision a88c423b511666d7ff7a4dcaee651669312bea97.
A header-only GGUF audit reconciled every serialized tensor byte without reading tensor payloads:
| Tensor family or type | Stored payload |
|---|---|
| Routed experts | 72.562 GiB, 89.845% of the file |
IQ2_XXS tensors | 44.344 GiB |
Q2_K tensors | 28.219 GiB |
Q8_0 tensors | 6.146 GiB |
F16 tensors | 2.041 GiB |
| Attention projections | 5.458 GiB |
| Shared experts | 1.071 GiB |
The model routes 6 of 256 experts per layer. This does not imply that only 6/256 of stored expert bytes are transferred by the runtime, but it explains why selective low-bit expert storage matters more than uniformly reducing every tensor. See measurement/header-byte-accounting.json.
Important limitation: the imatrix was collected for the earlier V4 Flash checkpoint and reused for 0731. This release has not established broad quality equivalence to the official mixed-precision checkpoint.
Measured on one NVIDIA GB10 with 128 GB unified memory using target-only decoding. Both checkpoints used the same DS4 binary, prompts, tokenizer/template path, context, sampler settings, seed, power setting, and GPU budget.
Protocol:
260729ABBAAB measured order| Workload | Earlier V4 Flash | V4 Flash 0731 | Change |
|---|---|---|---|
| Code generation | 16.86 tok/s | 16.91 tok/s | +0.30% |
| Technical prose | 16.83 tok/s | 16.84 tok/s | +0.06% |
Across all 16 warm-up and measured runs:
The output text differed between the two checkpoints, as expected after post-training. The benchmark prompts were capped at 256 generated tokens and therefore do not support a broad quality-retention claim. See benchmark/protocol.json and benchmark/summary.json.
The matched benchmark ran before a same-length provenance-only metadata sanitization. That edit replaced the local quantize.imatrix.file string without changing tensor payload bytes, tensor offsets, or file size. The uploaded artifact was rehashed, header-audited, and smoke-tested afterward.
The final GGUF loaded through the CUDA backend on an NVIDIA GB10 and generated:
This sanitized local model response was generated on an NVIDIA GX10.
For that post-sanitization one-shot smoke test:
The matched three-run benchmark above is the better throughput reference.
hf download tekosML/DeepSeek-V4-Flash-0731-GGUF-GX10 \
DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix.gguf \
--local-dir .
Build DS4 at the tested commit using its documented CUDA Spark target, then run:
./ds4 \
-m DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix.gguf \
--prompt-file prompt.txt \
--ctx 1024 \
-n 256 \
--temp 0 \
--seed 260729 \
--nothink \
--power 100 \
--gpu-vram 105
Start conservatively. Larger context windows require more KV memory and were not validated for this release.
See:
provenance.json for revisions, sizes, and hashesREADY.json for the exercised smoke-test resultbenchmark/protocol.json and benchmark/summary.json for the matched GX10 benchmarkmeasurement/header-byte-accounting.json for exact GGUF payload and tensor-family accountingconversion/resume-notes.md for the boundary-validated write-resume provenance after the original converter stopped at tensor 133The resume helper changed file continuation behavior only. It validated the complete GGUF header, every planned tensor's name, shape, type, size, and offset, and required the existing file to end at an exact tensor boundary before appending the remaining tensors. The final artifact was independently hashed after completion.
The source checkpoint is MIT licensed. The original license is included in this repository. Preserve the license and attribution when redistributing this derivative.