local-inference-lab/rtx6kpro

RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

1,024

stars

536

commits

Python

primary language

Sep 9, 2026

updated

README

RTX PRO 6000 Blackwell LLM Wiki

This repository is a field wiki for running frontier LLMs on NVIDIA RTX PRO 6000 Blackwell / SM120 PCIe systems. It is more than a few launch snippets: it contains reproducible Docker builds, exact vLLM and SGLang runbooks, benchmark tables, KLD quality checks, quantization notes, DCP/MTP/DSpark/DFlash debugging, PCIe topology work, and regression history.

Community workbench for RTX PRO 6000 / SM120 serving: https://discord.gg/X54jjmcxWJ

Start Here

If you just want to run a model, use these stable hub pages first:

Model familyStart hereScope
GLM-5.3-FlashGLM-5.3-FlashJovian Judgement Community source-locked NVIDIA 4-bit floating-point (NVFP4) target with no-spec, MTP:3, and Microscaling 8-bit floating-point (MXFP8) DFlash2 modes, qualified TP4/DCP1 performance, AA-LCR and KLD evaluation, optional DCP4 full-CKV prefill, and Docker launch.
GLM-5.2GLM-5.2 Runbook HubFathomless vLLM, NVFP4, online FP8/MXFP8, B12X, DCP, MTP, KLD.
DeepSeek-V4-Flash / DSparkDeepSeek-V4-Flash Runbook HubText and Vision checkpoints, DSpark, B12X, LMCache, Lucifer, and CUTLASS.
KimiKimi Runbook HubKimi-K2.7-Code, DFlash, parser/tool-call runtime.
Xiaomi MiMoMiMo Runbook HubMiMo V2.5 Pro FP4-DFlash.
Qwen3.8-27BQwen3.8-27B on RTX PRO 6000 Blackwell, readable QSRT K5 training result, exact QSRT K5 specificationTP1, TP2, and TP4 throughput evidence plus the QSRT K5 training interpretation, artifact, fidelity, runtime, and source contract.
GLM-5.1GLM-5.1 Runbook HubHistorical GLM-5.1, KLD methodology, older B12X/SGLang work.
Legacy / secondary modelsLegacy Model RunbooksDeepSeek-V4-Pro, GLM-4.7, Qwen, MiniMax, older Kimi pages.

Need the complete map of every Markdown page?

IndexUse
Full Wiki IndexComplete generated catalog of every page in this repository.
Glossary And Acronym GuideAcronym expansions and writing rules for newcomer-friendly docs.
Newcomer OnboardingHow to ask useful questions without lowering the technical signal.

What Is In This Repository?

NeedWhere
Copy/paste production launch commandsModel hubs and current versioned model pages.
Rebuild the Docker imageEldritch Docker, current model image sections, and build scripts in scripts.
Compare backend speedModel benchmark tables plus Benchmark Results.
Check quantization fidelityGeneral KLD methodology, GLM-5.2 KLD, and model-specific KLD sections.
Understand MTP, DSpark, or DFlashSpeculative Decoding, DS4/Kimi/MiMo pages.
Debug topology or PCIe behaviorTopology, PCIe Bandwidth, GPU Configurations.
Avoid known runtime footgunsCommon Issues, model caveats, and daily summaries.
Understand old measurementsHistorical versioned pages and Daily Summaries.
AreaPageWhy it matters
GLM-5.3-Flash serving stackGLM-5.3-FlashQualified Jovian Judgement Community TP4 image with DCP1 no-spec, MTP:3, and MXFP8 DFlash2 measurements plus DCP4 full-CKV prefill evidence.
GLM-5.2 serving stackGLM-5.2 Infernal Invocation r18Source-qualified CUDA 13.3 profiles with sparse-prefill row validation, projection-mixed EXL3 TP4, online MCG K6, and NVFP4 TP8.
GLM-5.2 MXFP4GLM-5.2 FP8 + MXFP4 ExpertsNative MXFP4 expert checkpoint path and A8 serving notes.
DS4 text and Vision serving profilesDeepSeek-V4-Flash Jovian Judgement r8Source-locked TP2 fixed-K5 text and fixed-K3 Vision serving with measured GPU KV admission, InstantTensor loading, and optional engine-driven LMCache.
DS4 full referenceDS4 DSpark v9Full DSpark and standard MTP sweep reference.
Kimi-K2.7-CodeKimi-K2.7-Code v3Fathomless Kimi DFlash validation.
MiMo FP4-DFlashMiMo FP4-DFlash v3MiMo DFlash validation and fix notes.

Older pages are intentionally preserved. Prefer the hub page for each model family unless you are reproducing a specific old result.

Core Topics

TopicPage
Docker images and release linesDocker Images
PCIe oneshot all-reducePCIe oneshot all-reduce
NCCL tuning and empty graph-file failuresNCCL tuning
Speculative decodingSpeculative decoding
NVFP4 quantizationNVFP4 quantization
Hybrid NVFP4 assemblyHybrid NVFP4 assembly
B12X FP8 / DeepGEMM comparisonB12X dense FP8 GEMM vs DeepGEMM
B12X W4A8 tiny decodeB12X W4A8 MX tiny decode
DSpark upstream consolidationDSpark upstream consolidation
I/O tuningI/O tuning

Benchmarks And Quality

AreaPage
Consolidated throughputBenchmark Results
vLLM vs SGLang throughputInference throughput
GLM-5.2 KLD and quant qualityGLM-5.2 KLD Evaluation
General KLD methodologyMeasuring quantization distribution fidelity in vLLM
MTP quality checksMTP Quality Evaluation
NVFP4 quantization comparisonNVFP4 Quantization Comparison
GLM-5.3-Flash behavioral fidelityVerifier-backed BF16, NVFP4, and QAD comparison

KLD is a regression and quantization-sanity tool, not a complete quality metric. Use it together with long-context decode, coding probes, acceptance-rate checks, and task-level benchmarks.

Hardware And Topology

Most current measurements target RTX PRO 6000 Blackwell / GB202 / SM120 cards: 96 GB GDDR7 per GPU, PCIe 5.0 x16, no NVLink, usually 4-GPU, 8-GPU, or 16-GPU PCIe-switch systems.

AreaPage
SM120 vs SM100SM120 vs SM100 Architecture
PCIe topologyTopology
PCIe bandwidthPCIe Bandwidth
GPU configsGPU Configurations
ASUS ESC8000A-E13PASUS ESC8000A-E13P + Broadcom Switches
ASRockRack Turin 16 GPUASRockRack + EPYC Turin + 4x c-payne
ASRock WRX90 16 GPUASRock WRX90 + 4x c-payne
Power tuningBlackwell power limit sweep

Inference Engines

EnginePageCurrent role
vLLMvLLMPrimary runtime for current GLM-5.2, DS4, Kimi, and MiMo pages.
FlashInferFlashInferSM120 sparse MLA, CUTLASS MoE, sampler, and kernel integration notes.
SGLangSGLangHistorical and alternate runtime notes, especially older GLM/MiMo paths.

Acronym Policy

The wiki uses many acronyms: DCP, MTP, DSpark, DFlash, MLA, MoE, KLD, TP, CC, P2P, NVFP4, MXFP8, and more. To make pages readable for newcomers:

  • Expand important acronyms on first use: Decode Context Parallelism (DCP).
  • Do not expand acronyms inside commands, Docker tags, environment variables, JSON, file paths, or raw logs.
  • Use Glossary And Acronym Guide as the source of truth.
  • Run python3 scripts/check-acronyms.py before polishing a major page.

Keeping The Community Useful

The wiki should reduce accidental gatekeeping: newcomers should be able to find the right runbook, decode acronyms, and reproduce a known-good launch without needing tribal knowledge. That does not mean every Discord question can be answered from memory.

Use Newcomer Onboarding as the support contract: bring the model page, Docker image, full launch command, GPU layout, TP/DCP/MTP or DSpark settings, client command, and logs. That keeps the server welcoming without turning it into an unstructured support queue.

Common Operational Rules

  • Do not launch with NCCL_GRAPH_FILE= set to an empty string. Unset it if no real XML graph file is used.
  • Reuse cache directories while debugging; otherwise TileLang, Triton, CuTe, and FlashInfer rebuilds dominate iteration time.
  • For quick smoke tests, use small MAX_NUM_SEQS and graph caps. For published tables, use the graph sizes documented in the model page.
  • For DFlash and DSpark, confirm backend markers and acceptance rates before trusting throughput numbers.
  • For GLM-5.2, keep the exact index_topk_pattern and DCP policy from the relevant runbook; a truncated pattern can silently degrade output.

Maintaining The Wiki

When adding a page:

  • Link it from the relevant model hub.
  • Add a short status block if it is a current runbook.
  • Keep exact Docker image tags, source commits, model snapshot IDs, GPU layout, backend choices, and benchmark commands.
  • Regenerate the full index:
python3 scripts/generate-wiki-index.py > INDEX.md

For performance claims, include both the server launch config and the client command so results can be reproduced on another PCIe-only Blackwell host.

Maintained from community Discord experiments through July 2026.

Contributors

voipmonitor

526 commits

malaiwah

7 commits

lukealonso

1 commits

local-inference-lab/rtx6kpro

RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

1,024

stars

536

commits

Python

primary language

Sep 9, 2026

updated

README

RTX PRO 6000 Blackwell LLM Wiki

This repository is a field wiki for running frontier LLMs on NVIDIA RTX PRO 6000 Blackwell / SM120 PCIe systems. It is more than a few launch snippets: it contains reproducible Docker builds, exact vLLM and SGLang runbooks, benchmark tables, KLD quality checks, quantization notes, DCP/MTP/DSpark/DFlash debugging, PCIe topology work, and regression history.

Community workbench for RTX PRO 6000 / SM120 serving: https://discord.gg/X54jjmcxWJ

Start Here

If you just want to run a model, use these stable hub pages first:

Model familyStart hereScope
GLM-5.3-FlashGLM-5.3-FlashJovian Judgement Community source-locked NVIDIA 4-bit floating-point (NVFP4) target with no-spec, MTP:3, and Microscaling 8-bit floating-point (MXFP8) DFlash2 modes, qualified TP4/DCP1 performance, AA-LCR and KLD evaluation, optional DCP4 full-CKV prefill, and Docker launch.
GLM-5.2GLM-5.2 Runbook HubFathomless vLLM, NVFP4, online FP8/MXFP8, B12X, DCP, MTP, KLD.
DeepSeek-V4-Flash / DSparkDeepSeek-V4-Flash Runbook HubText and Vision checkpoints, DSpark, B12X, LMCache, Lucifer, and CUTLASS.
KimiKimi Runbook HubKimi-K2.7-Code, DFlash, parser/tool-call runtime.
Xiaomi MiMoMiMo Runbook HubMiMo V2.5 Pro FP4-DFlash.
Qwen3.8-27BQwen3.8-27B on RTX PRO 6000 Blackwell, readable QSRT K5 training result, exact QSRT K5 specificationTP1, TP2, and TP4 throughput evidence plus the QSRT K5 training interpretation, artifact, fidelity, runtime, and source contract.
GLM-5.1GLM-5.1 Runbook HubHistorical GLM-5.1, KLD methodology, older B12X/SGLang work.
Legacy / secondary modelsLegacy Model RunbooksDeepSeek-V4-Pro, GLM-4.7, Qwen, MiniMax, older Kimi pages.

Need the complete map of every Markdown page?

IndexUse
Full Wiki IndexComplete generated catalog of every page in this repository.
Glossary And Acronym GuideAcronym expansions and writing rules for newcomer-friendly docs.
Newcomer OnboardingHow to ask useful questions without lowering the technical signal.

What Is In This Repository?

NeedWhere
Copy/paste production launch commandsModel hubs and current versioned model pages.
Rebuild the Docker imageEldritch Docker, current model image sections, and build scripts in scripts.
Compare backend speedModel benchmark tables plus Benchmark Results.
Check quantization fidelityGeneral KLD methodology, GLM-5.2 KLD, and model-specific KLD sections.
Understand MTP, DSpark, or DFlashSpeculative Decoding, DS4/Kimi/MiMo pages.
Debug topology or PCIe behaviorTopology, PCIe Bandwidth, GPU Configurations.
Avoid known runtime footgunsCommon Issues, model caveats, and daily summaries.
Understand old measurementsHistorical versioned pages and Daily Summaries.
AreaPageWhy it matters
GLM-5.3-Flash serving stackGLM-5.3-FlashQualified Jovian Judgement Community TP4 image with DCP1 no-spec, MTP:3, and MXFP8 DFlash2 measurements plus DCP4 full-CKV prefill evidence.
GLM-5.2 serving stackGLM-5.2 Infernal Invocation r18Source-qualified CUDA 13.3 profiles with sparse-prefill row validation, projection-mixed EXL3 TP4, online MCG K6, and NVFP4 TP8.
GLM-5.2 MXFP4GLM-5.2 FP8 + MXFP4 ExpertsNative MXFP4 expert checkpoint path and A8 serving notes.
DS4 text and Vision serving profilesDeepSeek-V4-Flash Jovian Judgement r8Source-locked TP2 fixed-K5 text and fixed-K3 Vision serving with measured GPU KV admission, InstantTensor loading, and optional engine-driven LMCache.
DS4 full referenceDS4 DSpark v9Full DSpark and standard MTP sweep reference.
Kimi-K2.7-CodeKimi-K2.7-Code v3Fathomless Kimi DFlash validation.
MiMo FP4-DFlashMiMo FP4-DFlash v3MiMo DFlash validation and fix notes.

Older pages are intentionally preserved. Prefer the hub page for each model family unless you are reproducing a specific old result.

Core Topics

TopicPage
Docker images and release linesDocker Images
PCIe oneshot all-reducePCIe oneshot all-reduce
NCCL tuning and empty graph-file failuresNCCL tuning
Speculative decodingSpeculative decoding
NVFP4 quantizationNVFP4 quantization
Hybrid NVFP4 assemblyHybrid NVFP4 assembly
B12X FP8 / DeepGEMM comparisonB12X dense FP8 GEMM vs DeepGEMM
B12X W4A8 tiny decodeB12X W4A8 MX tiny decode
DSpark upstream consolidationDSpark upstream consolidation
I/O tuningI/O tuning

Benchmarks And Quality

AreaPage
Consolidated throughputBenchmark Results
vLLM vs SGLang throughputInference throughput
GLM-5.2 KLD and quant qualityGLM-5.2 KLD Evaluation
General KLD methodologyMeasuring quantization distribution fidelity in vLLM
MTP quality checksMTP Quality Evaluation
NVFP4 quantization comparisonNVFP4 Quantization Comparison
GLM-5.3-Flash behavioral fidelityVerifier-backed BF16, NVFP4, and QAD comparison

KLD is a regression and quantization-sanity tool, not a complete quality metric. Use it together with long-context decode, coding probes, acceptance-rate checks, and task-level benchmarks.

Hardware And Topology

Most current measurements target RTX PRO 6000 Blackwell / GB202 / SM120 cards: 96 GB GDDR7 per GPU, PCIe 5.0 x16, no NVLink, usually 4-GPU, 8-GPU, or 16-GPU PCIe-switch systems.

AreaPage
SM120 vs SM100SM120 vs SM100 Architecture
PCIe topologyTopology
PCIe bandwidthPCIe Bandwidth
GPU configsGPU Configurations
ASUS ESC8000A-E13PASUS ESC8000A-E13P + Broadcom Switches
ASRockRack Turin 16 GPUASRockRack + EPYC Turin + 4x c-payne
ASRock WRX90 16 GPUASRock WRX90 + 4x c-payne
Power tuningBlackwell power limit sweep

Inference Engines

EnginePageCurrent role
vLLMvLLMPrimary runtime for current GLM-5.2, DS4, Kimi, and MiMo pages.
FlashInferFlashInferSM120 sparse MLA, CUTLASS MoE, sampler, and kernel integration notes.
SGLangSGLangHistorical and alternate runtime notes, especially older GLM/MiMo paths.

Acronym Policy

The wiki uses many acronyms: DCP, MTP, DSpark, DFlash, MLA, MoE, KLD, TP, CC, P2P, NVFP4, MXFP8, and more. To make pages readable for newcomers:

  • Expand important acronyms on first use: Decode Context Parallelism (DCP).
  • Do not expand acronyms inside commands, Docker tags, environment variables, JSON, file paths, or raw logs.
  • Use Glossary And Acronym Guide as the source of truth.
  • Run python3 scripts/check-acronyms.py before polishing a major page.

Keeping The Community Useful

The wiki should reduce accidental gatekeeping: newcomers should be able to find the right runbook, decode acronyms, and reproduce a known-good launch without needing tribal knowledge. That does not mean every Discord question can be answered from memory.

Use Newcomer Onboarding as the support contract: bring the model page, Docker image, full launch command, GPU layout, TP/DCP/MTP or DSpark settings, client command, and logs. That keeps the server welcoming without turning it into an unstructured support queue.

Common Operational Rules

  • Do not launch with NCCL_GRAPH_FILE= set to an empty string. Unset it if no real XML graph file is used.
  • Reuse cache directories while debugging; otherwise TileLang, Triton, CuTe, and FlashInfer rebuilds dominate iteration time.
  • For quick smoke tests, use small MAX_NUM_SEQS and graph caps. For published tables, use the graph sizes documented in the model page.
  • For DFlash and DSpark, confirm backend markers and acceptance rates before trusting throughput numbers.
  • For GLM-5.2, keep the exact index_topk_pattern and DCP policy from the relevant runbook; a truncated pattern can silently degrade output.

Maintaining The Wiki

When adding a page:

  • Link it from the relevant model hub.
  • Add a short status block if it is a current runbook.
  • Keep exact Docker image tags, source commits, model snapshot IDs, GPU layout, backend choices, and benchmark commands.
  • Regenerate the full index:
python3 scripts/generate-wiki-index.py > INDEX.md

For performance claims, include both the server launch config and the client command so results can be reproduced on another PCIe-only Blackwell host.

Maintained from community Discord experiments through July 2026.

Contributors

voipmonitor

526 commits

malaiwah

7 commits

lukealonso

1 commits

Languages

Python

71.6%

Shell

26.9%

Dockerfile

1.5%