SemiAnalysisAI/InferenceX

Open Source Inference Research Platform Standard / 开源推理研究平台

Python

1,782

2,473 commits

updated Sep 29, 2026

See the code

README

InferenceX™, Open Source Inference Research Platform / 开源推理研究平台

License PRs Welcome Dashboard Ask DeepWiki GitHub Stars

English | 中文

Trusted by Operators of Trillion Dollar Token Factories such as OpenAI, Meta, Microsoft, Oracle, etc, & ML Community such as PyTorch Foundation, vLLM, SGLang, Tri Dao

Projects

DirectoryContents
InferenceX-e2e/🚀 End-to-end Inference Serving Benchmarks
CollectiveX/🌐 Networking & Collective Communication Benchmarks (Experimental Beta)
OperatorX/⚙️ Operator & Kernel Level Benchmarks (Experimental Beta)
power_model/⚡ OSS System Level Power Modelling
shared/🧩 Home for shared components
experimental/🧪 Remaining experiments

News

  • [2026/09] DeepSeek V4.1 Flash: added AgentX benchmarks dashboard
  • [2026/08] Qwen3.8-Flash-Next: added AgentX benchmarks with native multi-token prediction (MTP) dashboard
  • [2026/08] 🔥 AgentX: World's First Fully Open Source Apache 2.0 Realistic 1Mil+ Long Context, Multi Turn Benchmark Live dashboard
  • [2026/08] 🔥 GLM5.3: continuous agentic benchmarks live too dashboard
  • [2026/07] 🔥 Kimi K3 2.8T: continuous benchmarks live since Day 0
  • [2026/06] 🔥 MiniMax M3: continuous benchmarks live since Day 0 dashboard
  • [2026/04] 🔥 DeepSeek V4 Pro 1.6T: continuous benchmarks live since Day 0 article, dashboard
  • [2026/03] 🔥 Qwen3.5 397B: continuous benchmarks live since Day 0 dashboard
  • [2026/03] Added Kimi K2.5 (same architecture as Kimi 2.7-Code), GLM5 (same arch as GLM5.1), and MiniMax M2.5 (same arch as MiniMax M2.7) dashboard
  • [2026/02] GB300 NVL72: added to InferenceX & continuously benchmarked SGLang Maintainer Lmsys Blog
  • [2026/02] 🔥 InferenceX v2 launch comparing NVIDIA Blackwell, AMD, and Hopper article
  • [2025/10] 🔥 InferenceX (formerly InferenceMAX) v1 launch article

Introduction

InferenceX™ (formerly InferenceMAX) is an inference performance research platform dedicated to continually analyzing & benchmarking the world’s most popular open-source inference frameworks used by major token factories and models to track real performance in real time. As these software stacks improve, InferenceX™ captures that progress in near real-time, providing a live indicator of inference performance progress. A open sourced live dashboard is available for free publicly at https://inferencex.com/.

[!IMPORTANT] Only SemiAnalysisAI/InferenceX repo contains the Official InferenceX™ result, all other forks & repos are Unofficial. The benchmark setup & quality of machines/clouds in unofficial repos may be differ leading to subpar benchmarking. Unofficial must be explicitly labelled as Unofficial. Forks may not remove this disclaimer

InferenceX DeepSeekv4 MXFP4 Performance Curve

Why?

InferenceX™, an open-source, under Apache2 license, automated benchmark designed to move at the same rapid speed as the software ecosystem itself, is built to address this challenge.

LLM Inference performance is driven by two pillars, hardware and software. While hardware innovation drives step jumps in performance every year through the release of new GPUs/XPUs and new systems, software evolves every single day, delivering continuous performance gains on top of these step jumps. Speed is the Moat 🚀

AI software like SGLang, vLLM, TensorRT-LLM, CUDA, ROCm and achieve this continuous improvement in performance through kernel-level optimizations, distributed inference strategies, and scheduling innovations that increase the pareto frontier of performance in incremental releases that can be just days apart.

This pace of software advancement creates a challenge: benchmarks conducted at a fixed point in time quickly go stale and do not represent the performance that can be achieved with the latest software packages.

Officially Supported Hardware

SKUStatus
Vera Rubin NVL72✅
GB300 NVL72✅
GB200 NVL72✅
MI355X✅
B300✅
B200✅
MI325X✅
MI300X✅
H200✅
H100✅
TPUv7x Ironwood Ghostfish✅
RTX PRO 6000 Server✅
MI455 UALoE72Coming Soon 🔜
Rubin NVL8Coming Soon 🔜
Chip #1 from Hardware Vendor #1Coming Soon 🔜
Chip #2 from Hardware Vendor #1Coming Soon 🔜
Chip #1 from Hardware Vendor #2Coming Soon 🔜
Chip #1 from Hardware Vendor #3Coming Soon 🔜
Chip #1 from Hardware Vendor #4Coming Soon 🔜

Contributing

PRs are welcome! See CONTRIBUTING.md for more details on the PR review flow, the PR Review Checklist, and the merge process. For the maintainer and agent documentation map, start with docs/index.md. It links the architecture, configuration, workflow, eval, runner, and troubleshooting references.

To benchmark an existing server without CI or Slurm, see Run AgentX-Harness Standalone for client installation and a direct aiperf profile command.

Acknowledgements & Supporters

Thank you to Lisa Su and Anush Elangovan for providing the MI355X and CDNA3 GPUs for this free and open-source project. We want to recognize the many AMD contributors for their responsiveness and for debugging, optimizing, and validating performance across AMD GPUs. We’re also grateful to Jensen Huang and Ian Buck for supporting this open source with access to a GB200 NVL72 rack (through OCI) and B200 GPUs. Thank you to the many NVIDIA contributors from the NVIDIA inference team, NVIDIA Dynamo team.

We also want to recognize the SGLang, vLLM, and TensorRT-LLM maintainers for building a world-class software stack and open sourcing it to the entire world. Finally, we’re grateful to Crusoe, CoreWeave, Nebius, TensorWave, Oracle and TogetherAI for supporting open-source innovation through compute resources, enabling this.

Full list of supporters & quotes: https://inferencex.semianalysis.com/quotes

image
agentic
agentic-ai
amd
benchmarking
benchmarks
cuda
deepseek
gpu
inference
kimi
llm-inference
llm-serving
mlops
nvidia
pytorch
qwen
rocm
sglang
tensorrt-llm
vllm

Significant stargazers

(top 24 of 34)

Paco Xu

818 followers · starred May 2026

CYJiang

234 followers · starred Apr 2026

samsja

332 followers · starred Oct 2025

samzong

204 followers · starred Jul 2026

SemiAnalysisAI/InferenceX

Open Source Inference Research Platform Standard / 开源推理研究平台

Python

1,782

2,473 commits

updated Sep 29, 2026

See the code

README

InferenceX™, Open Source Inference Research Platform / 开源推理研究平台

License PRs Welcome Dashboard Ask DeepWiki GitHub Stars

English | 中文

Trusted by Operators of Trillion Dollar Token Factories such as OpenAI, Meta, Microsoft, Oracle, etc, & ML Community such as PyTorch Foundation, vLLM, SGLang, Tri Dao

Projects

DirectoryContents
InferenceX-e2e/🚀 End-to-end Inference Serving Benchmarks
CollectiveX/🌐 Networking & Collective Communication Benchmarks (Experimental Beta)
OperatorX/⚙️ Operator & Kernel Level Benchmarks (Experimental Beta)
power_model/⚡ OSS System Level Power Modelling
shared/🧩 Home for shared components
experimental/🧪 Remaining experiments

News

  • [2026/09] DeepSeek V4.1 Flash: added AgentX benchmarks dashboard
  • [2026/08] Qwen3.8-Flash-Next: added AgentX benchmarks with native multi-token prediction (MTP) dashboard
  • [2026/08] 🔥 AgentX: World's First Fully Open Source Apache 2.0 Realistic 1Mil+ Long Context, Multi Turn Benchmark Live dashboard
  • [2026/08] 🔥 GLM5.3: continuous agentic benchmarks live too dashboard
  • [2026/07] 🔥 Kimi K3 2.8T: continuous benchmarks live since Day 0
  • [2026/06] 🔥 MiniMax M3: continuous benchmarks live since Day 0 dashboard
  • [2026/04] 🔥 DeepSeek V4 Pro 1.6T: continuous benchmarks live since Day 0 article, dashboard
  • [2026/03] 🔥 Qwen3.5 397B: continuous benchmarks live since Day 0 dashboard
  • [2026/03] Added Kimi K2.5 (same architecture as Kimi 2.7-Code), GLM5 (same arch as GLM5.1), and MiniMax M2.5 (same arch as MiniMax M2.7) dashboard
  • [2026/02] GB300 NVL72: added to InferenceX & continuously benchmarked SGLang Maintainer Lmsys Blog
  • [2026/02] 🔥 InferenceX v2 launch comparing NVIDIA Blackwell, AMD, and Hopper article
  • [2025/10] 🔥 InferenceX (formerly InferenceMAX) v1 launch article

Introduction

InferenceX™ (formerly InferenceMAX) is an inference performance research platform dedicated to continually analyzing & benchmarking the world’s most popular open-source inference frameworks used by major token factories and models to track real performance in real time. As these software stacks improve, InferenceX™ captures that progress in near real-time, providing a live indicator of inference performance progress. A open sourced live dashboard is available for free publicly at https://inferencex.com/.

[!IMPORTANT] Only SemiAnalysisAI/InferenceX repo contains the Official InferenceX™ result, all other forks & repos are Unofficial. The benchmark setup & quality of machines/clouds in unofficial repos may be differ leading to subpar benchmarking. Unofficial must be explicitly labelled as Unofficial. Forks may not remove this disclaimer

InferenceX DeepSeekv4 MXFP4 Performance Curve

Why?

InferenceX™, an open-source, under Apache2 license, automated benchmark designed to move at the same rapid speed as the software ecosystem itself, is built to address this challenge.

LLM Inference performance is driven by two pillars, hardware and software. While hardware innovation drives step jumps in performance every year through the release of new GPUs/XPUs and new systems, software evolves every single day, delivering continuous performance gains on top of these step jumps. Speed is the Moat 🚀

AI software like SGLang, vLLM, TensorRT-LLM, CUDA, ROCm and achieve this continuous improvement in performance through kernel-level optimizations, distributed inference strategies, and scheduling innovations that increase the pareto frontier of performance in incremental releases that can be just days apart.

This pace of software advancement creates a challenge: benchmarks conducted at a fixed point in time quickly go stale and do not represent the performance that can be achieved with the latest software packages.

Officially Supported Hardware

SKUStatus
Vera Rubin NVL72✅
GB300 NVL72✅
GB200 NVL72✅
MI355X✅
B300✅
B200✅
MI325X✅
MI300X✅
H200✅
H100✅
TPUv7x Ironwood Ghostfish✅
RTX PRO 6000 Server✅
MI455 UALoE72Coming Soon 🔜
Rubin NVL8Coming Soon 🔜
Chip #1 from Hardware Vendor #1Coming Soon 🔜
Chip #2 from Hardware Vendor #1Coming Soon 🔜
Chip #1 from Hardware Vendor #2Coming Soon 🔜
Chip #1 from Hardware Vendor #3Coming Soon 🔜
Chip #1 from Hardware Vendor #4Coming Soon 🔜

Contributing

PRs are welcome! See CONTRIBUTING.md for more details on the PR review flow, the PR Review Checklist, and the merge process. For the maintainer and agent documentation map, start with docs/index.md. It links the architecture, configuration, workflow, eval, runner, and troubleshooting references.

To benchmark an existing server without CI or Slurm, see Run AgentX-Harness Standalone for client installation and a direct aiperf profile command.

Acknowledgements & Supporters

Thank you to Lisa Su and Anush Elangovan for providing the MI355X and CDNA3 GPUs for this free and open-source project. We want to recognize the many AMD contributors for their responsiveness and for debugging, optimizing, and validating performance across AMD GPUs. We’re also grateful to Jensen Huang and Ian Buck for supporting this open source with access to a GB200 NVL72 rack (through OCI) and B200 GPUs. Thank you to the many NVIDIA contributors from the NVIDIA inference team, NVIDIA Dynamo team.

We also want to recognize the SGLang, vLLM, and TensorRT-LLM maintainers for building a world-class software stack and open sourcing it to the entire world. Finally, we’re grateful to Crusoe, CoreWeave, Nebius, TensorWave, Oracle and TogetherAI for supporting open-source innovation through compute resources, enabling this.

Full list of supporters & quotes: https://inferencex.semianalysis.com/quotes

image
agentic
agentic-ai
amd
benchmarking
benchmarks
cuda
deepseek
gpu
inference
kimi
llm-inference
llm-serving
mlops
nvidia
pytorch
qwen
rocm
sglang
tensorrt-llm
vllm

Significant stargazers

(top 24 of 34)

Paco Xu

818 followers · starred May 2026

CYJiang

234 followers · starred Apr 2026

samsja

332 followers · starred Oct 2025

samzong

204 followers · starred Jul 2026

Languages

Python

78.1%

Shell

21.4%