Run chat, coding, image and video models on AMD Radeon with Paiton's native GPU runtimes, integrated with vLLM, Diffusers and ComfyUI.
Using Qwen3.8? Start with its MXFP4 or 3-bit quickstart.
Download the weights, then pick how to run it:
MXFP4 for the highest accuracy, 3-bit for speed, --mode long for one long document, --mode long-kv4 for coding
agents (prefix caching, about 570K tokens of cache), opt-in --mode long-512k for up to 524K, --vision for images.
Prefer a desktop app? Paiton Studio provides chat, image and video tools for the supported models.
Models · Quick start · Requirements · Your own vLLM · Docs
The container launchers include each model's supported inference environment. Choose a model below and follow its guide. For the smallest language model, MiniCPM5-2B:
git clone --depth 1 https://github.com/Eliovp-BV/paiton-vllm-plugin.git
cd paiton-vllm-plugin
./models/MiniCPM5-2B/serve-docker.sh
The first launch downloads its weights and prepares runtime components; later launches reuse the caches. Its guide includes chat and API examples. Qwen3.8 and Meeting need weight preparation before launch. Run one model at a time on the tested single-GPU setup.
Model links lead to setup, launch options and support limits. Benchmark links include the tested settings and quality tradeoffs; results from different models or profiles are separate comparisons.
These models serve an OpenAI-compatible API through vLLM.
| Model and setup | Use | Benchmark |
|---|---|---|
| MiniCPM5-2B | Lightweight chat, coding and tools · 8K | Results |
| Qwen3.8 27B MXFP4 / 3-bit + DFlash2 | Chat and coding · MXFP4 (most accurate) or 3-bit (fastest) · 65K by default; 262K with --mode long (one long document, fast follow-ups) or --mode long-kv4 (coding agents, about 570K tokens of reusable cache); 524K opt-in (--mode long-512k) · --vision for images · Pick how to run it | Results |
| Qwen3.8 27B Qronos | Chat, coding and optional reasoning · 8K · text | Results |
| Qwen3.8 NEO CODER MAX 27B | Coding and visual chat · 8K · text plus one image | Results |
| Qwen3-Coder 30B A3B | Code writing, review, testing and tools · 4K | Results |
| GPT-OSS-20B | Reasoning, coding, tools and JSON schemas · 8K | Results |
| Ornith 1.5 35B A3B | Chat and optional reasoning · 8K · text | Results |
| Model and setup | Use | Benchmark |
|---|---|---|
| FLUX.2 klein 4B | Text to image · ComfyUI, web or CLI | Results |
| Qwen-Image-2.1 MXFP4 | Image generation and editing · HTTP API or CLI | Results |
| Model and setup | Use | Benchmark |
|---|---|---|
| FastWan FullAttn 5B | Text to silent video · ComfyUI | Results |
| Wan2.2 TI2V-5B | Text or image to silent video · ComfyUI | Results |
| MiniMax H3 | Video with stereo audio · optional first/last frame · ComfyUI | Results |
Meeting is a review candidate for transcripts, speaker labels and timestamped notes from recordings. Notes need review against the recording; live Teams capture and Studio integration are unsupported. Benchmark.
The pip-installable plugin can run supported presets inside an existing matching vLLM build. It checks the checkpoint, runtime and native artifacts; it does not install or replace vLLM, PyTorch or ROCm.
Activate the preset's supported environment, then:
python -m pip install https://github.com/Eliovp-BV/paiton-vllm-plugin/releases/download/v0.3.4/paiton_vllm_plugin-0.3.4-py3-none-any.whl
paiton doctor
paiton serve minicpm5
See native execution for all presets, exact runtime
requirements, local weights and offline use. The native qwen38-nvfp4 preset
runs without DFlash2 or the 3-bit weights; use the Qwen3.8 container guide
for those release profiles and their benchmark results.
Cloning this repository gets the launchers and guides. Follow your model's guide to download its pinned weights or reuse an existing copy. See weights and caches for local folders and Docker mounts.
gfx1201. Other GPUs have not been qualified./dev/kfd and /dev/dri; ComfyUI launchers also need Docker Compose.| Looking for | Go to |
|---|---|
| Weights and cache setup | Model weights |
| Plugin presets, runtime requirements and offline use | Native execution |
| Bundle packaging and the older compatibility launcher | Native packaging · Existing vLLM |
| Runtime and wheel downloads | GitHub Releases · Containers |
| Published weights | Hugging Face · EliovpAI |
| Plugin and CLI source | paiton_vllm_plugin |
| Questions and bug reports | Issues |
This repository contains public integrations and compiled runtime artifacts. The compiler is proprietary. The plugin is Apache-2.0 licensed; model weights and bundled components retain their own licenses. See third-party notices and each model's notices.
Python
94.7%
HTML
3.1%
Run chat, coding, image and video models on AMD Radeon with Paiton's native GPU runtimes, integrated with vLLM, Diffusers and ComfyUI.
Using Qwen3.8? Start with its MXFP4 or 3-bit quickstart.
Download the weights, then pick how to run it:
MXFP4 for the highest accuracy, 3-bit for speed, --mode long for one long document, --mode long-kv4 for coding
agents (prefix caching, about 570K tokens of cache), opt-in --mode long-512k for up to 524K, --vision for images.
Prefer a desktop app? Paiton Studio provides chat, image and video tools for the supported models.
Models · Quick start · Requirements · Your own vLLM · Docs
The container launchers include each model's supported inference environment. Choose a model below and follow its guide. For the smallest language model, MiniCPM5-2B:
git clone --depth 1 https://github.com/Eliovp-BV/paiton-vllm-plugin.git
cd paiton-vllm-plugin
./models/MiniCPM5-2B/serve-docker.sh
The first launch downloads its weights and prepares runtime components; later launches reuse the caches. Its guide includes chat and API examples. Qwen3.8 and Meeting need weight preparation before launch. Run one model at a time on the tested single-GPU setup.
Model links lead to setup, launch options and support limits. Benchmark links include the tested settings and quality tradeoffs; results from different models or profiles are separate comparisons.
These models serve an OpenAI-compatible API through vLLM.
| Model and setup | Use | Benchmark |
|---|---|---|
| MiniCPM5-2B | Lightweight chat, coding and tools · 8K | Results |
| Qwen3.8 27B MXFP4 / 3-bit + DFlash2 | Chat and coding · MXFP4 (most accurate) or 3-bit (fastest) · 65K by default; 262K with --mode long (one long document, fast follow-ups) or --mode long-kv4 (coding agents, about 570K tokens of reusable cache); 524K opt-in (--mode long-512k) · --vision for images · Pick how to run it | Results |
| Qwen3.8 27B Qronos | Chat, coding and optional reasoning · 8K · text | Results |
| Qwen3.8 NEO CODER MAX 27B | Coding and visual chat · 8K · text plus one image | Results |
| Qwen3-Coder 30B A3B | Code writing, review, testing and tools · 4K | Results |
| GPT-OSS-20B | Reasoning, coding, tools and JSON schemas · 8K | Results |
| Ornith 1.5 35B A3B | Chat and optional reasoning · 8K · text | Results |
| Model and setup | Use | Benchmark |
|---|---|---|
| FLUX.2 klein 4B | Text to image · ComfyUI, web or CLI | Results |
| Qwen-Image-2.1 MXFP4 | Image generation and editing · HTTP API or CLI | Results |
| Model and setup | Use | Benchmark |
|---|---|---|
| FastWan FullAttn 5B | Text to silent video · ComfyUI | Results |
| Wan2.2 TI2V-5B | Text or image to silent video · ComfyUI | Results |
| MiniMax H3 | Video with stereo audio · optional first/last frame · ComfyUI | Results |
Meeting is a review candidate for transcripts, speaker labels and timestamped notes from recordings. Notes need review against the recording; live Teams capture and Studio integration are unsupported. Benchmark.
The pip-installable plugin can run supported presets inside an existing matching vLLM build. It checks the checkpoint, runtime and native artifacts; it does not install or replace vLLM, PyTorch or ROCm.
Activate the preset's supported environment, then:
python -m pip install https://github.com/Eliovp-BV/paiton-vllm-plugin/releases/download/v0.3.4/paiton_vllm_plugin-0.3.4-py3-none-any.whl
paiton doctor
paiton serve minicpm5
See native execution for all presets, exact runtime
requirements, local weights and offline use. The native qwen38-nvfp4 preset
runs without DFlash2 or the 3-bit weights; use the Qwen3.8 container guide
for those release profiles and their benchmark results.
Cloning this repository gets the launchers and guides. Follow your model's guide to download its pinned weights or reuse an existing copy. See weights and caches for local folders and Docker mounts.
gfx1201. Other GPUs have not been qualified./dev/kfd and /dev/dri; ComfyUI launchers also need Docker Compose.| Looking for | Go to |
|---|---|
| Weights and cache setup | Model weights |
| Plugin presets, runtime requirements and offline use | Native execution |
| Bundle packaging and the older compatibility launcher | Native packaging · Existing vLLM |
| Runtime and wheel downloads | GitHub Releases · Containers |
| Published weights | Hugging Face · EliovpAI |
| Plugin and CLI source | paiton_vllm_plugin |
| Questions and bug reports | Issues |
This repository contains public integrations and compiled runtime artifacts. The compiler is proprietary. The plugin is Apache-2.0 licensed; model weights and bundled components retain their own licenses. See third-party notices and each model's notices.
Python
94.7%
HTML
3.1%