Eliovp-BV/paiton-vllm-plugin

Python

47

96 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen3.8 27B | 1 x R9700: 262K context, half a million tokens of reusable cache, ~180 tok/s. And yes, let's talk about the "3-bit" :) (r/LocalLLaMA)

**TL;DR:** Qwen3.8 27B on a single AMD Radeon AI PRO R9700 (32 GB, 300 W), with speculative decoding and our 3-bit weights (not a blanket 3-bit quant, see below), now keeps **569,878 tokens of reusable cache** in its new coding mode. Every request gets 262,144 tokens of context, and two full-length…

0

Oct 4, 2026

README

Paiton

Run chat, coding, image and video models on AMD Radeon with Paiton's native GPU runtimes, integrated with vLLM, Diffusers and ComfyUI.

Using Qwen3.8? Start with its MXFP4 or 3-bit quickstart. Download the weights, then pick how to run it: MXFP4 for the highest accuracy, 3-bit for speed, --mode long for one long document, --mode long-kv4 for coding agents (prefix caching, about 570K tokens of cache), opt-in --mode long-512k for up to 524K, --vision for images.

Prefer a desktop app? Paiton Studio provides chat, image and video tools for the supported models.

Models · Quick start · Requirements · Your own vLLM · Docs

Quick start

The container launchers include each model's supported inference environment. Choose a model below and follow its guide. For the smallest language model, MiniCPM5-2B:

git clone --depth 1 https://github.com/Eliovp-BV/paiton-vllm-plugin.git
cd paiton-vllm-plugin
./models/MiniCPM5-2B/serve-docker.sh

The first launch downloads its weights and prepares runtime components; later launches reuse the caches. Its guide includes chat and API examples. Qwen3.8 and Meeting need weight preparation before launch. Run one model at a time on the tested single-GPU setup.

Models

Model links lead to setup, launch options and support limits. Benchmark links include the tested settings and quality tradeoffs; results from different models or profiles are separate comparisons.

Chat, reasoning and coding

These models serve an OpenAI-compatible API through vLLM.

Model and setupUseBenchmark
MiniCPM5-2BLightweight chat, coding and tools · 8KResults
Qwen3.8 27B MXFP4 / 3-bit + DFlash2Chat and coding · MXFP4 (most accurate) or 3-bit (fastest) · 65K by default; 262K with --mode long (one long document, fast follow-ups) or --mode long-kv4 (coding agents, about 570K tokens of reusable cache); 524K opt-in (--mode long-512k) · --vision for images · Pick how to run itResults
Qwen3.8 27B QronosChat, coding and optional reasoning · 8K · textResults
Qwen3.8 NEO CODER MAX 27BCoding and visual chat · 8K · text plus one imageResults
Qwen3-Coder 30B A3BCode writing, review, testing and tools · 4KResults
GPT-OSS-20BReasoning, coding, tools and JSON schemas · 8KResults
Ornith 1.5 35B A3BChat and optional reasoning · 8K · textResults

Image generation

Model and setupUseBenchmark
FLUX.2 klein 4BText to image · ComfyUI, web or CLIResults
Qwen-Image-2.1 MXFP4Image generation and editing · HTTP API or CLIResults

Video generation

Model and setupUseBenchmark
FastWan FullAttn 5BText to silent video · ComfyUIResults
Wan2.2 TI2V-5BText or image to silent video · ComfyUIResults
MiniMax H3Video with stereo audio · optional first/last frame · ComfyUIResults

Meeting notes

Meeting is a review candidate for transcripts, speaker labels and timestamped notes from recordings. Notes need review against the recording; live Teams capture and Studio integration are unsupported. Benchmark.

Use your own vLLM environment

The pip-installable plugin can run supported presets inside an existing matching vLLM build. It checks the checkpoint, runtime and native artifacts; it does not install or replace vLLM, PyTorch or ROCm.

Activate the preset's supported environment, then:

python -m pip install https://github.com/Eliovp-BV/paiton-vllm-plugin/releases/download/v0.3.4/paiton_vllm_plugin-0.3.4-py3-none-any.whl
paiton doctor
paiton serve minicpm5

See native execution for all presets, exact runtime requirements, local weights and offline use. The native qwen38-nvfp4 preset runs without DFlash2 or the 3-bit weights; use the Qwen3.8 container guide for those release profiles and their benchmark results.

Model weights and existing downloads

Cloning this repository gets the launchers and guides. Follow your model's guide to download its pinned weights or reuse an existing copy. See weights and caches for local folders and Docker mounts.

Requirements

  • Tested GPU: one Radeon AI PRO R9700, 32 GB, RDNA4 / gfx1201. Other GPUs have not been qualified.
  • Containers: Linux, Docker and AMD device access through /dev/kfd and /dev/dri; ComfyUI launchers also need Docker Compose.
  • Native plugin: the preset's exact supported environment, listed in native execution.
  • Memory and storage: depend on the model, context, concurrency and image/video settings; check the model guide before downloading.

Where to find things

Looking forGo to
Weights and cache setupModel weights
Plugin presets, runtime requirements and offline useNative execution
Bundle packaging and the older compatibility launcherNative packaging · Existing vLLM
Runtime and wheel downloadsGitHub Releases · Containers
Published weightsHugging Face · EliovpAI
Plugin and CLI sourcepaiton_vllm_plugin
Questions and bug reportsIssues

About Paiton

This repository contains public integrations and compiled runtime artifacts. The compiler is proprietary. The plugin is Apache-2.0 licensed; model weights and bundled components retain their own licenses. See third-party notices and each model's notices.

More about Paiton.

Eliovp-BV/paiton-vllm-plugin

Python

47

96 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen3.8 27B | 1 x R9700: 262K context, half a million tokens of reusable cache, ~180 tok/s. And yes, let's talk about the "3-bit" :) (r/LocalLLaMA)

**TL;DR:** Qwen3.8 27B on a single AMD Radeon AI PRO R9700 (32 GB, 300 W), with speculative decoding and our 3-bit weights (not a blanket 3-bit quant, see below), now keeps **569,878 tokens of reusable cache** in its new coding mode. Every request gets 262,144 tokens of context, and two full-length…

0

Oct 4, 2026

README

Paiton

Run chat, coding, image and video models on AMD Radeon with Paiton's native GPU runtimes, integrated with vLLM, Diffusers and ComfyUI.

Using Qwen3.8? Start with its MXFP4 or 3-bit quickstart. Download the weights, then pick how to run it: MXFP4 for the highest accuracy, 3-bit for speed, --mode long for one long document, --mode long-kv4 for coding agents (prefix caching, about 570K tokens of cache), opt-in --mode long-512k for up to 524K, --vision for images.

Prefer a desktop app? Paiton Studio provides chat, image and video tools for the supported models.

Models · Quick start · Requirements · Your own vLLM · Docs

Quick start

The container launchers include each model's supported inference environment. Choose a model below and follow its guide. For the smallest language model, MiniCPM5-2B:

git clone --depth 1 https://github.com/Eliovp-BV/paiton-vllm-plugin.git
cd paiton-vllm-plugin
./models/MiniCPM5-2B/serve-docker.sh

The first launch downloads its weights and prepares runtime components; later launches reuse the caches. Its guide includes chat and API examples. Qwen3.8 and Meeting need weight preparation before launch. Run one model at a time on the tested single-GPU setup.

Models

Model links lead to setup, launch options and support limits. Benchmark links include the tested settings and quality tradeoffs; results from different models or profiles are separate comparisons.

Chat, reasoning and coding

These models serve an OpenAI-compatible API through vLLM.

Model and setupUseBenchmark
MiniCPM5-2BLightweight chat, coding and tools · 8KResults
Qwen3.8 27B MXFP4 / 3-bit + DFlash2Chat and coding · MXFP4 (most accurate) or 3-bit (fastest) · 65K by default; 262K with --mode long (one long document, fast follow-ups) or --mode long-kv4 (coding agents, about 570K tokens of reusable cache); 524K opt-in (--mode long-512k) · --vision for images · Pick how to run itResults
Qwen3.8 27B QronosChat, coding and optional reasoning · 8K · textResults
Qwen3.8 NEO CODER MAX 27BCoding and visual chat · 8K · text plus one imageResults
Qwen3-Coder 30B A3BCode writing, review, testing and tools · 4KResults
GPT-OSS-20BReasoning, coding, tools and JSON schemas · 8KResults
Ornith 1.5 35B A3BChat and optional reasoning · 8K · textResults

Image generation

Model and setupUseBenchmark
FLUX.2 klein 4BText to image · ComfyUI, web or CLIResults
Qwen-Image-2.1 MXFP4Image generation and editing · HTTP API or CLIResults

Video generation

Model and setupUseBenchmark
FastWan FullAttn 5BText to silent video · ComfyUIResults
Wan2.2 TI2V-5BText or image to silent video · ComfyUIResults
MiniMax H3Video with stereo audio · optional first/last frame · ComfyUIResults

Meeting notes

Meeting is a review candidate for transcripts, speaker labels and timestamped notes from recordings. Notes need review against the recording; live Teams capture and Studio integration are unsupported. Benchmark.

Use your own vLLM environment

The pip-installable plugin can run supported presets inside an existing matching vLLM build. It checks the checkpoint, runtime and native artifacts; it does not install or replace vLLM, PyTorch or ROCm.

Activate the preset's supported environment, then:

python -m pip install https://github.com/Eliovp-BV/paiton-vllm-plugin/releases/download/v0.3.4/paiton_vllm_plugin-0.3.4-py3-none-any.whl
paiton doctor
paiton serve minicpm5

See native execution for all presets, exact runtime requirements, local weights and offline use. The native qwen38-nvfp4 preset runs without DFlash2 or the 3-bit weights; use the Qwen3.8 container guide for those release profiles and their benchmark results.

Model weights and existing downloads

Cloning this repository gets the launchers and guides. Follow your model's guide to download its pinned weights or reuse an existing copy. See weights and caches for local folders and Docker mounts.

Requirements

  • Tested GPU: one Radeon AI PRO R9700, 32 GB, RDNA4 / gfx1201. Other GPUs have not been qualified.
  • Containers: Linux, Docker and AMD device access through /dev/kfd and /dev/dri; ComfyUI launchers also need Docker Compose.
  • Native plugin: the preset's exact supported environment, listed in native execution.
  • Memory and storage: depend on the model, context, concurrency and image/video settings; check the model guide before downloading.

Where to find things

Looking forGo to
Weights and cache setupModel weights
Plugin presets, runtime requirements and offline useNative execution
Bundle packaging and the older compatibility launcherNative packaging · Existing vLLM
Runtime and wheel downloadsGitHub Releases · Containers
Published weightsHugging Face · EliovpAI
Plugin and CLI sourcepaiton_vllm_plugin
Questions and bug reportsIssues

About Paiton

This repository contains public integrations and compiled runtime artifacts. The compiler is proprietary. The plugin is Apache-2.0 licensed; model weights and bundled components retain their own licenses. See third-party notices and each model's notices.

More about Paiton.

Languages

Python

94.7%

HTML

3.1%