julianmb/halofpx

High-performance Lemonade alternative for AMD Strix Halo & Radeon — latest MoE models (Ling-3.0-Flash, Qwen 3.8 Flash Next, Ornith 1.5), bleeding-edge RDNA 3.5 Wave64/ROCmFP4 kernels, and silicon-tuned profiles.

Python

84

94 commits

updated Sep 24, 2026

See the code

README

HaloFPX — High-Performance Model Server & Zoo for AMD Strix Halo

Hardware Vulkan FastAPI License

The high-performance, cutting-edge Lemonade alternative engineered for AMD Strix Halo (Ryzen AI Max APUs / 64GB–128GB UMA) and AMD Radeon GPUs.

HaloFPX delivers the familiar developer experience and single-endpoint architecture of Lemonade, but supercharged with the latest frontier models, bleeding-edge RDNA 3.5 matrix kernels, and silicon-tuned profiles not available in standard Lemonade.


🍋 HaloFPX vs. Standard Lemonade — Why HaloFPX?

HaloFPX is engineered from the silicon up for AMD Strix Halo (gfx1151) and high-end Radeon GPUs. While maintaining complete CLI and model-cache interoperability with Lemonade, it unlocks performance and architectures standard Lemonade cannot reach:

CapabilityStandard LemonadeHaloFPX 🚀
Next-Gen MoE ModelsGeneric catalog; misses latest large MoE quantsDay-0 support: Ling-3.0-Flash (124B), Qwen 3.8 Flash Next (125B), DeepSeek V4 Flash (284B), Ornith 1.5 (35B)
Compute KernelsGeneric upstream llama.cpp / stock ROCm kernelsMesa RADV Wave64 cooperative matrices (KHR_coopmat) on RDNA 3.5 + Dual-Backend ROCm/Vulkan
Quantization FormatsStock GGUF (Q4_K_M, standard quants)Custom ROCmFP4 / ROCmFP4_FAST (direct hardware block mapping, −16.7% size, +13.5% decode speed)
Silicon-Tuned ProfilesOne-size-fits-all server flagsHardware-benchmarked run_config: tuned MTP speculation, TurboQuant KV (q8_0), -ctxcp, -cram, zero flag guessing
Long Context ScalingStandard context limits; risks OOM on unified APUsValidated 262K context (524K tokens across 4 slots) with TurboQuant KV in unified memory
CLI & Workflow Paritylemonade {run, list, chat, backends, delete}100% command parity: halofpx run/chat/list/backends/delete, native drop-in, zero retraining
Lemonade Cache SharingStandalone cache (/var/lib/lemonade/models)Bi-directional cache sharing: reads Lemonade user models, shares weights, resolves aliases seamlessly
Multimodal VisionManual projector configurationAutomatic vision discovery & loading (mmproj) verified over standard OpenAI /v1/chat/completions

⚡ Quick Instructions — How to Run a Model

Run any model in seconds using familiar Lemonade-compatible commands or launch the OpenAI-compatible single-endpoint server. Silicon-tuned configurations (run_config, TurboQuant KV, Wave64 cooperative matrices) are applied automatically with zero flag guessing.

1. Interactive Terminal Chat (halofpx run)

# List available models and check download status
halofpx list

# Launch an interactive chat session with any model (auto-loads tuned config):
halofpx run Ling-3.0-Flash

# Or run Qwen 3.8 Flash Next or Ornith 1.5:
halofpx run qwen38-flash-next
halofpx run ornith-1.5-35b

2. Start OpenAI-Compatible Server (halofpx serve)

# Start server with an auto-loaded model on http://localhost:8010
halofpx serve -m Ling-3.0-Flash

# Query via standard OpenAI /v1/chat/completions:
curl http://localhost:8010/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Ling-3.0-Flash",
    "messages": [{"role": "user", "content": "Explain quantum computing concisely."}]
  }'

3. Dynamic Model Switching & Telemetry

# Hot-swap to a different model dynamically without restarting the server:
halofpx load qwen38-flash-next

# Check APU memory residency and real-time status:
halofpx status

# Unload from memory when finished:
halofpx unload

  • julianmb/q38rocm: Dedicated single-model deep-dive and standalone deployment package specifically for Qwen 3.8 27B on AMD Strix Halo (up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64).
  • julianmb/haloq38flash: Dedicated optimization deep-dive for Qwen 3.8 Flash Next 125B MoE on AMD Strix Halo (split-PLE quant, up to 56 tok/s, 262K context).

🚀 Key Features

  • 📦 Unified Model Zoo: Download, verify, and serve pre-quantized models (Ornith 1.5 35B, Qwen 3.8 27B, Nemotron 3.5 30B, DeepSeek V4 Flash, Laguna S 2.1) directly from Hugging Face.
  • 🎯 Per-Model Optimization Profiles: Each model carries a benchmarked run_config — quant preset, KV-cache types, backend preference, MTP on/off — applied automatically on load. No flag archaeology.
  • 👁️ Automatic Vision (Multimodal): Models with a projector (ornith-1.5-35b) pull and verify their mmproj alongside the weights; image prompts work over the standard OpenAI API.
  • 📏 Validated Long Context: Ornith validated at the full 262K training context; TurboQuant KV enables 4-slot × 131K contexts (524K total tokens) in unified memory with zero OOM.
  • 🎮 Dynamic AMD Hardware Detection: Auto-detects compute targets (gfx1151, gfx1201, etc.) and applies hardware-specific execution flags.
  • 🔄 Hot-Swappable Memory Management: Dynamically load and unload models into available unified memory or dedicated VRAM with automatic GPU memory reclamation.
  • ⚡ Dual-Backend Hardware Acceleration:
    • Vulkan0 (Mesa RADV Wave64): Fastest token decode and MTP speculative tree verification (up to 36 tok/s on 27B).
    • ROCm0 (HIP): High-throughput prompt evaluation / prefill processing (up to 390+ tok/s).
  • 🚀 Measured Speed Increase Over Standard GGUF: ROCmFP4/ROCmFP4_FAST quants beat stock Q4_K_M in decode throughput and size on Strix Halo (gfx1151). See the benchmark table below.
  • 🔒 Optional API Key Authentication: Secure your endpoints via HALOFPX_API_KEY (disabled by default for local development).
  • 🦙 Ollama-Compatible API: /api/tags, /api/chat, /api/generate and /api/version let existing Ollama clients and tools work against halofpx drop-in.
  • 🌐 Standard OpenAI API & Management API: Standard /v1/chat/completions (with streaming SSE) plus /api/v1/{pull, load, unload, status, system-info} endpoints on a single port (8010).
  • 🐳 Modular Docker Compose: Run lightweight standalone or pair with Open WebUI via --profile webui.

📦 Model Zoo & Verified Hardware Benchmarks

All models ship pre-quantized with silicon-tuned execution profiles. Measurements taken directly on AMD Ryzen AI Max+ 395 (40 CU Radeon 8060S @ 2.9 GHz, 128 GB 256-bit LPDDR5X, Linux 7.0, Mesa 26.0 RADV Wave64 / Vulkan0):

Model Name & CLI IDCategoryQuant & SizeMin VRAMPrefill (pp512)Decode (Bare / MTP)Verified DateHF Repository & Quick Command
Qwen 3.8 Flash Next 125B MoE ⚡
qwen38-flash-next
Next-Gen MoE / Hybrid AttentionROCmFP4_ple16
(87.1 GiB)
68 GB413.09 tok/s (CM1)28.96 / 🔥 34.60 tok/s2026-09-10unsloth/Qwen3.8-Flash-Next-GGUF
halofpx run qwen38-flash-next
Ling 3.0 Flash 124B MoE ⚡
ling3-flash
Frontier 124B MoE / ReasoningQ4_K_M
(71.7 GiB)
80 GB340.50 tok/s45.23 / 🔥 55.10 tok/s2026-09-02inclusionAI/Ling-3.0-flash-GGUF
halofpx run Ling-3.0-Flash
Nex N2.5 Mini 35B-A3B MoE ⭐
nex-n2.5-mini
Agentic Reasoning / VisionROCmFP4_LEAN
(17.3 GiB)
22 GB1180.20 tok/s76.92 tok/s2026-08-30julianmb/Nex-N2.5-mini-ROCmFP4-GGUF
halofpx run nex-n2.5-mini
Ornith 1.5 35B-A3B MoE ⭐
ornith-1.5-35b
Agentic Coding / Vision MoEROCmFP4
(18.2 GiB)
22 GB1190.40 tok/s76.90 / 🔥 105.60 tok/s2026-08-28julianmb/Ornith-1.5-35B-A3B-ROCmFP4-GGUF
halofpx run ornith-1.5-35b
Qwen 3.8 / 27B UltraQuality
qwen38-27b
Dense / ReasoningROCmFP4_FAST
(13.5 GiB)
16 GB385.65 tok/s14.02 / 🔥 33.80 tok/s2026-08-25julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
halofpx run qwen38-27b
NVIDIA Nemotron 3.5 Lightning 30B
nemotron-3.5-30b
High-Speed MoEROCmFP4_FAST
(14.8 GiB)
16 GB1286.24 tok/s52.40 / 🔥 95.20 tok/s2026-08-22julianmb/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-ROCmFP4-GGUF
halofpx pull nemotron-3.5-30b
DeepSeek V4 Flash 284B MoE
deepseek-v4-flash
Ultra-Scale MoEIQ2_XXS
(86.7 GiB)
90 GB195.40 tok/s22.50 / 32.00 tok/s2026-08-20julianmb/DeepSeek-V4-Flash-0731-IQ2XXS-STRIX
halofpx pull deepseek-v4-flash
Laguna S 2.1 StrixKVSpine v4
laguna-s21
General ChatROCmFP4_Spine
(60.9 GiB)
64 GB390.13 tok/s34.05 tok/s2026-08-15julianmb/Laguna-S-2.1-ROCmFP4-StrixKVSpine-v4
halofpx pull laguna-s21
Ornith 1.0 35B ROCmFPX
ornith-35b
Multi-Slot AgentROCmFPX_Speed
(18.4 GiB)
22 GB1194.19 tok/s11.20 / 115.0+ tok/s (16s)2026-08-15julianmb/Ornith-1.0-35B-ROCmFPX-StrixHalo
halofpx pull ornith-35b

⭐ = vision-capable (mmproj included). ⚡ = cutting-edge large MoE architecture. Full methodology: docs/BENCHMARKS.md.

👉 See Hardware Support & VRAM Sizing Guide (docs/HARDWARE_SUPPORT.md) for memory sizing tables across AMD APUs and discrete GPUs.

🔬 How to Run & Record Benchmarks

HaloFPX provides automated benchmarking runners to measure prompt prefill throughput, token decode speed, TTFT, and MTP draft acceptance rates:

# 1. Automated performance benchmark (auto-exports timestamped Markdown & JSON reports)
python3 scripts/benchmark.py --model qwen38-27b --device Vulkan0

# 2. Context scaling benchmark up to 262K tokens
python3 scripts/context_scaling_benchmark.py --model ornith-1.5-35b

# 3. Low-level engine benchmark using llama-bench
llama-bench -m models/qwen38-flash-next/Qwen3.8-Flash-Next-ROCmFP4-FAST-v2-ple16.gguf -p 512 -n 128 -ngl 99

Benchmark outputs are stored with timestamps under benchmarks/benchmark_YYYYMMDD_HHMMSS.md and benchmarks/benchmark_YYYYMMDD_HHMMSS.json.


🚀 Speed Increase Over Standard GGUF

Measured on AMD Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, Mesa RADV Wave64) with identical prompts — ROCmFP4-family quants beat stock Q4_K_M on decode speed and model size:

ModelStock Q4_K_MROCmFP4 / ROCmFP4_FASTDecode SpeedupSize Savings
Qwen 3.8 27B15.92 GiB — 12.35 tok/s13.55 GiB — 14.02 tok/s+13.5%−14.9%
Ornith 1.5 35B-A3B21.80 GiB — 71.5–71.7 tok/s18.16 GiB — 76.9 tok/s+7.5%−16.7%
Nex N2.5 Mini 35B-A3B19.71 GiB — 72.76 tok/s17.32 GiB — 76.92 tok/s+5.7%−12.1%

Additional gains over stock GGUF on Strix Halo:

  • Prefill: ROCmFP4 quant blocks map directly to RDNA 3.5 cooperative-matrix (KHR_coopmat) operands — faster prompt evaluation at equal context, without the multi-scale dequantization overhead of Q4_K blocks.
  • Combined with MTP speculative decoding (Qwen 3.8 27B, n4/p0.0): 33.8 tok/s sustained = 2.40× over stock baseline (12.35 tok/s).
  • KV cache: TurboQuant KV (q8_0) shrinks memory footprint, enabling 4-slot × 131K contexts (524K total tokens) in unified memory with zero OOM.

Full methodology and raw numbers: docs/BENCHMARKS.md.


🛠️ Installation & Advanced Setup

1. Installation

# Clone the repository
git clone https://github.com/julianmb/halofpx.git
cd halofpx

# Install Python requirements and CLI
pip install -r requirements.txt
pip install -e .

# Set up environment variables for AMD Strix Halo
source ./scripts/setup_env.sh

2. Backend Inspection & Lemonade Cache Sync

# Check available hardware acceleration backends (ROCm HIP, Vulkan RADV Wave64)
halofpx backends

# Inspect configuration or synchronize cache with Lemonade daemon
halofpx config show
halofpx config sync-lemonade

3. Advanced Workload Tuning & Concurrency

# Parallel multi-agent concurrency (4 slots -> ~40.5 tok/s aggregate)
halofpx load qwen38-27b --slots 4 --draft-n 6 --draft-p 0.60

# Single-user interactive chat with burst MTP
halofpx load qwen38-27b --draft-n 5 --draft-p 0.50

# Long context scaling up to 262K tokens with TurboQuant KV in unified memory
halofpx load ornith-1.5-35b --ctx 262144

🔌 Client & IDE Integration

Connect your local tools to http://localhost:8010/v1:

  • Open WebUI: Set Base URL to http://localhost:8010/v1 and API Key to sk-no-key.
  • Continue.dev: Add halofpx as provider in ~/.continue/config.json.
  • Cursor IDE: Override OpenAI Base URL to http://localhost:8010/v1.

👉 See the complete Client Integration Guide (docs/CLIENT_INTEGRATION.md).


🐳 Docker Deployment Options

Option A: Lightweight Standalone Server (Default)

Runs only the high-performance HaloFPX server (zero extra RAM overhead for web frontends):

docker compose up -d
  • API Endpoint: http://localhost:8010/v1

Option B: Server + Open WebUI Chat Interface

Runs both the backend server and Open WebUI in a unified stack:

docker compose --profile webui up -d
  • HaloFPX API: http://localhost:8010/v1
  • Open WebUI: http://localhost:3000

Option C: Direct docker run

docker run -d -p 8010:8010 \
  --device=/dev/dri \
  --group-add video --group-add render \
  --ipc=host \
  -v $(pwd)/models:/app/models \
  -v ~/.cache/huggingface/hub:/root/.cache/huggingface/hub \
  --name halofpx-server \
  ghcr.io/julianmb/halofpx:latest
# Note: For optional ROCm prefill acceleration, also add --device=/dev/kfd and use ghcr.io/julianmb/halofpx:rocm

👉 See the complete Docker Deployment Guide (docs/DOCKER_GUIDE.md) for GPU passthrough prerequisites, container CLI commands, and local builds.


🤝 Upstream Integration & Engine Core

HaloFPX wraps and orchestrates the charlie12345/ROCmFPX engine, compiling directly against pinned builds (e87d53e (213)) or downloading pre-compiled Strix Halo binaries via ./scripts/build_engine.sh --prebuilt.


📄 License

Apache 2.0 License.

amd-ryzen-ai
fp4
gguf
halofpx
lemonade
llama-cpp
llm-inference
local-llm
model-server
moe
multimodal
openai-api
ornith
quantization
radv
rocm
speculative-decoding
strix-halo
vulkan
wave64

Significant stargazers

Goni Zahavy

23 followers · starred Aug 2026

julianmb/halofpx

High-performance Lemonade alternative for AMD Strix Halo & Radeon — latest MoE models (Ling-3.0-Flash, Qwen 3.8 Flash Next, Ornith 1.5), bleeding-edge RDNA 3.5 Wave64/ROCmFP4 kernels, and silicon-tuned profiles.

Python

84

94 commits

updated Sep 24, 2026

See the code

README

HaloFPX — High-Performance Model Server & Zoo for AMD Strix Halo

Hardware Vulkan FastAPI License

The high-performance, cutting-edge Lemonade alternative engineered for AMD Strix Halo (Ryzen AI Max APUs / 64GB–128GB UMA) and AMD Radeon GPUs.

HaloFPX delivers the familiar developer experience and single-endpoint architecture of Lemonade, but supercharged with the latest frontier models, bleeding-edge RDNA 3.5 matrix kernels, and silicon-tuned profiles not available in standard Lemonade.


🍋 HaloFPX vs. Standard Lemonade — Why HaloFPX?

HaloFPX is engineered from the silicon up for AMD Strix Halo (gfx1151) and high-end Radeon GPUs. While maintaining complete CLI and model-cache interoperability with Lemonade, it unlocks performance and architectures standard Lemonade cannot reach:

CapabilityStandard LemonadeHaloFPX 🚀
Next-Gen MoE ModelsGeneric catalog; misses latest large MoE quantsDay-0 support: Ling-3.0-Flash (124B), Qwen 3.8 Flash Next (125B), DeepSeek V4 Flash (284B), Ornith 1.5 (35B)
Compute KernelsGeneric upstream llama.cpp / stock ROCm kernelsMesa RADV Wave64 cooperative matrices (KHR_coopmat) on RDNA 3.5 + Dual-Backend ROCm/Vulkan
Quantization FormatsStock GGUF (Q4_K_M, standard quants)Custom ROCmFP4 / ROCmFP4_FAST (direct hardware block mapping, −16.7% size, +13.5% decode speed)
Silicon-Tuned ProfilesOne-size-fits-all server flagsHardware-benchmarked run_config: tuned MTP speculation, TurboQuant KV (q8_0), -ctxcp, -cram, zero flag guessing
Long Context ScalingStandard context limits; risks OOM on unified APUsValidated 262K context (524K tokens across 4 slots) with TurboQuant KV in unified memory
CLI & Workflow Paritylemonade {run, list, chat, backends, delete}100% command parity: halofpx run/chat/list/backends/delete, native drop-in, zero retraining
Lemonade Cache SharingStandalone cache (/var/lib/lemonade/models)Bi-directional cache sharing: reads Lemonade user models, shares weights, resolves aliases seamlessly
Multimodal VisionManual projector configurationAutomatic vision discovery & loading (mmproj) verified over standard OpenAI /v1/chat/completions

⚡ Quick Instructions — How to Run a Model

Run any model in seconds using familiar Lemonade-compatible commands or launch the OpenAI-compatible single-endpoint server. Silicon-tuned configurations (run_config, TurboQuant KV, Wave64 cooperative matrices) are applied automatically with zero flag guessing.

1. Interactive Terminal Chat (halofpx run)

# List available models and check download status
halofpx list

# Launch an interactive chat session with any model (auto-loads tuned config):
halofpx run Ling-3.0-Flash

# Or run Qwen 3.8 Flash Next or Ornith 1.5:
halofpx run qwen38-flash-next
halofpx run ornith-1.5-35b

2. Start OpenAI-Compatible Server (halofpx serve)

# Start server with an auto-loaded model on http://localhost:8010
halofpx serve -m Ling-3.0-Flash

# Query via standard OpenAI /v1/chat/completions:
curl http://localhost:8010/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Ling-3.0-Flash",
    "messages": [{"role": "user", "content": "Explain quantum computing concisely."}]
  }'

3. Dynamic Model Switching & Telemetry

# Hot-swap to a different model dynamically without restarting the server:
halofpx load qwen38-flash-next

# Check APU memory residency and real-time status:
halofpx status

# Unload from memory when finished:
halofpx unload

  • julianmb/q38rocm: Dedicated single-model deep-dive and standalone deployment package specifically for Qwen 3.8 27B on AMD Strix Halo (up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64).
  • julianmb/haloq38flash: Dedicated optimization deep-dive for Qwen 3.8 Flash Next 125B MoE on AMD Strix Halo (split-PLE quant, up to 56 tok/s, 262K context).

🚀 Key Features

  • 📦 Unified Model Zoo: Download, verify, and serve pre-quantized models (Ornith 1.5 35B, Qwen 3.8 27B, Nemotron 3.5 30B, DeepSeek V4 Flash, Laguna S 2.1) directly from Hugging Face.
  • 🎯 Per-Model Optimization Profiles: Each model carries a benchmarked run_config — quant preset, KV-cache types, backend preference, MTP on/off — applied automatically on load. No flag archaeology.
  • 👁️ Automatic Vision (Multimodal): Models with a projector (ornith-1.5-35b) pull and verify their mmproj alongside the weights; image prompts work over the standard OpenAI API.
  • 📏 Validated Long Context: Ornith validated at the full 262K training context; TurboQuant KV enables 4-slot × 131K contexts (524K total tokens) in unified memory with zero OOM.
  • 🎮 Dynamic AMD Hardware Detection: Auto-detects compute targets (gfx1151, gfx1201, etc.) and applies hardware-specific execution flags.
  • 🔄 Hot-Swappable Memory Management: Dynamically load and unload models into available unified memory or dedicated VRAM with automatic GPU memory reclamation.
  • ⚡ Dual-Backend Hardware Acceleration:
    • Vulkan0 (Mesa RADV Wave64): Fastest token decode and MTP speculative tree verification (up to 36 tok/s on 27B).
    • ROCm0 (HIP): High-throughput prompt evaluation / prefill processing (up to 390+ tok/s).
  • 🚀 Measured Speed Increase Over Standard GGUF: ROCmFP4/ROCmFP4_FAST quants beat stock Q4_K_M in decode throughput and size on Strix Halo (gfx1151). See the benchmark table below.
  • 🔒 Optional API Key Authentication: Secure your endpoints via HALOFPX_API_KEY (disabled by default for local development).
  • 🦙 Ollama-Compatible API: /api/tags, /api/chat, /api/generate and /api/version let existing Ollama clients and tools work against halofpx drop-in.
  • 🌐 Standard OpenAI API & Management API: Standard /v1/chat/completions (with streaming SSE) plus /api/v1/{pull, load, unload, status, system-info} endpoints on a single port (8010).
  • 🐳 Modular Docker Compose: Run lightweight standalone or pair with Open WebUI via --profile webui.

📦 Model Zoo & Verified Hardware Benchmarks

All models ship pre-quantized with silicon-tuned execution profiles. Measurements taken directly on AMD Ryzen AI Max+ 395 (40 CU Radeon 8060S @ 2.9 GHz, 128 GB 256-bit LPDDR5X, Linux 7.0, Mesa 26.0 RADV Wave64 / Vulkan0):

Model Name & CLI IDCategoryQuant & SizeMin VRAMPrefill (pp512)Decode (Bare / MTP)Verified DateHF Repository & Quick Command
Qwen 3.8 Flash Next 125B MoE ⚡
qwen38-flash-next
Next-Gen MoE / Hybrid AttentionROCmFP4_ple16
(87.1 GiB)
68 GB413.09 tok/s (CM1)28.96 / 🔥 34.60 tok/s2026-09-10unsloth/Qwen3.8-Flash-Next-GGUF
halofpx run qwen38-flash-next
Ling 3.0 Flash 124B MoE ⚡
ling3-flash
Frontier 124B MoE / ReasoningQ4_K_M
(71.7 GiB)
80 GB340.50 tok/s45.23 / 🔥 55.10 tok/s2026-09-02inclusionAI/Ling-3.0-flash-GGUF
halofpx run Ling-3.0-Flash
Nex N2.5 Mini 35B-A3B MoE ⭐
nex-n2.5-mini
Agentic Reasoning / VisionROCmFP4_LEAN
(17.3 GiB)
22 GB1180.20 tok/s76.92 tok/s2026-08-30julianmb/Nex-N2.5-mini-ROCmFP4-GGUF
halofpx run nex-n2.5-mini
Ornith 1.5 35B-A3B MoE ⭐
ornith-1.5-35b
Agentic Coding / Vision MoEROCmFP4
(18.2 GiB)
22 GB1190.40 tok/s76.90 / 🔥 105.60 tok/s2026-08-28julianmb/Ornith-1.5-35B-A3B-ROCmFP4-GGUF
halofpx run ornith-1.5-35b
Qwen 3.8 / 27B UltraQuality
qwen38-27b
Dense / ReasoningROCmFP4_FAST
(13.5 GiB)
16 GB385.65 tok/s14.02 / 🔥 33.80 tok/s2026-08-25julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
halofpx run qwen38-27b
NVIDIA Nemotron 3.5 Lightning 30B
nemotron-3.5-30b
High-Speed MoEROCmFP4_FAST
(14.8 GiB)
16 GB1286.24 tok/s52.40 / 🔥 95.20 tok/s2026-08-22julianmb/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-ROCmFP4-GGUF
halofpx pull nemotron-3.5-30b
DeepSeek V4 Flash 284B MoE
deepseek-v4-flash
Ultra-Scale MoEIQ2_XXS
(86.7 GiB)
90 GB195.40 tok/s22.50 / 32.00 tok/s2026-08-20julianmb/DeepSeek-V4-Flash-0731-IQ2XXS-STRIX
halofpx pull deepseek-v4-flash
Laguna S 2.1 StrixKVSpine v4
laguna-s21
General ChatROCmFP4_Spine
(60.9 GiB)
64 GB390.13 tok/s34.05 tok/s2026-08-15julianmb/Laguna-S-2.1-ROCmFP4-StrixKVSpine-v4
halofpx pull laguna-s21
Ornith 1.0 35B ROCmFPX
ornith-35b
Multi-Slot AgentROCmFPX_Speed
(18.4 GiB)
22 GB1194.19 tok/s11.20 / 115.0+ tok/s (16s)2026-08-15julianmb/Ornith-1.0-35B-ROCmFPX-StrixHalo
halofpx pull ornith-35b

⭐ = vision-capable (mmproj included). ⚡ = cutting-edge large MoE architecture. Full methodology: docs/BENCHMARKS.md.

👉 See Hardware Support & VRAM Sizing Guide (docs/HARDWARE_SUPPORT.md) for memory sizing tables across AMD APUs and discrete GPUs.

🔬 How to Run & Record Benchmarks

HaloFPX provides automated benchmarking runners to measure prompt prefill throughput, token decode speed, TTFT, and MTP draft acceptance rates:

# 1. Automated performance benchmark (auto-exports timestamped Markdown & JSON reports)
python3 scripts/benchmark.py --model qwen38-27b --device Vulkan0

# 2. Context scaling benchmark up to 262K tokens
python3 scripts/context_scaling_benchmark.py --model ornith-1.5-35b

# 3. Low-level engine benchmark using llama-bench
llama-bench -m models/qwen38-flash-next/Qwen3.8-Flash-Next-ROCmFP4-FAST-v2-ple16.gguf -p 512 -n 128 -ngl 99

Benchmark outputs are stored with timestamps under benchmarks/benchmark_YYYYMMDD_HHMMSS.md and benchmarks/benchmark_YYYYMMDD_HHMMSS.json.


🚀 Speed Increase Over Standard GGUF

Measured on AMD Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, Mesa RADV Wave64) with identical prompts — ROCmFP4-family quants beat stock Q4_K_M on decode speed and model size:

ModelStock Q4_K_MROCmFP4 / ROCmFP4_FASTDecode SpeedupSize Savings
Qwen 3.8 27B15.92 GiB — 12.35 tok/s13.55 GiB — 14.02 tok/s+13.5%−14.9%
Ornith 1.5 35B-A3B21.80 GiB — 71.5–71.7 tok/s18.16 GiB — 76.9 tok/s+7.5%−16.7%
Nex N2.5 Mini 35B-A3B19.71 GiB — 72.76 tok/s17.32 GiB — 76.92 tok/s+5.7%−12.1%

Additional gains over stock GGUF on Strix Halo:

  • Prefill: ROCmFP4 quant blocks map directly to RDNA 3.5 cooperative-matrix (KHR_coopmat) operands — faster prompt evaluation at equal context, without the multi-scale dequantization overhead of Q4_K blocks.
  • Combined with MTP speculative decoding (Qwen 3.8 27B, n4/p0.0): 33.8 tok/s sustained = 2.40× over stock baseline (12.35 tok/s).
  • KV cache: TurboQuant KV (q8_0) shrinks memory footprint, enabling 4-slot × 131K contexts (524K total tokens) in unified memory with zero OOM.

Full methodology and raw numbers: docs/BENCHMARKS.md.


🛠️ Installation & Advanced Setup

1. Installation

# Clone the repository
git clone https://github.com/julianmb/halofpx.git
cd halofpx

# Install Python requirements and CLI
pip install -r requirements.txt
pip install -e .

# Set up environment variables for AMD Strix Halo
source ./scripts/setup_env.sh

2. Backend Inspection & Lemonade Cache Sync

# Check available hardware acceleration backends (ROCm HIP, Vulkan RADV Wave64)
halofpx backends

# Inspect configuration or synchronize cache with Lemonade daemon
halofpx config show
halofpx config sync-lemonade

3. Advanced Workload Tuning & Concurrency

# Parallel multi-agent concurrency (4 slots -> ~40.5 tok/s aggregate)
halofpx load qwen38-27b --slots 4 --draft-n 6 --draft-p 0.60

# Single-user interactive chat with burst MTP
halofpx load qwen38-27b --draft-n 5 --draft-p 0.50

# Long context scaling up to 262K tokens with TurboQuant KV in unified memory
halofpx load ornith-1.5-35b --ctx 262144

🔌 Client & IDE Integration

Connect your local tools to http://localhost:8010/v1:

  • Open WebUI: Set Base URL to http://localhost:8010/v1 and API Key to sk-no-key.
  • Continue.dev: Add halofpx as provider in ~/.continue/config.json.
  • Cursor IDE: Override OpenAI Base URL to http://localhost:8010/v1.

👉 See the complete Client Integration Guide (docs/CLIENT_INTEGRATION.md).


🐳 Docker Deployment Options

Option A: Lightweight Standalone Server (Default)

Runs only the high-performance HaloFPX server (zero extra RAM overhead for web frontends):

docker compose up -d
  • API Endpoint: http://localhost:8010/v1

Option B: Server + Open WebUI Chat Interface

Runs both the backend server and Open WebUI in a unified stack:

docker compose --profile webui up -d
  • HaloFPX API: http://localhost:8010/v1
  • Open WebUI: http://localhost:3000

Option C: Direct docker run

docker run -d -p 8010:8010 \
  --device=/dev/dri \
  --group-add video --group-add render \
  --ipc=host \
  -v $(pwd)/models:/app/models \
  -v ~/.cache/huggingface/hub:/root/.cache/huggingface/hub \
  --name halofpx-server \
  ghcr.io/julianmb/halofpx:latest
# Note: For optional ROCm prefill acceleration, also add --device=/dev/kfd and use ghcr.io/julianmb/halofpx:rocm

👉 See the complete Docker Deployment Guide (docs/DOCKER_GUIDE.md) for GPU passthrough prerequisites, container CLI commands, and local builds.


🤝 Upstream Integration & Engine Core

HaloFPX wraps and orchestrates the charlie12345/ROCmFPX engine, compiling directly against pinned builds (e87d53e (213)) or downloading pre-compiled Strix Halo binaries via ./scripts/build_engine.sh --prebuilt.


📄 License

Apache 2.0 License.

amd-ryzen-ai
fp4
gguf
halofpx
lemonade
llama-cpp
llm-inference
local-llm
model-server
moe
multimodal
openai-api
ornith
quantization
radv
rocm
speculative-decoding
strix-halo
vulkan
wave64

Significant stargazers

Goni Zahavy

23 followers · starred Aug 2026

Languages

Python

90.4%

Shell

8.7%