High-performance Lemonade alternative for AMD Strix Halo & Radeon — latest MoE models (Ling-3.0-Flash, Qwen 3.8 Flash Next, Ornith 1.5), bleeding-edge RDNA 3.5 Wave64/ROCmFP4 kernels, and silicon-tuned profiles.
See the codeThe high-performance, cutting-edge Lemonade alternative engineered for AMD Strix Halo (Ryzen AI Max APUs / 64GB–128GB UMA) and AMD Radeon GPUs.
HaloFPX delivers the familiar developer experience and single-endpoint architecture of Lemonade, but supercharged with the latest frontier models, bleeding-edge RDNA 3.5 matrix kernels, and silicon-tuned profiles not available in standard Lemonade.
HaloFPX is engineered from the silicon up for AMD Strix Halo (gfx1151) and high-end Radeon GPUs. While maintaining complete CLI and model-cache interoperability with Lemonade, it unlocks performance and architectures standard Lemonade cannot reach:
| Capability | Standard Lemonade | HaloFPX 🚀 |
|---|---|---|
| Next-Gen MoE Models | Generic catalog; misses latest large MoE quants | Day-0 support: Ling-3.0-Flash (124B), Qwen 3.8 Flash Next (125B), DeepSeek V4 Flash (284B), Ornith 1.5 (35B) |
| Compute Kernels | Generic upstream llama.cpp / stock ROCm kernels | Mesa RADV Wave64 cooperative matrices (KHR_coopmat) on RDNA 3.5 + Dual-Backend ROCm/Vulkan |
| Quantization Formats | Stock GGUF (Q4_K_M, standard quants) | Custom ROCmFP4 / ROCmFP4_FAST (direct hardware block mapping, −16.7% size, +13.5% decode speed) |
| Silicon-Tuned Profiles | One-size-fits-all server flags | Hardware-benchmarked run_config: tuned MTP speculation, TurboQuant KV (q8_0), -ctxcp, -cram, zero flag guessing |
| Long Context Scaling | Standard context limits; risks OOM on unified APUs | Validated 262K context (524K tokens across 4 slots) with TurboQuant KV in unified memory |
| CLI & Workflow Parity | lemonade {run, list, chat, backends, delete} | 100% command parity: halofpx run/chat/list/backends/delete, native drop-in, zero retraining |
| Lemonade Cache Sharing | Standalone cache (/var/lib/lemonade/models) | Bi-directional cache sharing: reads Lemonade user models, shares weights, resolves aliases seamlessly |
| Multimodal Vision | Manual projector configuration | Automatic vision discovery & loading (mmproj) verified over standard OpenAI /v1/chat/completions |
Run any model in seconds using familiar Lemonade-compatible commands or launch the OpenAI-compatible single-endpoint server. Silicon-tuned configurations (run_config, TurboQuant KV, Wave64 cooperative matrices) are applied automatically with zero flag guessing.
halofpx run)# List available models and check download status
halofpx list
# Launch an interactive chat session with any model (auto-loads tuned config):
halofpx run Ling-3.0-Flash
# Or run Qwen 3.8 Flash Next or Ornith 1.5:
halofpx run qwen38-flash-next
halofpx run ornith-1.5-35b
halofpx serve)# Start server with an auto-loaded model on http://localhost:8010
halofpx serve -m Ling-3.0-Flash
# Query via standard OpenAI /v1/chat/completions:
curl http://localhost:8010/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Ling-3.0-Flash",
"messages": [{"role": "user", "content": "Explain quantum computing concisely."}]
}'
# Hot-swap to a different model dynamically without restarting the server:
halofpx load qwen38-flash-next
# Check APU memory residency and real-time status:
halofpx status
# Unload from memory when finished:
halofpx unload
run_config — quant preset, KV-cache types, backend preference, MTP on/off — applied automatically on load. No flag archaeology.ornith-1.5-35b) pull and verify their mmproj alongside the weights; image prompts work over the standard OpenAI API.gfx1151, gfx1201, etc.) and applies hardware-specific execution flags.Q4_K_M in decode throughput and size on Strix Halo (gfx1151). See the benchmark table below.HALOFPX_API_KEY (disabled by default for local development)./api/tags, /api/chat, /api/generate and /api/version let existing Ollama clients and tools work against halofpx drop-in./v1/chat/completions (with streaming SSE) plus /api/v1/{pull, load, unload, status, system-info} endpoints on a single port (8010).--profile webui.All models ship pre-quantized with silicon-tuned execution profiles. Measurements taken directly on AMD Ryzen AI Max+ 395 (40 CU Radeon 8060S @ 2.9 GHz, 128 GB 256-bit LPDDR5X, Linux 7.0, Mesa 26.0 RADV Wave64 / Vulkan0):
| Model Name & CLI ID | Category | Quant & Size | Min VRAM | Prefill (pp512) | Decode (Bare / MTP) | Verified Date | HF Repository & Quick Command |
|---|---|---|---|---|---|---|---|
Qwen 3.8 Flash Next 125B MoE ⚡qwen38-flash-next | Next-Gen MoE / Hybrid Attention | ROCmFP4_ple16(87.1 GiB) | 68 GB | 413.09 tok/s (CM1) | 28.96 / 🔥 34.60 tok/s | 2026-09-10 | unsloth/Qwen3.8-Flash-Next-GGUFhalofpx run qwen38-flash-next |
Ling 3.0 Flash 124B MoE ⚡ling3-flash | Frontier 124B MoE / Reasoning | Q4_K_M(71.7 GiB) | 80 GB | 340.50 tok/s | 45.23 / 🔥 55.10 tok/s | 2026-09-02 | inclusionAI/Ling-3.0-flash-GGUFhalofpx run Ling-3.0-Flash |
Nex N2.5 Mini 35B-A3B MoE ⭐nex-n2.5-mini | Agentic Reasoning / Vision | ROCmFP4_LEAN(17.3 GiB) | 22 GB | 1180.20 tok/s | 76.92 tok/s | 2026-08-30 | julianmb/Nex-N2.5-mini-ROCmFP4-GGUFhalofpx run nex-n2.5-mini |
Ornith 1.5 35B-A3B MoE ⭐ornith-1.5-35b | Agentic Coding / Vision MoE | ROCmFP4(18.2 GiB) | 22 GB | 1190.40 tok/s | 76.90 / 🔥 105.60 tok/s | 2026-08-28 | julianmb/Ornith-1.5-35B-A3B-ROCmFP4-GGUFhalofpx run ornith-1.5-35b |
Qwen 3.8 / 27B UltraQualityqwen38-27b | Dense / Reasoning | ROCmFP4_FAST(13.5 GiB) | 16 GB | 385.65 tok/s | 14.02 / 🔥 33.80 tok/s | 2026-08-25 | julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUFhalofpx run qwen38-27b |
NVIDIA Nemotron 3.5 Lightning 30Bnemotron-3.5-30b | High-Speed MoE | ROCmFP4_FAST(14.8 GiB) | 16 GB | 1286.24 tok/s | 52.40 / 🔥 95.20 tok/s | 2026-08-22 | julianmb/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-ROCmFP4-GGUFhalofpx pull nemotron-3.5-30b |
DeepSeek V4 Flash 284B MoEdeepseek-v4-flash | Ultra-Scale MoE | IQ2_XXS(86.7 GiB) | 90 GB | 195.40 tok/s | 22.50 / 32.00 tok/s | 2026-08-20 | julianmb/DeepSeek-V4-Flash-0731-IQ2XXS-STRIXhalofpx pull deepseek-v4-flash |
Laguna S 2.1 StrixKVSpine v4laguna-s21 | General Chat | ROCmFP4_Spine(60.9 GiB) | 64 GB | 390.13 tok/s | 34.05 tok/s | 2026-08-15 | julianmb/Laguna-S-2.1-ROCmFP4-StrixKVSpine-v4halofpx pull laguna-s21 |
Ornith 1.0 35B ROCmFPXornith-35b | Multi-Slot Agent | ROCmFPX_Speed(18.4 GiB) | 22 GB | 1194.19 tok/s | 11.20 / 115.0+ tok/s (16s) | 2026-08-15 | julianmb/Ornith-1.0-35B-ROCmFPX-StrixHalohalofpx pull ornith-35b |
⭐ = vision-capable (mmproj included). ⚡ = cutting-edge large MoE architecture. Full methodology: docs/BENCHMARKS.md.
👉 See Hardware Support & VRAM Sizing Guide (docs/HARDWARE_SUPPORT.md) for memory sizing tables across AMD APUs and discrete GPUs.
HaloFPX provides automated benchmarking runners to measure prompt prefill throughput, token decode speed, TTFT, and MTP draft acceptance rates:
# 1. Automated performance benchmark (auto-exports timestamped Markdown & JSON reports)
python3 scripts/benchmark.py --model qwen38-27b --device Vulkan0
# 2. Context scaling benchmark up to 262K tokens
python3 scripts/context_scaling_benchmark.py --model ornith-1.5-35b
# 3. Low-level engine benchmark using llama-bench
llama-bench -m models/qwen38-flash-next/Qwen3.8-Flash-Next-ROCmFP4-FAST-v2-ple16.gguf -p 512 -n 128 -ngl 99
Benchmark outputs are stored with timestamps under benchmarks/benchmark_YYYYMMDD_HHMMSS.md and benchmarks/benchmark_YYYYMMDD_HHMMSS.json.
Measured on AMD Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, Mesa RADV Wave64) with identical prompts — ROCmFP4-family quants beat stock Q4_K_M on decode speed and model size:
| Model | Stock Q4_K_M | ROCmFP4 / ROCmFP4_FAST | Decode Speedup | Size Savings |
|---|---|---|---|---|
| Qwen 3.8 27B | 15.92 GiB — 12.35 tok/s | 13.55 GiB — 14.02 tok/s | +13.5% | −14.9% |
| Ornith 1.5 35B-A3B | 21.80 GiB — 71.5–71.7 tok/s | 18.16 GiB — 76.9 tok/s | +7.5% | −16.7% |
| Nex N2.5 Mini 35B-A3B | 19.71 GiB — 72.76 tok/s | 17.32 GiB — 76.92 tok/s | +5.7% | −12.1% |
Additional gains over stock GGUF on Strix Halo:
KHR_coopmat) operands — faster prompt evaluation at equal context, without the multi-scale dequantization overhead of Q4_K blocks.n4/p0.0): 33.8 tok/s sustained = 2.40× over stock baseline (12.35 tok/s).q8_0) shrinks memory footprint, enabling 4-slot × 131K contexts (524K total tokens) in unified memory with zero OOM.Full methodology and raw numbers: docs/BENCHMARKS.md.
# Clone the repository
git clone https://github.com/julianmb/halofpx.git
cd halofpx
# Install Python requirements and CLI
pip install -r requirements.txt
pip install -e .
# Set up environment variables for AMD Strix Halo
source ./scripts/setup_env.sh
# Check available hardware acceleration backends (ROCm HIP, Vulkan RADV Wave64)
halofpx backends
# Inspect configuration or synchronize cache with Lemonade daemon
halofpx config show
halofpx config sync-lemonade
# Parallel multi-agent concurrency (4 slots -> ~40.5 tok/s aggregate)
halofpx load qwen38-27b --slots 4 --draft-n 6 --draft-p 0.60
# Single-user interactive chat with burst MTP
halofpx load qwen38-27b --draft-n 5 --draft-p 0.50
# Long context scaling up to 262K tokens with TurboQuant KV in unified memory
halofpx load ornith-1.5-35b --ctx 262144
Connect your local tools to http://localhost:8010/v1:
http://localhost:8010/v1 and API Key to sk-no-key.halofpx as provider in ~/.continue/config.json.http://localhost:8010/v1.👉 See the complete Client Integration Guide (docs/CLIENT_INTEGRATION.md).
Runs only the high-performance HaloFPX server (zero extra RAM overhead for web frontends):
docker compose up -d
http://localhost:8010/v1Runs both the backend server and Open WebUI in a unified stack:
docker compose --profile webui up -d
http://localhost:8010/v1http://localhost:3000docker rundocker run -d -p 8010:8010 \
--device=/dev/dri \
--group-add video --group-add render \
--ipc=host \
-v $(pwd)/models:/app/models \
-v ~/.cache/huggingface/hub:/root/.cache/huggingface/hub \
--name halofpx-server \
ghcr.io/julianmb/halofpx:latest
# Note: For optional ROCm prefill acceleration, also add --device=/dev/kfd and use ghcr.io/julianmb/halofpx:rocm
👉 See the complete Docker Deployment Guide (docs/DOCKER_GUIDE.md) for GPU passthrough prerequisites, container CLI commands, and local builds.
HaloFPX wraps and orchestrates the charlie12345/ROCmFPX engine, compiling directly against pinned builds (e87d53e (213)) or downloading pre-compiled Strix Halo binaries via ./scripts/build_engine.sh --prebuilt.
Apache 2.0 License.
23 followers · starred Aug 2026
Python
90.4%
Shell
8.7%
High-performance Lemonade alternative for AMD Strix Halo & Radeon — latest MoE models (Ling-3.0-Flash, Qwen 3.8 Flash Next, Ornith 1.5), bleeding-edge RDNA 3.5 Wave64/ROCmFP4 kernels, and silicon-tuned profiles.
See the codeThe high-performance, cutting-edge Lemonade alternative engineered for AMD Strix Halo (Ryzen AI Max APUs / 64GB–128GB UMA) and AMD Radeon GPUs.
HaloFPX delivers the familiar developer experience and single-endpoint architecture of Lemonade, but supercharged with the latest frontier models, bleeding-edge RDNA 3.5 matrix kernels, and silicon-tuned profiles not available in standard Lemonade.
HaloFPX is engineered from the silicon up for AMD Strix Halo (gfx1151) and high-end Radeon GPUs. While maintaining complete CLI and model-cache interoperability with Lemonade, it unlocks performance and architectures standard Lemonade cannot reach:
| Capability | Standard Lemonade | HaloFPX 🚀 |
|---|---|---|
| Next-Gen MoE Models | Generic catalog; misses latest large MoE quants | Day-0 support: Ling-3.0-Flash (124B), Qwen 3.8 Flash Next (125B), DeepSeek V4 Flash (284B), Ornith 1.5 (35B) |
| Compute Kernels | Generic upstream llama.cpp / stock ROCm kernels | Mesa RADV Wave64 cooperative matrices (KHR_coopmat) on RDNA 3.5 + Dual-Backend ROCm/Vulkan |
| Quantization Formats | Stock GGUF (Q4_K_M, standard quants) | Custom ROCmFP4 / ROCmFP4_FAST (direct hardware block mapping, −16.7% size, +13.5% decode speed) |
| Silicon-Tuned Profiles | One-size-fits-all server flags | Hardware-benchmarked run_config: tuned MTP speculation, TurboQuant KV (q8_0), -ctxcp, -cram, zero flag guessing |
| Long Context Scaling | Standard context limits; risks OOM on unified APUs | Validated 262K context (524K tokens across 4 slots) with TurboQuant KV in unified memory |
| CLI & Workflow Parity | lemonade {run, list, chat, backends, delete} | 100% command parity: halofpx run/chat/list/backends/delete, native drop-in, zero retraining |
| Lemonade Cache Sharing | Standalone cache (/var/lib/lemonade/models) | Bi-directional cache sharing: reads Lemonade user models, shares weights, resolves aliases seamlessly |
| Multimodal Vision | Manual projector configuration | Automatic vision discovery & loading (mmproj) verified over standard OpenAI /v1/chat/completions |
Run any model in seconds using familiar Lemonade-compatible commands or launch the OpenAI-compatible single-endpoint server. Silicon-tuned configurations (run_config, TurboQuant KV, Wave64 cooperative matrices) are applied automatically with zero flag guessing.
halofpx run)# List available models and check download status
halofpx list
# Launch an interactive chat session with any model (auto-loads tuned config):
halofpx run Ling-3.0-Flash
# Or run Qwen 3.8 Flash Next or Ornith 1.5:
halofpx run qwen38-flash-next
halofpx run ornith-1.5-35b
halofpx serve)# Start server with an auto-loaded model on http://localhost:8010
halofpx serve -m Ling-3.0-Flash
# Query via standard OpenAI /v1/chat/completions:
curl http://localhost:8010/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Ling-3.0-Flash",
"messages": [{"role": "user", "content": "Explain quantum computing concisely."}]
}'
# Hot-swap to a different model dynamically without restarting the server:
halofpx load qwen38-flash-next
# Check APU memory residency and real-time status:
halofpx status
# Unload from memory when finished:
halofpx unload
run_config — quant preset, KV-cache types, backend preference, MTP on/off — applied automatically on load. No flag archaeology.ornith-1.5-35b) pull and verify their mmproj alongside the weights; image prompts work over the standard OpenAI API.gfx1151, gfx1201, etc.) and applies hardware-specific execution flags.Q4_K_M in decode throughput and size on Strix Halo (gfx1151). See the benchmark table below.HALOFPX_API_KEY (disabled by default for local development)./api/tags, /api/chat, /api/generate and /api/version let existing Ollama clients and tools work against halofpx drop-in./v1/chat/completions (with streaming SSE) plus /api/v1/{pull, load, unload, status, system-info} endpoints on a single port (8010).--profile webui.All models ship pre-quantized with silicon-tuned execution profiles. Measurements taken directly on AMD Ryzen AI Max+ 395 (40 CU Radeon 8060S @ 2.9 GHz, 128 GB 256-bit LPDDR5X, Linux 7.0, Mesa 26.0 RADV Wave64 / Vulkan0):
| Model Name & CLI ID | Category | Quant & Size | Min VRAM | Prefill (pp512) | Decode (Bare / MTP) | Verified Date | HF Repository & Quick Command |
|---|---|---|---|---|---|---|---|
Qwen 3.8 Flash Next 125B MoE ⚡qwen38-flash-next | Next-Gen MoE / Hybrid Attention | ROCmFP4_ple16(87.1 GiB) | 68 GB | 413.09 tok/s (CM1) | 28.96 / 🔥 34.60 tok/s | 2026-09-10 | unsloth/Qwen3.8-Flash-Next-GGUFhalofpx run qwen38-flash-next |
Ling 3.0 Flash 124B MoE ⚡ling3-flash | Frontier 124B MoE / Reasoning | Q4_K_M(71.7 GiB) | 80 GB | 340.50 tok/s | 45.23 / 🔥 55.10 tok/s | 2026-09-02 | inclusionAI/Ling-3.0-flash-GGUFhalofpx run Ling-3.0-Flash |
Nex N2.5 Mini 35B-A3B MoE ⭐nex-n2.5-mini | Agentic Reasoning / Vision | ROCmFP4_LEAN(17.3 GiB) | 22 GB | 1180.20 tok/s | 76.92 tok/s | 2026-08-30 | julianmb/Nex-N2.5-mini-ROCmFP4-GGUFhalofpx run nex-n2.5-mini |
Ornith 1.5 35B-A3B MoE ⭐ornith-1.5-35b | Agentic Coding / Vision MoE | ROCmFP4(18.2 GiB) | 22 GB | 1190.40 tok/s | 76.90 / 🔥 105.60 tok/s | 2026-08-28 | julianmb/Ornith-1.5-35B-A3B-ROCmFP4-GGUFhalofpx run ornith-1.5-35b |
Qwen 3.8 / 27B UltraQualityqwen38-27b | Dense / Reasoning | ROCmFP4_FAST(13.5 GiB) | 16 GB | 385.65 tok/s | 14.02 / 🔥 33.80 tok/s | 2026-08-25 | julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUFhalofpx run qwen38-27b |
NVIDIA Nemotron 3.5 Lightning 30Bnemotron-3.5-30b | High-Speed MoE | ROCmFP4_FAST(14.8 GiB) | 16 GB | 1286.24 tok/s | 52.40 / 🔥 95.20 tok/s | 2026-08-22 | julianmb/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-ROCmFP4-GGUFhalofpx pull nemotron-3.5-30b |
DeepSeek V4 Flash 284B MoEdeepseek-v4-flash | Ultra-Scale MoE | IQ2_XXS(86.7 GiB) | 90 GB | 195.40 tok/s | 22.50 / 32.00 tok/s | 2026-08-20 | julianmb/DeepSeek-V4-Flash-0731-IQ2XXS-STRIXhalofpx pull deepseek-v4-flash |
Laguna S 2.1 StrixKVSpine v4laguna-s21 | General Chat | ROCmFP4_Spine(60.9 GiB) | 64 GB | 390.13 tok/s | 34.05 tok/s | 2026-08-15 | julianmb/Laguna-S-2.1-ROCmFP4-StrixKVSpine-v4halofpx pull laguna-s21 |
Ornith 1.0 35B ROCmFPXornith-35b | Multi-Slot Agent | ROCmFPX_Speed(18.4 GiB) | 22 GB | 1194.19 tok/s | 11.20 / 115.0+ tok/s (16s) | 2026-08-15 | julianmb/Ornith-1.0-35B-ROCmFPX-StrixHalohalofpx pull ornith-35b |
⭐ = vision-capable (mmproj included). ⚡ = cutting-edge large MoE architecture. Full methodology: docs/BENCHMARKS.md.
👉 See Hardware Support & VRAM Sizing Guide (docs/HARDWARE_SUPPORT.md) for memory sizing tables across AMD APUs and discrete GPUs.
HaloFPX provides automated benchmarking runners to measure prompt prefill throughput, token decode speed, TTFT, and MTP draft acceptance rates:
# 1. Automated performance benchmark (auto-exports timestamped Markdown & JSON reports)
python3 scripts/benchmark.py --model qwen38-27b --device Vulkan0
# 2. Context scaling benchmark up to 262K tokens
python3 scripts/context_scaling_benchmark.py --model ornith-1.5-35b
# 3. Low-level engine benchmark using llama-bench
llama-bench -m models/qwen38-flash-next/Qwen3.8-Flash-Next-ROCmFP4-FAST-v2-ple16.gguf -p 512 -n 128 -ngl 99
Benchmark outputs are stored with timestamps under benchmarks/benchmark_YYYYMMDD_HHMMSS.md and benchmarks/benchmark_YYYYMMDD_HHMMSS.json.
Measured on AMD Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, Mesa RADV Wave64) with identical prompts — ROCmFP4-family quants beat stock Q4_K_M on decode speed and model size:
| Model | Stock Q4_K_M | ROCmFP4 / ROCmFP4_FAST | Decode Speedup | Size Savings |
|---|---|---|---|---|
| Qwen 3.8 27B | 15.92 GiB — 12.35 tok/s | 13.55 GiB — 14.02 tok/s | +13.5% | −14.9% |
| Ornith 1.5 35B-A3B | 21.80 GiB — 71.5–71.7 tok/s | 18.16 GiB — 76.9 tok/s | +7.5% | −16.7% |
| Nex N2.5 Mini 35B-A3B | 19.71 GiB — 72.76 tok/s | 17.32 GiB — 76.92 tok/s | +5.7% | −12.1% |
Additional gains over stock GGUF on Strix Halo:
KHR_coopmat) operands — faster prompt evaluation at equal context, without the multi-scale dequantization overhead of Q4_K blocks.n4/p0.0): 33.8 tok/s sustained = 2.40× over stock baseline (12.35 tok/s).q8_0) shrinks memory footprint, enabling 4-slot × 131K contexts (524K total tokens) in unified memory with zero OOM.Full methodology and raw numbers: docs/BENCHMARKS.md.
# Clone the repository
git clone https://github.com/julianmb/halofpx.git
cd halofpx
# Install Python requirements and CLI
pip install -r requirements.txt
pip install -e .
# Set up environment variables for AMD Strix Halo
source ./scripts/setup_env.sh
# Check available hardware acceleration backends (ROCm HIP, Vulkan RADV Wave64)
halofpx backends
# Inspect configuration or synchronize cache with Lemonade daemon
halofpx config show
halofpx config sync-lemonade
# Parallel multi-agent concurrency (4 slots -> ~40.5 tok/s aggregate)
halofpx load qwen38-27b --slots 4 --draft-n 6 --draft-p 0.60
# Single-user interactive chat with burst MTP
halofpx load qwen38-27b --draft-n 5 --draft-p 0.50
# Long context scaling up to 262K tokens with TurboQuant KV in unified memory
halofpx load ornith-1.5-35b --ctx 262144
Connect your local tools to http://localhost:8010/v1:
http://localhost:8010/v1 and API Key to sk-no-key.halofpx as provider in ~/.continue/config.json.http://localhost:8010/v1.👉 See the complete Client Integration Guide (docs/CLIENT_INTEGRATION.md).
Runs only the high-performance HaloFPX server (zero extra RAM overhead for web frontends):
docker compose up -d
http://localhost:8010/v1Runs both the backend server and Open WebUI in a unified stack:
docker compose --profile webui up -d
http://localhost:8010/v1http://localhost:3000docker rundocker run -d -p 8010:8010 \
--device=/dev/dri \
--group-add video --group-add render \
--ipc=host \
-v $(pwd)/models:/app/models \
-v ~/.cache/huggingface/hub:/root/.cache/huggingface/hub \
--name halofpx-server \
ghcr.io/julianmb/halofpx:latest
# Note: For optional ROCm prefill acceleration, also add --device=/dev/kfd and use ghcr.io/julianmb/halofpx:rocm
👉 See the complete Docker Deployment Guide (docs/DOCKER_GUIDE.md) for GPU passthrough prerequisites, container CLI commands, and local builds.
HaloFPX wraps and orchestrates the charlie12345/ROCmFPX engine, compiling directly against pinned builds (e87d53e (213)) or downloading pre-compiled Strix Halo binaries via ./scripts/build_engine.sh --prebuilt.
Apache 2.0 License.
23 followers · starred Aug 2026
Python
90.4%
Shell
8.7%