The fastest way to run an LLM on your machine: Flux measures your GPUs, CPU, and storage once, splits the model across them, and serves it OpenAI-style in seconds.
Rust
2
57 commits
updated Oct 7, 2026
Squeezes every last token out of your machine.
Flux is a measured execution planner and runtime for LLM inference. It measures your GPUs, CPU, and storage, searches placements of a model's layers and mixture-of-experts weights across them, times the fastest candidates on real prompts, and saves the winner as an immutable plan. When a model outgrows RAM, you can let the plan stream the remaining weights from an NVMe drive. flux serve then runs that plan behind an OpenAI-compatible API.
Cost models only rank and prune candidates; Flux always saves the plan that measured fastest. It runs on a pinned, patched build of llama.cpp and can also plan for llama-server or any OpenAI-compatible engine you configure.
Planning is model compilation. A compiler fits a program to the processor that will run it; flux plan fits a model to the machine that will serve it. It leaves the weights untouched and decides where each layer and expert runs. Like a binary, the plan is built once and run many times, and a new GPU, driver, or backend build calls for a new one.
| Flux | Hand-tuned single-model engines | |
|---|---|---|
| Models | Any of the 148 architectures the pinned llama.cpp implements | The one model they were tuned for |
| Ready to serve | In seconds: weights are memory-mapped and page-locked in the background; 15 s for Qwen3.8-Flash-Next on an RTX 2060 + RTX 3060 | Minutes: the experts are read into RAM before the first request, longest on a cold start |
| CPU's share of the work | Measured on your machine: RAM bandwidth per thread count and the model's own kernels | Fixed constants; an optional calibration run adjusts a few of them |
| Layer and expert placement | Searched across the CPU and GPUs; the finalists are timed on real prompts | Fixed rules decide each token's split |
| Several GPUs | Splits layers across GPUs and caches experts on them, in the same plan | A layer split or extra expert caches, one at a time |
| Engine | Compares its native engine, llama-server, and any engine you register, then keeps the fastest | One engine |
| Models bigger than RAM | Measures the drive that holds the model and predicts the decode limit of reading weights from it; opt in with --allow-storage-streaming | A low-RAM mode chosen from the amount of RAM, without testing the drive |
| Changes while serving | Watches decode speed and can replan in place | Calibration runs only when you start it |
| Component | Requirement |
|---|---|
| OS | Linux x86_64 |
| Rust | 1.85 or newer |
| Build tools | CMake, a C++17 compiler, and git |
| GPU | An NVIDIA GPU and the CUDA toolkit for the default build; set GGML_CUDA=OFF to build for CPU only |
| Python | Python 3 with torch, numpy, and transformers, needed only by flux convert |
Flux runs any GGUF file whose general.architecture the pinned llama.cpp implements: 148 architectures at the current pin, including the Llama, Qwen3, Qwen3-Next, Gemma 3, DeepSeek, gpt-oss, GLM, and MiniCPM families. The list moves with backend.pin, because Flux reads it from llama.cpp's source at build time. Run flux inspect <model> to check a file before planning.
| Model format | How Flux runs it |
|---|---|
| GGUF with a supported architecture | The native engine or llama-server |
| Hugging Face safetensors | Convert it to GGUF with flux convert, or register an engine for it in flux.toml |
| EXL3, GPTQ, AWQ, or FP8 | Register an engine that serves the format, such as TabbyAPI or vLLM, in flux.toml |
Expert caching and the second-GPU tier apply only to mixture-of-experts models; dense models get measured layer placement. A plan cannot exceed the model's trained context. Flux has been tested end to end on Qwen3.8-Flash-Next, Qwen3.6-35B-A3B, and MiniCPM5.
The CPU is a device in every plan, not a fallback for what the GPUs cannot hold. Flux times its RAM bandwidth and the model's own kernels on it, places whole layers or a layer's experts on it, and fills spare GPU memory with the experts that save the most time per byte. The finalists run on real prompts, CPU work included, so the CPU's share and its thread count are measured rather than assumed.
While serving, the CPU computes its experts while the GPU computes the rest of the layer. Prompt chunks of 32 tokens or more copy those experts to the GPU instead, where the larger batch runs faster.
flux plan gives each model its trained context: 262,144 tokens for Qwen3.8-Flash-Next. The KV cache commits memory as a conversation grows, not up front:
flux serve gives KV pages the RAM available when it starts, minus the weights it pins and the host reserve (host_reserve_percent, 10%). flux plan caps the context only when that RAM cannot hold the trained one.A request that does not fit fails at admission with HTTP 400 and OpenAI's context_length_exceeded error, which clients such as Qwen Code answer by compressing their history; a running step never fails for lack of memory. Each request must leave room for min(max_tokens, serve.min_reply) reply tokens (min_reply defaults to 4,096). /v1/models reports the limit as context_length, max_model_len and meta.n_ctx.
When a prompt replaces a conversation, the worker parks the old one's state in RAM, so an agent that returns to it after a side request skips recomputing the prefix. flux show <plan> prints the page budgets, and /flux/stats reports memory by use.
git clone --recurse-submodules https://github.com/cyqlelabs/flux.git
cd flux
# Optional: install Python dependencies if you intend to run `flux convert`
pip install -r requirements.txt
scripts/build-backend.sh
cargo build --release -p flux-cli -p flux-worker
build-backend.sh ensures third_party/llama.cpp matches backend.pin, applies the patches in patches/llama.cpp/, and builds llama.cpp for CUDA compute capabilities 7.5 and 8.6 (RTX 20 and RTX 30 series). For other GPUs set CUDA_ARCHS, for example CUDA_ARCHS=89 scripts/build-backend.sh for the RTX 40 series. The flux and flux-worker binaries land in target/release/.
scripts/package.sh builds a relocatable tarball in dist/ that bundles both binaries, the llama.cpp libraries, and a starter flux.toml.
export PATH="$PWD/target/release:$PATH"
flux plan path/to/model.gguf
flux serve <plan-id>
The first flux plan takes several minutes; flux serve then starts in seconds. Planning probes the hardware, downloads the WikiText-2 prompt corpus, and times the finalist placements within a 600-second tuning budget (--budget-s changes it). It validates the top two on held-out prompts and on a built-in set of coding-agent conversations, then saves the winner. Later runs for the same model, machine, and workload reuse the saved plan; pass --replan to measure again. flux plans lists saved plans, and flux serve accepts any unique prefix of a plan id.
The server listens on 127.0.0.1:8090:
curl http://127.0.0.1:8090/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 64}'
| Command | Purpose |
|---|---|
flux inspect <path> | Identify a GGUF file or Hugging Face checkpoint: metadata, shards, hashes, and compatibility |
flux fetch <repo> --include <glob> | Download and verify Hugging Face files at a pinned revision |
flux convert <dir> | Convert a Hugging Face checkpoint to GGUF with the pinned converter |
flux corpus | Download the WikiText-2 prompt corpus |
flux probe | Measure copy curves, contention, kernel shapes, and CPU and storage bandwidth |
flux plan <model> | Find, measure, and save the fastest validated plan |
flux plans, flux show <id> | List saved plans, or show one with its decisions and measurements |
flux trace <plan> | Attribute decode-step time to devices and operations |
flux serve <plan> | Serve a plan over the OpenAI-compatible API |
flux bench <suite> | Benchmark speed, quality, conformance, and resilience |
Run flux <command> --help for every flag.
flux plan flags| Flag | Default | Effect |
|---|---|---|
--ctx | The model's trained context, capped only when RAM cannot hold its KV pages | Tokens per sequence the plan must hold, prompt plus output |
--concurrency | 1 | Concurrent sequences the plan must hold |
--serving-p95-ms | off | Optimize aggregate tokens per second under this p95 per-token latency |
--engines | native,llama-server | Engines to compare, including any named in flux.toml |
--kv | f16 | KV cache type; any other type is labelled as a separate quality profile |
--speculation | off | Also measure drafting with --draft-model; models with next-token heads are measured without it |
--draft-model, --heads | none | Draft with a separate model, or graft next-token (MTP) heads onto the model |
--allow-storage-streaming | off | Let a plan read the weights that do not fit in RAM from the drive; without it, Flux rejects such placements and reports their decode limit |
--replan, --reprobe | off | Ignore the saved plan or the saved probe report |
flux bench suites| Suite | Purpose |
|---|---|
run | Race Flux against the strongest tuned baseline in paired, randomized trials |
quality | Compare a plan with a reference placement by KL divergence, top-1 agreement, and perplexity |
conformance | Check that the native engine and llama-server agree on templates, special tokens, sampling, and stops |
soak | Hit flux serve with overload, cancellations, and injected faults |
arrivals | Replay Poisson arrivals and a long-running chat on Flux and on llama.cpp auto-fit |
ablate | Compare the full plan with variants that each remove one optimization |
show, matrix | Summarize one suite, or print a pass/fail matrix over every saved plan |
flux serve --host and --port override the default address.
| Route | Purpose |
|---|---|
GET /v1/models | List the served model |
POST /v1/chat/completions | Create chat completions, streamed or not; replies are split into reasoning_content, content, and tool_calls as llama-server does |
POST /v1/completions | Create text completions, streamed or not |
GET /health | Return 200 when the worker is ready and admission is open, otherwise 503 |
GET /live | Return 200 while the HTTP server is alive |
GET /flux/plan | Return the plan being served |
GET /flux/stats | Report admission, decode drift, and worker counters |
POST /flux/replan | Drain, plan again, and switch; restore the old plan if the new one fails to load |
POST /flux/tokenize | Tokenize text with the model's vocabulary |
POST /flux/admission | Adjust the host-memory reserve with {"host_reserve_bytes": N} |
GET /flux/requests/{id} | Return a journaled request's status and text |
GET /flux/requests/{id}/stream?after=N | Read retained output starting at event index N (default 0) |
A request's id is its x-request-id header or, when that is absent, the id in the response. Output is journaled in memory for 600 seconds after a request completes (earlier under memory pressure), so a client that drops its connection can resume the stream: each event carries an index, and after=N+1 continues after event N. An expired request returns 404.
A worker that fails while idle restarts on its own. One that fails mid-request ends that request with an explicit error; start a new request to generate again.
Flux reads $FLUX_CONFIG, then ~/.config/flux/flux.toml. Every field has a default, so the file only lists overrides:
models_dir = "/data/models" # where flux fetch writes
cache_dir = "/data/flux-cache" # hash index, prepared artifacts, corpora
host_reserve_percent = 10 # percentage of total RAM kept free
[plan]
tuning_budget_s = 600
[serve]
port = 8090
prefill_timeout_s = 300
decode_timeout_s = 120
client_write_timeout_s = 30
replan_timeout_s = 3600
journal_max_bytes = 268435456
min_reply = 4096
[serve.worker_timeouts]
hello_s = 30
load_s = 1800
rpc_s = 120
write_s = 30
# Any OpenAI-compatible engine; {port}, {model} and {ctx} are substituted.
[engines.my-engine]
command = ["my-engine", "serve"]
args = ["--port", "{port}", "--model", "{model}", "--ctx-size", "{ctx}"]
architectures = ["qwen3moe"]
A registered engine must stream per-token logprobs and completion usage, which Flux needs to time it.
Plans, probe reports, logs, and benchmark results live in $XDG_DATA_HOME/flux, which defaults to ~/.local/share/flux. flux serve writes worker output to logs/serve-worker.log there. All defaults are in crates/flux-core/src/config.rs.
Planning and serving share the host-memory reserve. POST /flux/admission overrides it until the server restarts, and /flux/stats reports the current value.
| Variable | Effect |
|---|---|
FLUX_CONFIG | Path of flux.toml |
FLUX_LOG | Tracing filter for flux (default info) |
FLUX_NATIVE_LOG | Minimum ggml log level the bridge prints: debug, info, warn (default) or error |
FLUX_WORKER | Path of the worker binary (default: flux-worker next to flux) |
FLUX_MOE_HOST_PROFILE=1 | Time host (CPU) expert work per step |
FLUX_CUDA_OP_PROFILE=1 | Time each CUDA operation |
FLUX_CUDA_OP_PROFILE=2 | Time every CUDA operation without truncating the op table |
FLUX_MOE_HOST_SYNC=1 | Stop overlapping CPU experts with the GPU |
FLUX_MOE_CACHE_FREEZE=1 | Stop the GPU expert cache from adapting |
FLUX_MOE_TIER_WAIT=1 | Wait for a second GPU's expert tier instead of recomputing its late work on the CPU, so greedy outputs repeat exactly |
FLUX_MOE_ROUTE_DUMP=<path> | Write each MoE layer's expert selection counts to this file as 16-bit (layer, expert, count) records |
GGML_OP_OFFLOAD_MIN_BATCH | Batch size at which ops on host weights move to a GPU (default 32) |
Throughput depends on everything else the machine is doing: a busy browser can halve decode speed, and a cold page cache slows the first prompts. Plan and benchmark on an idle machine.
Only flux-worker links llama.cpp, so a native crash never takes down flux. Planning and probing run the worker as one-shot jobs. Serving and plan validation keep a flux-worker serve process alive and talk to it over a versioned JSON-lines protocol.
| Crate | Role |
|---|---|
flux-cli | The flux binary |
flux-ingest | GGUF and Hugging Face manifests, hashing, fetching, and conversion |
flux-probe | Hardware measurements, saved per topology |
flux-plan | Placement search, expert cache sizing, finalist measurement, and the plan store |
flux-serve | The OpenAI API, admission control, request journal, and drift-triggered replanning |
flux-bench | Paired trials, quality, conformance, and soak tests |
flux-core | Shared types: plan, config, worker protocol, and supervisor |
flux-native | C ABI bridge to llama.cpp that exchanges JSON for complex values |
flux-worker | The flux-worker binary |
A saved plan never changes. It is keyed by the model's file hashes, the hardware, the backend build, the driver, and the planning flags and settings, so a change to any of them needs a new plan.
If a GPU has less free memory when flux serve starts than the plan was measured to need, the server drops that GPU's least-routed cached experts and logs a warning. When the cached experts cannot cover the shortfall, it refuses the plan; free the memory or run flux plan <model> --replan.
cargo test --release # whole workspace
cargo test --release -p flux-plan experts # one crate, filtered by test name
cargo fmt
Every crate needs the submodule checked out. Crates that link flux-native also need the backend built.
Flux's backend changes live in patches/llama.cpp/ as git diff output against backend.pin; the submodule's working tree is dirty by design. Never commit inside the submodule. Edit third_party/llama.cpp in place, rebuild with scripts/build-backend.sh, then regenerate both patches:
git -C third_party/llama.cpp diff -- ggml/src/ggml-cpu/arch-fallback.h ggml/src/ggml-cpu/arch/x86/quants.c \
> patches/llama.cpp/0002-flux-q2_0-avx2.patch
git -C third_party/llama.cpp diff -- . ':!ggml/src/ggml-cpu/arch-fallback.h' ':!ggml/src/ggml-cpu/arch/x86/quants.c' \
> patches/llama.cpp/0001-flux-backend-extensions.patch
Every plan's key includes a hash of the patches, the built libraries, and the flux-native and flux-worker sources, so a change to any of them invalidates all saved plans and probe reports. Run flux plan again afterwards. The default build compiles only the CPU and CUDA backends, so patch edits to Metal, Vulkan, SYCL, and other backends go unchecked.
MIT
The fastest way to run an LLM on your machine: Flux measures your GPUs, CPU, and storage once, splits the model across them, and serves it OpenAI-style in seconds.
Rust
2
57 commits
updated Oct 7, 2026
Squeezes every last token out of your machine.
Flux is a measured execution planner and runtime for LLM inference. It measures your GPUs, CPU, and storage, searches placements of a model's layers and mixture-of-experts weights across them, times the fastest candidates on real prompts, and saves the winner as an immutable plan. When a model outgrows RAM, you can let the plan stream the remaining weights from an NVMe drive. flux serve then runs that plan behind an OpenAI-compatible API.
Cost models only rank and prune candidates; Flux always saves the plan that measured fastest. It runs on a pinned, patched build of llama.cpp and can also plan for llama-server or any OpenAI-compatible engine you configure.
Planning is model compilation. A compiler fits a program to the processor that will run it; flux plan fits a model to the machine that will serve it. It leaves the weights untouched and decides where each layer and expert runs. Like a binary, the plan is built once and run many times, and a new GPU, driver, or backend build calls for a new one.
| Flux | Hand-tuned single-model engines | |
|---|---|---|
| Models | Any of the 148 architectures the pinned llama.cpp implements | The one model they were tuned for |
| Ready to serve | In seconds: weights are memory-mapped and page-locked in the background; 15 s for Qwen3.8-Flash-Next on an RTX 2060 + RTX 3060 | Minutes: the experts are read into RAM before the first request, longest on a cold start |
| CPU's share of the work | Measured on your machine: RAM bandwidth per thread count and the model's own kernels | Fixed constants; an optional calibration run adjusts a few of them |
| Layer and expert placement | Searched across the CPU and GPUs; the finalists are timed on real prompts | Fixed rules decide each token's split |
| Several GPUs | Splits layers across GPUs and caches experts on them, in the same plan | A layer split or extra expert caches, one at a time |
| Engine | Compares its native engine, llama-server, and any engine you register, then keeps the fastest | One engine |
| Models bigger than RAM | Measures the drive that holds the model and predicts the decode limit of reading weights from it; opt in with --allow-storage-streaming | A low-RAM mode chosen from the amount of RAM, without testing the drive |
| Changes while serving | Watches decode speed and can replan in place | Calibration runs only when you start it |
| Component | Requirement |
|---|---|
| OS | Linux x86_64 |
| Rust | 1.85 or newer |
| Build tools | CMake, a C++17 compiler, and git |
| GPU | An NVIDIA GPU and the CUDA toolkit for the default build; set GGML_CUDA=OFF to build for CPU only |
| Python | Python 3 with torch, numpy, and transformers, needed only by flux convert |
Flux runs any GGUF file whose general.architecture the pinned llama.cpp implements: 148 architectures at the current pin, including the Llama, Qwen3, Qwen3-Next, Gemma 3, DeepSeek, gpt-oss, GLM, and MiniCPM families. The list moves with backend.pin, because Flux reads it from llama.cpp's source at build time. Run flux inspect <model> to check a file before planning.
| Model format | How Flux runs it |
|---|---|
| GGUF with a supported architecture | The native engine or llama-server |
| Hugging Face safetensors | Convert it to GGUF with flux convert, or register an engine for it in flux.toml |
| EXL3, GPTQ, AWQ, or FP8 | Register an engine that serves the format, such as TabbyAPI or vLLM, in flux.toml |
Expert caching and the second-GPU tier apply only to mixture-of-experts models; dense models get measured layer placement. A plan cannot exceed the model's trained context. Flux has been tested end to end on Qwen3.8-Flash-Next, Qwen3.6-35B-A3B, and MiniCPM5.
The CPU is a device in every plan, not a fallback for what the GPUs cannot hold. Flux times its RAM bandwidth and the model's own kernels on it, places whole layers or a layer's experts on it, and fills spare GPU memory with the experts that save the most time per byte. The finalists run on real prompts, CPU work included, so the CPU's share and its thread count are measured rather than assumed.
While serving, the CPU computes its experts while the GPU computes the rest of the layer. Prompt chunks of 32 tokens or more copy those experts to the GPU instead, where the larger batch runs faster.
flux plan gives each model its trained context: 262,144 tokens for Qwen3.8-Flash-Next. The KV cache commits memory as a conversation grows, not up front:
flux serve gives KV pages the RAM available when it starts, minus the weights it pins and the host reserve (host_reserve_percent, 10%). flux plan caps the context only when that RAM cannot hold the trained one.A request that does not fit fails at admission with HTTP 400 and OpenAI's context_length_exceeded error, which clients such as Qwen Code answer by compressing their history; a running step never fails for lack of memory. Each request must leave room for min(max_tokens, serve.min_reply) reply tokens (min_reply defaults to 4,096). /v1/models reports the limit as context_length, max_model_len and meta.n_ctx.
When a prompt replaces a conversation, the worker parks the old one's state in RAM, so an agent that returns to it after a side request skips recomputing the prefix. flux show <plan> prints the page budgets, and /flux/stats reports memory by use.
git clone --recurse-submodules https://github.com/cyqlelabs/flux.git
cd flux
# Optional: install Python dependencies if you intend to run `flux convert`
pip install -r requirements.txt
scripts/build-backend.sh
cargo build --release -p flux-cli -p flux-worker
build-backend.sh ensures third_party/llama.cpp matches backend.pin, applies the patches in patches/llama.cpp/, and builds llama.cpp for CUDA compute capabilities 7.5 and 8.6 (RTX 20 and RTX 30 series). For other GPUs set CUDA_ARCHS, for example CUDA_ARCHS=89 scripts/build-backend.sh for the RTX 40 series. The flux and flux-worker binaries land in target/release/.
scripts/package.sh builds a relocatable tarball in dist/ that bundles both binaries, the llama.cpp libraries, and a starter flux.toml.
export PATH="$PWD/target/release:$PATH"
flux plan path/to/model.gguf
flux serve <plan-id>
The first flux plan takes several minutes; flux serve then starts in seconds. Planning probes the hardware, downloads the WikiText-2 prompt corpus, and times the finalist placements within a 600-second tuning budget (--budget-s changes it). It validates the top two on held-out prompts and on a built-in set of coding-agent conversations, then saves the winner. Later runs for the same model, machine, and workload reuse the saved plan; pass --replan to measure again. flux plans lists saved plans, and flux serve accepts any unique prefix of a plan id.
The server listens on 127.0.0.1:8090:
curl http://127.0.0.1:8090/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 64}'
| Command | Purpose |
|---|---|
flux inspect <path> | Identify a GGUF file or Hugging Face checkpoint: metadata, shards, hashes, and compatibility |
flux fetch <repo> --include <glob> | Download and verify Hugging Face files at a pinned revision |
flux convert <dir> | Convert a Hugging Face checkpoint to GGUF with the pinned converter |
flux corpus | Download the WikiText-2 prompt corpus |
flux probe | Measure copy curves, contention, kernel shapes, and CPU and storage bandwidth |
flux plan <model> | Find, measure, and save the fastest validated plan |
flux plans, flux show <id> | List saved plans, or show one with its decisions and measurements |
flux trace <plan> | Attribute decode-step time to devices and operations |
flux serve <plan> | Serve a plan over the OpenAI-compatible API |
flux bench <suite> | Benchmark speed, quality, conformance, and resilience |
Run flux <command> --help for every flag.
flux plan flags| Flag | Default | Effect |
|---|---|---|
--ctx | The model's trained context, capped only when RAM cannot hold its KV pages | Tokens per sequence the plan must hold, prompt plus output |
--concurrency | 1 | Concurrent sequences the plan must hold |
--serving-p95-ms | off | Optimize aggregate tokens per second under this p95 per-token latency |
--engines | native,llama-server | Engines to compare, including any named in flux.toml |
--kv | f16 | KV cache type; any other type is labelled as a separate quality profile |
--speculation | off | Also measure drafting with --draft-model; models with next-token heads are measured without it |
--draft-model, --heads | none | Draft with a separate model, or graft next-token (MTP) heads onto the model |
--allow-storage-streaming | off | Let a plan read the weights that do not fit in RAM from the drive; without it, Flux rejects such placements and reports their decode limit |
--replan, --reprobe | off | Ignore the saved plan or the saved probe report |
flux bench suites| Suite | Purpose |
|---|---|
run | Race Flux against the strongest tuned baseline in paired, randomized trials |
quality | Compare a plan with a reference placement by KL divergence, top-1 agreement, and perplexity |
conformance | Check that the native engine and llama-server agree on templates, special tokens, sampling, and stops |
soak | Hit flux serve with overload, cancellations, and injected faults |
arrivals | Replay Poisson arrivals and a long-running chat on Flux and on llama.cpp auto-fit |
ablate | Compare the full plan with variants that each remove one optimization |
show, matrix | Summarize one suite, or print a pass/fail matrix over every saved plan |
flux serve --host and --port override the default address.
| Route | Purpose |
|---|---|
GET /v1/models | List the served model |
POST /v1/chat/completions | Create chat completions, streamed or not; replies are split into reasoning_content, content, and tool_calls as llama-server does |
POST /v1/completions | Create text completions, streamed or not |
GET /health | Return 200 when the worker is ready and admission is open, otherwise 503 |
GET /live | Return 200 while the HTTP server is alive |
GET /flux/plan | Return the plan being served |
GET /flux/stats | Report admission, decode drift, and worker counters |
POST /flux/replan | Drain, plan again, and switch; restore the old plan if the new one fails to load |
POST /flux/tokenize | Tokenize text with the model's vocabulary |
POST /flux/admission | Adjust the host-memory reserve with {"host_reserve_bytes": N} |
GET /flux/requests/{id} | Return a journaled request's status and text |
GET /flux/requests/{id}/stream?after=N | Read retained output starting at event index N (default 0) |
A request's id is its x-request-id header or, when that is absent, the id in the response. Output is journaled in memory for 600 seconds after a request completes (earlier under memory pressure), so a client that drops its connection can resume the stream: each event carries an index, and after=N+1 continues after event N. An expired request returns 404.
A worker that fails while idle restarts on its own. One that fails mid-request ends that request with an explicit error; start a new request to generate again.
Flux reads $FLUX_CONFIG, then ~/.config/flux/flux.toml. Every field has a default, so the file only lists overrides:
models_dir = "/data/models" # where flux fetch writes
cache_dir = "/data/flux-cache" # hash index, prepared artifacts, corpora
host_reserve_percent = 10 # percentage of total RAM kept free
[plan]
tuning_budget_s = 600
[serve]
port = 8090
prefill_timeout_s = 300
decode_timeout_s = 120
client_write_timeout_s = 30
replan_timeout_s = 3600
journal_max_bytes = 268435456
min_reply = 4096
[serve.worker_timeouts]
hello_s = 30
load_s = 1800
rpc_s = 120
write_s = 30
# Any OpenAI-compatible engine; {port}, {model} and {ctx} are substituted.
[engines.my-engine]
command = ["my-engine", "serve"]
args = ["--port", "{port}", "--model", "{model}", "--ctx-size", "{ctx}"]
architectures = ["qwen3moe"]
A registered engine must stream per-token logprobs and completion usage, which Flux needs to time it.
Plans, probe reports, logs, and benchmark results live in $XDG_DATA_HOME/flux, which defaults to ~/.local/share/flux. flux serve writes worker output to logs/serve-worker.log there. All defaults are in crates/flux-core/src/config.rs.
Planning and serving share the host-memory reserve. POST /flux/admission overrides it until the server restarts, and /flux/stats reports the current value.
| Variable | Effect |
|---|---|
FLUX_CONFIG | Path of flux.toml |
FLUX_LOG | Tracing filter for flux (default info) |
FLUX_NATIVE_LOG | Minimum ggml log level the bridge prints: debug, info, warn (default) or error |
FLUX_WORKER | Path of the worker binary (default: flux-worker next to flux) |
FLUX_MOE_HOST_PROFILE=1 | Time host (CPU) expert work per step |
FLUX_CUDA_OP_PROFILE=1 | Time each CUDA operation |
FLUX_CUDA_OP_PROFILE=2 | Time every CUDA operation without truncating the op table |
FLUX_MOE_HOST_SYNC=1 | Stop overlapping CPU experts with the GPU |
FLUX_MOE_CACHE_FREEZE=1 | Stop the GPU expert cache from adapting |
FLUX_MOE_TIER_WAIT=1 | Wait for a second GPU's expert tier instead of recomputing its late work on the CPU, so greedy outputs repeat exactly |
FLUX_MOE_ROUTE_DUMP=<path> | Write each MoE layer's expert selection counts to this file as 16-bit (layer, expert, count) records |
GGML_OP_OFFLOAD_MIN_BATCH | Batch size at which ops on host weights move to a GPU (default 32) |
Throughput depends on everything else the machine is doing: a busy browser can halve decode speed, and a cold page cache slows the first prompts. Plan and benchmark on an idle machine.
Only flux-worker links llama.cpp, so a native crash never takes down flux. Planning and probing run the worker as one-shot jobs. Serving and plan validation keep a flux-worker serve process alive and talk to it over a versioned JSON-lines protocol.
| Crate | Role |
|---|---|
flux-cli | The flux binary |
flux-ingest | GGUF and Hugging Face manifests, hashing, fetching, and conversion |
flux-probe | Hardware measurements, saved per topology |
flux-plan | Placement search, expert cache sizing, finalist measurement, and the plan store |
flux-serve | The OpenAI API, admission control, request journal, and drift-triggered replanning |
flux-bench | Paired trials, quality, conformance, and soak tests |
flux-core | Shared types: plan, config, worker protocol, and supervisor |
flux-native | C ABI bridge to llama.cpp that exchanges JSON for complex values |
flux-worker | The flux-worker binary |
A saved plan never changes. It is keyed by the model's file hashes, the hardware, the backend build, the driver, and the planning flags and settings, so a change to any of them needs a new plan.
If a GPU has less free memory when flux serve starts than the plan was measured to need, the server drops that GPU's least-routed cached experts and logs a warning. When the cached experts cannot cover the shortfall, it refuses the plan; free the memory or run flux plan <model> --replan.
cargo test --release # whole workspace
cargo test --release -p flux-plan experts # one crate, filtered by test name
cargo fmt
Every crate needs the submodule checked out. Crates that link flux-native also need the backend built.
Flux's backend changes live in patches/llama.cpp/ as git diff output against backend.pin; the submodule's working tree is dirty by design. Never commit inside the submodule. Edit third_party/llama.cpp in place, rebuild with scripts/build-backend.sh, then regenerate both patches:
git -C third_party/llama.cpp diff -- ggml/src/ggml-cpu/arch-fallback.h ggml/src/ggml-cpu/arch/x86/quants.c \
> patches/llama.cpp/0002-flux-q2_0-avx2.patch
git -C third_party/llama.cpp diff -- . ':!ggml/src/ggml-cpu/arch-fallback.h' ':!ggml/src/ggml-cpu/arch/x86/quants.c' \
> patches/llama.cpp/0001-flux-backend-extensions.patch
Every plan's key includes a hash of the patches, the built libraries, and the flux-native and flux-worker sources, so a change to any of them invalidates all saved plans and probe reports. Run flux plan again afterwards. The default build compiles only the CPU and CUDA backends, so patch edits to Metal, Vulkan, SYCL, and other backends go unchecked.
MIT