cyqlelabs/flux

The fastest way to run an LLM on your machine: Flux measures your GPUs, CPU, and storage once, splits the model across them, and serves it OpenAI-style in seconds.

Rust

2

57 commits

updated Oct 7, 2026

See the code

See what people are saying

README

Flux logo: a cheerful cartoon elephant in a driver's cap squeezed onto a tiny red car

Flux

Squeezes every last token out of your machine.

Flux is a measured execution planner and runtime for LLM inference. It measures your GPUs, CPU, and storage, searches placements of a model's layers and mixture-of-experts weights across them, times the fastest candidates on real prompts, and saves the winner as an immutable plan. When a model outgrows RAM, you can let the plan stream the remaining weights from an NVMe drive. flux serve then runs that plan behind an OpenAI-compatible API.

Cost models only rank and prune candidates; Flux always saves the plan that measured fastest. It runs on a pinned, patched build of llama.cpp and can also plan for llama-server or any OpenAI-compatible engine you configure.

Planning is model compilation. A compiler fits a program to the processor that will run it; flux plan fits a model to the machine that will serve it. It leaves the weights untouched and decides where each layer and expert runs. Like a binary, the plan is built once and run many times, and a new GPU, driver, or backend build calls for a new one.

A GGUF or Hugging Face model goes through inspect, probe, and plan; the saved plan feeds serve and bench

Compared with hand-tuned engines

FluxHand-tuned single-model engines
ModelsAny of the 148 architectures the pinned llama.cpp implementsThe one model they were tuned for
Ready to serveIn seconds: weights are memory-mapped and page-locked in the background; 15 s for Qwen3.8-Flash-Next on an RTX 2060 + RTX 3060Minutes: the experts are read into RAM before the first request, longest on a cold start
CPU's share of the workMeasured on your machine: RAM bandwidth per thread count and the model's own kernelsFixed constants; an optional calibration run adjusts a few of them
Layer and expert placementSearched across the CPU and GPUs; the finalists are timed on real promptsFixed rules decide each token's split
Several GPUsSplits layers across GPUs and caches experts on them, in the same planA layer split or extra expert caches, one at a time
EngineCompares its native engine, llama-server, and any engine you register, then keeps the fastestOne engine
Models bigger than RAMMeasures the drive that holds the model and predicts the decode limit of reading weights from it; opt in with --allow-storage-streamingA low-RAM mode chosen from the amount of RAM, without testing the drive
Changes while servingWatches decode speed and can replan in placeCalibration runs only when you start it

Requirements

ComponentRequirement
OSLinux x86_64
Rust1.85 or newer
Build toolsCMake, a C++17 compiler, and git
GPUAn NVIDIA GPU and the CUDA toolkit for the default build; set GGML_CUDA=OFF to build for CPU only
PythonPython 3 with torch, numpy, and transformers, needed only by flux convert

Supported models

Flux runs any GGUF file whose general.architecture the pinned llama.cpp implements: 148 architectures at the current pin, including the Llama, Qwen3, Qwen3-Next, Gemma 3, DeepSeek, gpt-oss, GLM, and MiniCPM families. The list moves with backend.pin, because Flux reads it from llama.cpp's source at build time. Run flux inspect <model> to check a file before planning.

Model formatHow Flux runs it
GGUF with a supported architectureThe native engine or llama-server
Hugging Face safetensorsConvert it to GGUF with flux convert, or register an engine for it in flux.toml
EXL3, GPTQ, AWQ, or FP8Register an engine that serves the format, such as TabbyAPI or vLLM, in flux.toml

Expert caching and the second-GPU tier apply only to mixture-of-experts models; dense models get measured layer placement. A plan cannot exceed the model's trained context. Flux has been tested end to end on Qwen3.8-Flash-Next, Qwen3.6-35B-A3B, and MiniCPM5.

How Flux uses the CPU

The CPU is a device in every plan, not a fallback for what the GPUs cannot hold. Flux times its RAM bandwidth and the model's own kernels on it, places whole layers or a layer's experts on it, and fills spare GPU memory with the experts that save the most time per byte. The finalists run on real prompts, CPU work included, so the CPU's share and its thread count are measured rather than assumed.

While serving, the CPU computes its experts while the GPU computes the rest of the layer. Prompt chunks of 32 tokens or more copy those experts to the GPU instead, where the larger batch runs faster.

Long contexts

flux plan gives each model its trained context: 262,144 tokens for Qwen3.8-Flash-Next. The KV cache commits memory as a conversation grows, not up front:

  • The first 65,536 tokens of attention KV per sequence stay in VRAM, so conversations up to that length run at full speed.
  • Past that, KV pages go to pinned RAM on GPUs that can read it. Sparse attention, as in Flash-Next, reads only the cells it selects; dense models decode more slowly at those depths, and very long prompts prefill more slowly.
  • flux serve gives KV pages the RAM available when it starts, minus the weights it pins and the host reserve (host_reserve_percent, 10%). flux plan caps the context only when that RAM cannot hold the trained one.

A request that does not fit fails at admission with HTTP 400 and OpenAI's context_length_exceeded error, which clients such as Qwen Code answer by compressing their history; a running step never fails for lack of memory. Each request must leave room for min(max_tokens, serve.min_reply) reply tokens (min_reply defaults to 4,096). /v1/models reports the limit as context_length, max_model_len and meta.n_ctx.

When a prompt replaces a conversation, the worker parks the old one's state in RAM, so an agent that returns to it after a side request skips recomputing the prefix. flux show <plan> prints the page budgets, and /flux/stats reports memory by use.

Build

git clone --recurse-submodules https://github.com/cyqlelabs/flux.git
cd flux
# Optional: install Python dependencies if you intend to run `flux convert`
pip install -r requirements.txt

scripts/build-backend.sh
cargo build --release -p flux-cli -p flux-worker

build-backend.sh ensures third_party/llama.cpp matches backend.pin, applies the patches in patches/llama.cpp/, and builds llama.cpp for CUDA compute capabilities 7.5 and 8.6 (RTX 20 and RTX 30 series). For other GPUs set CUDA_ARCHS, for example CUDA_ARCHS=89 scripts/build-backend.sh for the RTX 40 series. The flux and flux-worker binaries land in target/release/.

scripts/package.sh builds a relocatable tarball in dist/ that bundles both binaries, the llama.cpp libraries, and a starter flux.toml.

Quick start

export PATH="$PWD/target/release:$PATH"

flux plan path/to/model.gguf
flux serve <plan-id>

The first flux plan takes several minutes; flux serve then starts in seconds. Planning probes the hardware, downloads the WikiText-2 prompt corpus, and times the finalist placements within a 600-second tuning budget (--budget-s changes it). It validates the top two on held-out prompts and on a built-in set of coding-agent conversations, then saves the winner. Later runs for the same model, machine, and workload reuse the saved plan; pass --replan to measure again. flux plans lists saved plans, and flux serve accepts any unique prefix of a plan id.

The server listens on 127.0.0.1:8090:

curl http://127.0.0.1:8090/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 64}'

Commands

CommandPurpose
flux inspect <path>Identify a GGUF file or Hugging Face checkpoint: metadata, shards, hashes, and compatibility
flux fetch <repo> --include <glob>Download and verify Hugging Face files at a pinned revision
flux convert <dir>Convert a Hugging Face checkpoint to GGUF with the pinned converter
flux corpusDownload the WikiText-2 prompt corpus
flux probeMeasure copy curves, contention, kernel shapes, and CPU and storage bandwidth
flux plan <model>Find, measure, and save the fastest validated plan
flux plans, flux show <id>List saved plans, or show one with its decisions and measurements
flux trace <plan>Attribute decode-step time to devices and operations
flux serve <plan>Serve a plan over the OpenAI-compatible API
flux bench <suite>Benchmark speed, quality, conformance, and resilience

Run flux <command> --help for every flag.

Common flux plan flags
FlagDefaultEffect
--ctxThe model's trained context, capped only when RAM cannot hold its KV pagesTokens per sequence the plan must hold, prompt plus output
--concurrency1Concurrent sequences the plan must hold
--serving-p95-msoffOptimize aggregate tokens per second under this p95 per-token latency
--enginesnative,llama-serverEngines to compare, including any named in flux.toml
--kvf16KV cache type; any other type is labelled as a separate quality profile
--speculationoffAlso measure drafting with --draft-model; models with next-token heads are measured without it
--draft-model, --headsnoneDraft with a separate model, or graft next-token (MTP) heads onto the model
--allow-storage-streamingoffLet a plan read the weights that do not fit in RAM from the drive; without it, Flux rejects such placements and reports their decode limit
--replan, --reprobeoffIgnore the saved plan or the saved probe report
flux bench suites
SuitePurpose
runRace Flux against the strongest tuned baseline in paired, randomized trials
qualityCompare a plan with a reference placement by KL divergence, top-1 agreement, and perplexity
conformanceCheck that the native engine and llama-server agree on templates, special tokens, sampling, and stops
soakHit flux serve with overload, cancellations, and injected faults
arrivalsReplay Poisson arrivals and a long-running chat on Flux and on llama.cpp auto-fit
ablateCompare the full plan with variants that each remove one optimization
show, matrixSummarize one suite, or print a pass/fail matrix over every saved plan

HTTP API

flux serve --host and --port override the default address.

RoutePurpose
GET /v1/modelsList the served model
POST /v1/chat/completionsCreate chat completions, streamed or not; replies are split into reasoning_content, content, and tool_calls as llama-server does
POST /v1/completionsCreate text completions, streamed or not
GET /healthReturn 200 when the worker is ready and admission is open, otherwise 503
GET /liveReturn 200 while the HTTP server is alive
GET /flux/planReturn the plan being served
GET /flux/statsReport admission, decode drift, and worker counters
POST /flux/replanDrain, plan again, and switch; restore the old plan if the new one fails to load
POST /flux/tokenizeTokenize text with the model's vocabulary
POST /flux/admissionAdjust the host-memory reserve with {"host_reserve_bytes": N}
GET /flux/requests/{id}Return a journaled request's status and text
GET /flux/requests/{id}/stream?after=NRead retained output starting at event index N (default 0)

A request's id is its x-request-id header or, when that is absent, the id in the response. Output is journaled in memory for 600 seconds after a request completes (earlier under memory pressure), so a client that drops its connection can resume the stream: each event carries an index, and after=N+1 continues after event N. An expired request returns 404.

A worker that fails while idle restarts on its own. One that fails mid-request ends that request with an explicit error; start a new request to generate again.

Configuration

Flux reads $FLUX_CONFIG, then ~/.config/flux/flux.toml. Every field has a default, so the file only lists overrides:

models_dir = "/data/models"       # where flux fetch writes
cache_dir = "/data/flux-cache"    # hash index, prepared artifacts, corpora
host_reserve_percent = 10        # percentage of total RAM kept free

[plan]
tuning_budget_s = 600

[serve]
port = 8090
prefill_timeout_s = 300
decode_timeout_s = 120
client_write_timeout_s = 30
replan_timeout_s = 3600
journal_max_bytes = 268435456
min_reply = 4096

[serve.worker_timeouts]
hello_s = 30
load_s = 1800
rpc_s = 120
write_s = 30

# Any OpenAI-compatible engine; {port}, {model} and {ctx} are substituted.
[engines.my-engine]
command = ["my-engine", "serve"]
args = ["--port", "{port}", "--model", "{model}", "--ctx-size", "{ctx}"]
architectures = ["qwen3moe"]

A registered engine must stream per-token logprobs and completion usage, which Flux needs to time it.

Plans, probe reports, logs, and benchmark results live in $XDG_DATA_HOME/flux, which defaults to ~/.local/share/flux. flux serve writes worker output to logs/serve-worker.log there. All defaults are in crates/flux-core/src/config.rs.

Planning and serving share the host-memory reserve. POST /flux/admission overrides it until the server restarts, and /flux/stats reports the current value.

Environment variables
VariableEffect
FLUX_CONFIGPath of flux.toml
FLUX_LOGTracing filter for flux (default info)
FLUX_NATIVE_LOGMinimum ggml log level the bridge prints: debug, info, warn (default) or error
FLUX_WORKERPath of the worker binary (default: flux-worker next to flux)
FLUX_MOE_HOST_PROFILE=1Time host (CPU) expert work per step
FLUX_CUDA_OP_PROFILE=1Time each CUDA operation
FLUX_CUDA_OP_PROFILE=2Time every CUDA operation without truncating the op table
FLUX_MOE_HOST_SYNC=1Stop overlapping CPU experts with the GPU
FLUX_MOE_CACHE_FREEZE=1Stop the GPU expert cache from adapting
FLUX_MOE_TIER_WAIT=1Wait for a second GPU's expert tier instead of recomputing its late work on the CPU, so greedy outputs repeat exactly
FLUX_MOE_ROUTE_DUMP=<path>Write each MoE layer's expert selection counts to this file as 16-bit (layer, expert, count) records
GGML_OP_OFFLOAD_MIN_BATCHBatch size at which ops on host weights move to a GPU (default 32)

Throughput depends on everything else the machine is doing: a busy browser can halve decode speed, and a cold page cache slows the first prompts. Plan and benchmark on an idle machine.

Architecture

Only flux-worker links llama.cpp, so a native crash never takes down flux. Planning and probing run the worker as one-shot jobs. Serving and plan validation keep a flux-worker serve process alive and talk to it over a versioned JSON-lines protocol.

An HTTP client calls flux, which talks JSON lines to flux-worker, which calls llama.cpp through the flux-native bridge

CrateRole
flux-cliThe flux binary
flux-ingestGGUF and Hugging Face manifests, hashing, fetching, and conversion
flux-probeHardware measurements, saved per topology
flux-planPlacement search, expert cache sizing, finalist measurement, and the plan store
flux-serveThe OpenAI API, admission control, request journal, and drift-triggered replanning
flux-benchPaired trials, quality, conformance, and soak tests
flux-coreShared types: plan, config, worker protocol, and supervisor
flux-nativeC ABI bridge to llama.cpp that exchanges JSON for complex values
flux-workerThe flux-worker binary

A saved plan never changes. It is keyed by the model's file hashes, the hardware, the backend build, the driver, and the planning flags and settings, so a change to any of them needs a new plan.

If a GPU has less free memory when flux serve starts than the plan was measured to need, the server drops that GPU's least-routed cached experts and logs a warning. When the cached experts cannot cover the shortfall, it refuses the plan; free the memory or run flux plan <model> --replan.

Development

cargo test --release                        # whole workspace
cargo test --release -p flux-plan experts   # one crate, filtered by test name
cargo fmt

Every crate needs the submodule checked out. Crates that link flux-native also need the backend built.

Changing the llama.cpp backend

Flux's backend changes live in patches/llama.cpp/ as git diff output against backend.pin; the submodule's working tree is dirty by design. Never commit inside the submodule. Edit third_party/llama.cpp in place, rebuild with scripts/build-backend.sh, then regenerate both patches:

git -C third_party/llama.cpp diff -- ggml/src/ggml-cpu/arch-fallback.h ggml/src/ggml-cpu/arch/x86/quants.c \
  > patches/llama.cpp/0002-flux-q2_0-avx2.patch
git -C third_party/llama.cpp diff -- . ':!ggml/src/ggml-cpu/arch-fallback.h' ':!ggml/src/ggml-cpu/arch/x86/quants.c' \
  > patches/llama.cpp/0001-flux-backend-extensions.patch

Every plan's key includes a hash of the patches, the built libraries, and the flux-native and flux-worker sources, so a change to any of them invalidates all saved plans and probe reports. Run flux plan again afterwards. The default build compiles only the CPU and CUDA backends, so patch edits to Metal, Vulkan, SYCL, and other backends go unchecked.

License

MIT

cyqlelabs/flux

The fastest way to run an LLM on your machine: Flux measures your GPUs, CPU, and storage once, splits the model across them, and serves it OpenAI-style in seconds.

Rust

2

57 commits

updated Oct 7, 2026

See the code

See what people are saying

README

Flux logo: a cheerful cartoon elephant in a driver's cap squeezed onto a tiny red car

Flux

Squeezes every last token out of your machine.

Flux is a measured execution planner and runtime for LLM inference. It measures your GPUs, CPU, and storage, searches placements of a model's layers and mixture-of-experts weights across them, times the fastest candidates on real prompts, and saves the winner as an immutable plan. When a model outgrows RAM, you can let the plan stream the remaining weights from an NVMe drive. flux serve then runs that plan behind an OpenAI-compatible API.

Cost models only rank and prune candidates; Flux always saves the plan that measured fastest. It runs on a pinned, patched build of llama.cpp and can also plan for llama-server or any OpenAI-compatible engine you configure.

Planning is model compilation. A compiler fits a program to the processor that will run it; flux plan fits a model to the machine that will serve it. It leaves the weights untouched and decides where each layer and expert runs. Like a binary, the plan is built once and run many times, and a new GPU, driver, or backend build calls for a new one.

A GGUF or Hugging Face model goes through inspect, probe, and plan; the saved plan feeds serve and bench

Compared with hand-tuned engines

FluxHand-tuned single-model engines
ModelsAny of the 148 architectures the pinned llama.cpp implementsThe one model they were tuned for
Ready to serveIn seconds: weights are memory-mapped and page-locked in the background; 15 s for Qwen3.8-Flash-Next on an RTX 2060 + RTX 3060Minutes: the experts are read into RAM before the first request, longest on a cold start
CPU's share of the workMeasured on your machine: RAM bandwidth per thread count and the model's own kernelsFixed constants; an optional calibration run adjusts a few of them
Layer and expert placementSearched across the CPU and GPUs; the finalists are timed on real promptsFixed rules decide each token's split
Several GPUsSplits layers across GPUs and caches experts on them, in the same planA layer split or extra expert caches, one at a time
EngineCompares its native engine, llama-server, and any engine you register, then keeps the fastestOne engine
Models bigger than RAMMeasures the drive that holds the model and predicts the decode limit of reading weights from it; opt in with --allow-storage-streamingA low-RAM mode chosen from the amount of RAM, without testing the drive
Changes while servingWatches decode speed and can replan in placeCalibration runs only when you start it

Requirements

ComponentRequirement
OSLinux x86_64
Rust1.85 or newer
Build toolsCMake, a C++17 compiler, and git
GPUAn NVIDIA GPU and the CUDA toolkit for the default build; set GGML_CUDA=OFF to build for CPU only
PythonPython 3 with torch, numpy, and transformers, needed only by flux convert

Supported models

Flux runs any GGUF file whose general.architecture the pinned llama.cpp implements: 148 architectures at the current pin, including the Llama, Qwen3, Qwen3-Next, Gemma 3, DeepSeek, gpt-oss, GLM, and MiniCPM families. The list moves with backend.pin, because Flux reads it from llama.cpp's source at build time. Run flux inspect <model> to check a file before planning.

Model formatHow Flux runs it
GGUF with a supported architectureThe native engine or llama-server
Hugging Face safetensorsConvert it to GGUF with flux convert, or register an engine for it in flux.toml
EXL3, GPTQ, AWQ, or FP8Register an engine that serves the format, such as TabbyAPI or vLLM, in flux.toml

Expert caching and the second-GPU tier apply only to mixture-of-experts models; dense models get measured layer placement. A plan cannot exceed the model's trained context. Flux has been tested end to end on Qwen3.8-Flash-Next, Qwen3.6-35B-A3B, and MiniCPM5.

How Flux uses the CPU

The CPU is a device in every plan, not a fallback for what the GPUs cannot hold. Flux times its RAM bandwidth and the model's own kernels on it, places whole layers or a layer's experts on it, and fills spare GPU memory with the experts that save the most time per byte. The finalists run on real prompts, CPU work included, so the CPU's share and its thread count are measured rather than assumed.

While serving, the CPU computes its experts while the GPU computes the rest of the layer. Prompt chunks of 32 tokens or more copy those experts to the GPU instead, where the larger batch runs faster.

Long contexts

flux plan gives each model its trained context: 262,144 tokens for Qwen3.8-Flash-Next. The KV cache commits memory as a conversation grows, not up front:

  • The first 65,536 tokens of attention KV per sequence stay in VRAM, so conversations up to that length run at full speed.
  • Past that, KV pages go to pinned RAM on GPUs that can read it. Sparse attention, as in Flash-Next, reads only the cells it selects; dense models decode more slowly at those depths, and very long prompts prefill more slowly.
  • flux serve gives KV pages the RAM available when it starts, minus the weights it pins and the host reserve (host_reserve_percent, 10%). flux plan caps the context only when that RAM cannot hold the trained one.

A request that does not fit fails at admission with HTTP 400 and OpenAI's context_length_exceeded error, which clients such as Qwen Code answer by compressing their history; a running step never fails for lack of memory. Each request must leave room for min(max_tokens, serve.min_reply) reply tokens (min_reply defaults to 4,096). /v1/models reports the limit as context_length, max_model_len and meta.n_ctx.

When a prompt replaces a conversation, the worker parks the old one's state in RAM, so an agent that returns to it after a side request skips recomputing the prefix. flux show <plan> prints the page budgets, and /flux/stats reports memory by use.

Build

git clone --recurse-submodules https://github.com/cyqlelabs/flux.git
cd flux
# Optional: install Python dependencies if you intend to run `flux convert`
pip install -r requirements.txt

scripts/build-backend.sh
cargo build --release -p flux-cli -p flux-worker

build-backend.sh ensures third_party/llama.cpp matches backend.pin, applies the patches in patches/llama.cpp/, and builds llama.cpp for CUDA compute capabilities 7.5 and 8.6 (RTX 20 and RTX 30 series). For other GPUs set CUDA_ARCHS, for example CUDA_ARCHS=89 scripts/build-backend.sh for the RTX 40 series. The flux and flux-worker binaries land in target/release/.

scripts/package.sh builds a relocatable tarball in dist/ that bundles both binaries, the llama.cpp libraries, and a starter flux.toml.

Quick start

export PATH="$PWD/target/release:$PATH"

flux plan path/to/model.gguf
flux serve <plan-id>

The first flux plan takes several minutes; flux serve then starts in seconds. Planning probes the hardware, downloads the WikiText-2 prompt corpus, and times the finalist placements within a 600-second tuning budget (--budget-s changes it). It validates the top two on held-out prompts and on a built-in set of coding-agent conversations, then saves the winner. Later runs for the same model, machine, and workload reuse the saved plan; pass --replan to measure again. flux plans lists saved plans, and flux serve accepts any unique prefix of a plan id.

The server listens on 127.0.0.1:8090:

curl http://127.0.0.1:8090/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 64}'

Commands

CommandPurpose
flux inspect <path>Identify a GGUF file or Hugging Face checkpoint: metadata, shards, hashes, and compatibility
flux fetch <repo> --include <glob>Download and verify Hugging Face files at a pinned revision
flux convert <dir>Convert a Hugging Face checkpoint to GGUF with the pinned converter
flux corpusDownload the WikiText-2 prompt corpus
flux probeMeasure copy curves, contention, kernel shapes, and CPU and storage bandwidth
flux plan <model>Find, measure, and save the fastest validated plan
flux plans, flux show <id>List saved plans, or show one with its decisions and measurements
flux trace <plan>Attribute decode-step time to devices and operations
flux serve <plan>Serve a plan over the OpenAI-compatible API
flux bench <suite>Benchmark speed, quality, conformance, and resilience

Run flux <command> --help for every flag.

Common flux plan flags
FlagDefaultEffect
--ctxThe model's trained context, capped only when RAM cannot hold its KV pagesTokens per sequence the plan must hold, prompt plus output
--concurrency1Concurrent sequences the plan must hold
--serving-p95-msoffOptimize aggregate tokens per second under this p95 per-token latency
--enginesnative,llama-serverEngines to compare, including any named in flux.toml
--kvf16KV cache type; any other type is labelled as a separate quality profile
--speculationoffAlso measure drafting with --draft-model; models with next-token heads are measured without it
--draft-model, --headsnoneDraft with a separate model, or graft next-token (MTP) heads onto the model
--allow-storage-streamingoffLet a plan read the weights that do not fit in RAM from the drive; without it, Flux rejects such placements and reports their decode limit
--replan, --reprobeoffIgnore the saved plan or the saved probe report
flux bench suites
SuitePurpose
runRace Flux against the strongest tuned baseline in paired, randomized trials
qualityCompare a plan with a reference placement by KL divergence, top-1 agreement, and perplexity
conformanceCheck that the native engine and llama-server agree on templates, special tokens, sampling, and stops
soakHit flux serve with overload, cancellations, and injected faults
arrivalsReplay Poisson arrivals and a long-running chat on Flux and on llama.cpp auto-fit
ablateCompare the full plan with variants that each remove one optimization
show, matrixSummarize one suite, or print a pass/fail matrix over every saved plan

HTTP API

flux serve --host and --port override the default address.

RoutePurpose
GET /v1/modelsList the served model
POST /v1/chat/completionsCreate chat completions, streamed or not; replies are split into reasoning_content, content, and tool_calls as llama-server does
POST /v1/completionsCreate text completions, streamed or not
GET /healthReturn 200 when the worker is ready and admission is open, otherwise 503
GET /liveReturn 200 while the HTTP server is alive
GET /flux/planReturn the plan being served
GET /flux/statsReport admission, decode drift, and worker counters
POST /flux/replanDrain, plan again, and switch; restore the old plan if the new one fails to load
POST /flux/tokenizeTokenize text with the model's vocabulary
POST /flux/admissionAdjust the host-memory reserve with {"host_reserve_bytes": N}
GET /flux/requests/{id}Return a journaled request's status and text
GET /flux/requests/{id}/stream?after=NRead retained output starting at event index N (default 0)

A request's id is its x-request-id header or, when that is absent, the id in the response. Output is journaled in memory for 600 seconds after a request completes (earlier under memory pressure), so a client that drops its connection can resume the stream: each event carries an index, and after=N+1 continues after event N. An expired request returns 404.

A worker that fails while idle restarts on its own. One that fails mid-request ends that request with an explicit error; start a new request to generate again.

Configuration

Flux reads $FLUX_CONFIG, then ~/.config/flux/flux.toml. Every field has a default, so the file only lists overrides:

models_dir = "/data/models"       # where flux fetch writes
cache_dir = "/data/flux-cache"    # hash index, prepared artifacts, corpora
host_reserve_percent = 10        # percentage of total RAM kept free

[plan]
tuning_budget_s = 600

[serve]
port = 8090
prefill_timeout_s = 300
decode_timeout_s = 120
client_write_timeout_s = 30
replan_timeout_s = 3600
journal_max_bytes = 268435456
min_reply = 4096

[serve.worker_timeouts]
hello_s = 30
load_s = 1800
rpc_s = 120
write_s = 30

# Any OpenAI-compatible engine; {port}, {model} and {ctx} are substituted.
[engines.my-engine]
command = ["my-engine", "serve"]
args = ["--port", "{port}", "--model", "{model}", "--ctx-size", "{ctx}"]
architectures = ["qwen3moe"]

A registered engine must stream per-token logprobs and completion usage, which Flux needs to time it.

Plans, probe reports, logs, and benchmark results live in $XDG_DATA_HOME/flux, which defaults to ~/.local/share/flux. flux serve writes worker output to logs/serve-worker.log there. All defaults are in crates/flux-core/src/config.rs.

Planning and serving share the host-memory reserve. POST /flux/admission overrides it until the server restarts, and /flux/stats reports the current value.

Environment variables
VariableEffect
FLUX_CONFIGPath of flux.toml
FLUX_LOGTracing filter for flux (default info)
FLUX_NATIVE_LOGMinimum ggml log level the bridge prints: debug, info, warn (default) or error
FLUX_WORKERPath of the worker binary (default: flux-worker next to flux)
FLUX_MOE_HOST_PROFILE=1Time host (CPU) expert work per step
FLUX_CUDA_OP_PROFILE=1Time each CUDA operation
FLUX_CUDA_OP_PROFILE=2Time every CUDA operation without truncating the op table
FLUX_MOE_HOST_SYNC=1Stop overlapping CPU experts with the GPU
FLUX_MOE_CACHE_FREEZE=1Stop the GPU expert cache from adapting
FLUX_MOE_TIER_WAIT=1Wait for a second GPU's expert tier instead of recomputing its late work on the CPU, so greedy outputs repeat exactly
FLUX_MOE_ROUTE_DUMP=<path>Write each MoE layer's expert selection counts to this file as 16-bit (layer, expert, count) records
GGML_OP_OFFLOAD_MIN_BATCHBatch size at which ops on host weights move to a GPU (default 32)

Throughput depends on everything else the machine is doing: a busy browser can halve decode speed, and a cold page cache slows the first prompts. Plan and benchmark on an idle machine.

Architecture

Only flux-worker links llama.cpp, so a native crash never takes down flux. Planning and probing run the worker as one-shot jobs. Serving and plan validation keep a flux-worker serve process alive and talk to it over a versioned JSON-lines protocol.

An HTTP client calls flux, which talks JSON lines to flux-worker, which calls llama.cpp through the flux-native bridge

CrateRole
flux-cliThe flux binary
flux-ingestGGUF and Hugging Face manifests, hashing, fetching, and conversion
flux-probeHardware measurements, saved per topology
flux-planPlacement search, expert cache sizing, finalist measurement, and the plan store
flux-serveThe OpenAI API, admission control, request journal, and drift-triggered replanning
flux-benchPaired trials, quality, conformance, and soak tests
flux-coreShared types: plan, config, worker protocol, and supervisor
flux-nativeC ABI bridge to llama.cpp that exchanges JSON for complex values
flux-workerThe flux-worker binary

A saved plan never changes. It is keyed by the model's file hashes, the hardware, the backend build, the driver, and the planning flags and settings, so a change to any of them needs a new plan.

If a GPU has less free memory when flux serve starts than the plan was measured to need, the server drops that GPU's least-routed cached experts and logs a warning. When the cached experts cannot cover the shortfall, it refuses the plan; free the memory or run flux plan <model> --replan.

Development

cargo test --release                        # whole workspace
cargo test --release -p flux-plan experts   # one crate, filtered by test name
cargo fmt

Every crate needs the submodule checked out. Crates that link flux-native also need the backend built.

Changing the llama.cpp backend

Flux's backend changes live in patches/llama.cpp/ as git diff output against backend.pin; the submodule's working tree is dirty by design. Never commit inside the submodule. Edit third_party/llama.cpp in place, rebuild with scripts/build-backend.sh, then regenerate both patches:

git -C third_party/llama.cpp diff -- ggml/src/ggml-cpu/arch-fallback.h ggml/src/ggml-cpu/arch/x86/quants.c \
  > patches/llama.cpp/0002-flux-q2_0-avx2.patch
git -C third_party/llama.cpp diff -- . ':!ggml/src/ggml-cpu/arch-fallback.h' ':!ggml/src/ggml-cpu/arch/x86/quants.c' \
  > patches/llama.cpp/0001-flux-backend-extensions.patch

Every plan's key includes a hash of the patches, the built libraries, and the flux-native and flux-worker sources, so a change to any of them invalidates all saved plans and probe reports. Run flux plan again afterwards. The default build compiles only the CPU and CUDA backends, so patch edits to Metal, Vulkan, SYCL, and other backends go unchecked.

License

MIT