notactuallytreyanastasio/shoehorn

88

stars

57

commits

Rust

primary language

Aug 19, 2026

updated

notactuallytreyanastasio.github.io/shoehorn/

README

shoehorn

Quantize a BF16 GGUF model with an imatrix so it fits exactly into your available VRAM, then run it with llama.cpp.

shoehorn's site has downloads and a browser-side "what fits your machine?" calculator.

Preset quantizations (Q4_K_M, Q5_K_S, ...) ignore your hardware. Pick one that fits your machine and you either leave hundreds of megabytes of quality unused or find out at load time that it didn't fit after all. shoehorn starts from the memory you actually have, subtracts what inference itself will need (KV cache, compute buffers), and solves a per-tensor mixed-precision assignment whose total size lands within a rounding error of the remainder. Every spare megabyte goes where the importance matrix says it buys the most model quality.

$ shoehorn fit unsloth/Qwen3-4B-GGUF --serve

That one command finds and downloads the BF16 GGUF from Hugging Face, picks up the repo's imatrix (or generates one locally), solves the quant mix for your machine, writes the file, and launches llama-server on it. The pieces are also available separately:

$ shoehorn vram
Apple M4 Pro: 17.76 GiB usable for GPU working set

$ shoehorn quantize -m Qwen3-0.6B-BF16.gguf -i qwen3.imatrix \
    --ctx 4096 --budget 1.75GiB -o fitted.gguf
...
weights: 519.2 MiB of 519.2 MiB budget (99.981% used, 103424 B slack) | overall 7.306 bpw

$ shoehorn run -m fitted.gguf --ctx 4096

The quantizer is implemented from scratch in this repo (Rust, no llama.cpp code linked). The output is standard GGUF v3 that any llama.cpp build, or anything downstream of it, loads directly. llama.cpp handles inference and doubles as an independent correctness oracle.


Contents


Quick start

shoehorn needs llama.cpp installed (inference backend + imatrix generation). For shoehorn itself, grab a prebuilt binary from the releases page or build it with a Rust toolchain as shown below.

macOS (Apple Silicon)

brew install notactuallytreyanastasio/shoehorn/shoehorn   # prebuilt binary + llama.cpp
shoehorn fit unsloth/Qwen3-4B-GGUF --serve

Or from source: brew install llama.cpp then cargo install --path .

Linux (NVIDIA)

# llama.cpp with CUDA — build from source (or grab a -cuda release binary
# from https://github.com/ggml-org/llama.cpp/releases and put it on PATH):
git clone https://github.com/ggml-org/llama.cpp
cmake -S llama.cpp -B llama.cpp/build -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j
sudo cmake --install llama.cpp/build

cargo install --path .
shoehorn fit unsloth/Qwen3-4B-GGUF --serve

The VRAM probe reads the first NVIDIA device's free memory through NVML (ships with the regular driver — nothing extra to install), falling back to rocm-smi for AMD. Intel GPUs aren't probed yet; pass --budget. On AMD, use a Vulkan or ROCm build of llama.cpp.

Windows (NVIDIA)

  1. Install Rust via rustup, then cargo install --path .
  2. Download a -cuda llama.cpp zip from the releases page, unzip it, and add the folder to PATH.
  3. shoehorn fit unsloth/Qwen3-4B-GGUF --serve

Windows support is compile-tested but not yet field-tested. One known gap: auto-generating an imatrix is skipped there (the calibration text is built from man pages), so fit a repo that publishes one — it's picked up automatically — or pass -i yourself.

fit does the whole pipeline: finds the BF16 GGUF in the repo, downloads it to ~/.cache/shoehorn (resumable), picks up or generates an imatrix, solves the mix for your machine, writes <model>-fit.gguf, and serves it. macOS on Apple Silicon and Linux/Windows with an NVIDIA GPU are probed automatically; --budget/--target work anywhere.

Prefer buttons to flags? See the web UI.

Doing the steps by hand instead:

# 1. Get a BF16 (or F16/F32) GGUF of your model.
#    Most HF quant repos (unsloth, bartowski, ggml-org) publish one.

# 2. Generate an importance matrix over calibration text:
llama-imatrix -m model-bf16.gguf -f calibration.txt -o model.imatrix -ngl 99

# 3. Solve and quantize to your machine's actual capacity:
shoehorn quantize -m model-bf16.gguf -i model.imatrix --ctx 8192 -o fitted.gguf

# 4. Serve it (execs llama-server with full GPU offload):
shoehorn run -m fitted.gguf --ctx 8192

shoehorn plan takes the same flags as quantize without -o and prints the solved per-tensor mix without writing anything, so you can preview what a budget implies before spending the encode time. ./demo/run.sh reproduces the full size/quality ladder on a small model in a few minutes.

The web UI

shoehorn ui

opens a local page (default http://127.0.0.1:7788) that drives the same pipeline with no flags: it shows the measured budget for your machine, takes a model name, and runs the fit with the log streamed into the page.

the fit form: your machine's measurement, a model field, and one button

Don't know which model to pick? What fits this machine? ranks Hugging Face's most-downloaded fittable models by what your budget affords:

discovery: ranked fittable models with Use buttons

While the fit runs, the budget renders as a tape measure — weights grow from the left while the fixed costs (KV cache, compute buffers, safety margin) hold the right edge. When it lands, Chat with it serves the result with llama-server and opens its chat page once the model is warm.

a finished fit: the tape-measure gauge at 99.998% used, and a Chat with it button

Step by step:

  1. shoehorn ui from the directory you want the fitted model written to — the output <model>-fit.gguf lands in the current directory, like the CLI. The browser opens by itself (--no-open to suppress, --port to move it off 7788).
  2. Type a model — or click What fits this machine? in the machine card to have discover rank fittable models for your budget, each with a Use button (also reachable directly at /#discover). The field suggests a few known-good repos and live-searches Hugging Face as you type. First fit of a repo downloads the BF16, which for big models is tens of GB — the download resumes if interrupted and is cached in ~/.cache/shoehorn.
  3. Pick how much conversation room you want. More context means a bigger KV cache, which comes straight out of the weight budget.
  4. Fit it to my machine — or Preview the mix first, which solves and shows the tape and the per-type breakdown without writing anything (fit --dry-run underneath), then offers Fit it for real. Downloads render as a progress bar; Stop aborts cleanly, and reloading the page mid-fit picks the run back up.
  5. Measure the cost (optional) runs shoehorn eval on the fit against the original model and reports the perplexity delta right in the card — the honest answer to "how much quality did I give up?"
  6. Chat with it starts llama-server on port 8093 and opens its chat page once the model is loaded. The fitted GGUF is a normal file — anything that speaks GGUF (LM Studio, ollama, llama.cpp on another machine) can use it afterward.
  7. Ctrl-C in the terminal running shoehorn ui shuts everything down.

Finished fits are remembered (in ~/.cache/shoehorn/fits.json) and listed in a "Previous fits" card with one-click Chat buttons, so a model fitted last week doesn't need re-fitting to be served again.

The advanced knobs (KV cache type, budget override, --calibrate) are under "More options". The page runs shoehorn fit as a subprocess and streams its output, so it can't drift from what the CLI does — everything the CLI would have printed is under "Show the work".

How it works

The pipeline runs in five stages.

Probing comes first. On Apple Silicon, "VRAM" is not RAM: Metal will only wire a fraction of unified memory for the GPU. shoehorn asks the Metal device for recommendedMaxWorkingSetSize (17.76 GiB on a 24 GB M4 Pro, about 75%). On other machines it asks NVML — or rocm-smi on AMD — for the first device's free VRAM. --budget overrides the probe, which also lets you quantize for a different machine ("make this fit my friend's 8 GiB M1") or for an artificial envelope.

Next it computes the budget. From the target context length and the model's own GGUF hyperparameters, shoehorn computes the KV cache size exactly and estimates the compute buffer, then subtracts both plus a safety --reserve. What remains is the weight budget. See The budget model.

Then it measures. Every quantizable tensor gets a candidate ladder — the IQ codebook formats (IQ2_XXS up through IQ4_XS), the K-quants, Q8_0, and F16, or the legacy 32-block formats when the row length isn't divisible by 256. Each (tensor, candidate) pair is scored by actually encoding and decoding a sample of rows and accumulating the imatrix-weighted squared error. That is the true end-to-end distortion under the decoder llama.cpp will use. The work parallelizes across all cores; on Qwen3-0.6B the whole measure and solve pass takes 0.7 s.

The solve itself is a multiple-choice knapsack: pick one type per tensor, minimize total weighted error subject to total bytes staying at or under budget. shoehorn uses Lagrangian relaxation (bisect the shadow price of a byte; each tensor independently picks the candidate minimizing err + λ·bytes), then a greedy pass that spends the slack the relaxation leaves behind: repeatedly apply the single-tensor upgrade with the best error reduction per byte that still fits. Utilization in practice exceeds 99.9% of the weight budget.

Finally it writes. Chosen types are re-encoded row-parallel and streamed out as a GGUF v3 with all source metadata preserved and general.file_type set to the dominant type for display purposes. Norms, biases, and anything else 1D stay F32, matching llama.cpp convention: they are tiny and numerically sensitive.

The BF16 source is mmap'd and never fully materialized. Peak memory is roughly one tensor per worker thread, so 30B-class models are fine on a laptop.

The math

Quantized formats represent a block of weights as one or two f16 scale factors plus low-bit integers: symmetric formats decode as x̂ = d·q, asymmetric ones as x̂ = d·q − m, with sub-block scales in the K-quants. The decoder is fixed (it's llama.cpp's dequantization), so the encoder's entire job is choosing d, m, and the integers to minimize error under a weighting that reflects how much each weight matters.

llama-imatrix supplies that weighting. It runs the model over calibration text and accumulates, for each matmul weight, the mean squared activation of each input column. Columns that see large activations amplify their weights' quantization error in the layer's output, so they deserve more of the bit budget. shoehorn uses the element weight

w[j] = imatrix[j] · sqrt(σ² + x[j]²)      σ² = mean square of the row

which is the same shaping ggml's imatrix-aware quantizers use. Tensors without an imatrix entry (typically token_embd, which is never a matmul input) fall back to the activation-agnostic sqrt(σ² + x²).

For symmetric formats the encoder mirrors ggml's make_qx_quants: try 19 candidate grids iscale = −(nmax + 0.1·is)/max for is in [−9, 9]; for each, round every element and evaluate the weighted least-squares objective; keep the grid maximizing (Σ w·x·l)² / Σ w·l², whose optimal scale is Σ w·x·l / Σ w·l². For asymmetric formats it mirrors make_qkx3_quants: over 37 candidate grids, solve the two-parameter weighted regression for scale and offset jointly, clamping the offset so the K-quant d·q − dmin·m convention keeps its positive-min invariant. K-quant super-blocks then quantize the 8 or 16 sub-block scales themselves to 6 or 8 bits and re-round every element against the quantized scales.

Why copy ggml's objective instead of inventing one? The reference quantizers have years of use against this exact decoder, and matching their semantics makes shoehorn's output directly comparable to llama-quantize's. The new work went into the solver.

The solver minimizes the sum over tensors of measured weighted error, with no per-tensor normalization: the imatrix magnitudes already encode relative importance across tensors, and total weighted distortion is exactly what the knapsack should minimize.

Supported formats

formatbits/weightblockencode notes
IQ2_XXS / IQ2_XS / IQ2_S2.06 / 2.31 / 2.56256E8-lattice codebook, neighbour search, 7-bit parity-packed or verbatim signs
IQ3_XXS / IQ3_S3.06 / 3.44256D4-lattice codebook, same search machinery
IQ4_XS4.25256nonlinear 16-value codebook, 6-bit sub-scales
Q4_K4.502568×32 sub-blocks, 6-bit scales+mins, weighted 2-param regression
Q5_K5.50256as Q4_K plus a high-bit plane
Q6_K6.5625616×16 sub-blocks, 8-bit signed scales, weighted grid search
Q8_08.5032absmax scaling (error is negligible; search unnecessary)
IQ4_NL4.5032nonlinear codebook, fallback for rows not divisible by 256
Q4_0 / Q4_14.50 / 5.0032legacy fallback for rows not divisible by 256
Q5_0 / Q5_15.50 / 6.0032legacy fallback, high-bit plane
F16 / BF16 / F3216 / 16 / 32n/apassthrough / conversion

Rows divisible by 256 get the full ladder from IQ2_XXS up to F16; rows divisible only by 32 get {IQ4_NL, Q4_0, Q4_1, Q5_0, Q5_1, Q8_0, F16}. token_embd.weight and output.weight are floored at 4-bit (IQ4_XS): the embedding has no imatrix data, weighted MSE understates LM-head sensitivity, and llama.cpp's own IQ2 presets apply the same guard. Every format has a matching in-crate decoder used for error measurement and tests.

The IQ ports follow ggml-quants.c at the exact commit the installed llama.cpp was built from, including its fudge constants and sign-parity rules; the lattice tables in src/iq_tables.rs are script-extracted from the reference, never hand-typed. ggml keeps two grids per IQ format (a true lattice for the encoder, tuned magnitudes for the decoder) and shoehorn reproduces that asymmetry — see DESIGN.md D11 for the details.

The budget model

weight_budget = usable_vram − kv_cache − compute_est − reserve

kv_cache    = n_layer · ctx · n_kv_heads · (key_len + value_len) · 2   (f16 K+V, exact)
compute_est = ubatch·n_vocab·4 + ubatch·n_embd·32                      (heuristic)
ubatch      = min(512, ctx)

All hyperparameters come from the model's own GGUF metadata (block_count, attention.head_count_kv, attention.key_length, ...), so grouped-query attention and unusual head sizes are handled per model rather than assumed.

Worked example, Qwen3-0.6B at --budget 1.75GiB --ctx 4096:

1.75 GiB − 448 MiB KV (28 layers · 4096 · 8 kv-heads · 256 · 2 B)
         − 313 MiB compute est (512·151936·4 logits dominate)
         − 512 MiB reserve
         = 519.2 MiB for weights → solver fills 519.1 MiB of it

The KV term is exact. The compute term is deliberately rough, since llama.cpp's graph allocation depends on flash-attention availability, batch shape, and version; the --reserve margin absorbs its error plus Metal shader buffers and the host process. If you have a measured number for your setup, --budget and --reserve let you dial it in precisely.

CLI reference

shoehorn fit       <path | owner/repo | url> [-i <imatrix>] [fit flags] [-o out.gguf] [--serve] [--dry-run]
shoehorn plan      -m <bf16.gguf> [-i <imatrix>] [fit flags]
shoehorn quantize  -m <bf16.gguf> [-i <imatrix>] [fit flags] -o <out.gguf>
shoehorn run       -m <model.gguf> [--ctx N] [--kv q8_0] [-- <llama-server args...>]
shoehorn vram
shoehorn ui        [--port 7788] [--no-open]
shoehorn eval      -m <model.gguf> [-f <text>] [--baseline <other.gguf>] [--ctx N]
shoehorn discover  [--ctx N] [--budget <size>] [--scan N]

fit accepts a local path, a Hugging Face repo id (it picks the largest BF16 file — falling back to F16, then F32 — downloads every shard of a split model, and grabs any imatrix in the repo), or a direct URL. Downloads land in ~/.cache/shoehorn and resume if interrupted. Without an imatrix, it generates one with llama-imatrix when the model fits the GPU. --dry-run stops after printing the solved plan — like plan, but with fit's fetching, so it works on repo ids and URLs too.

Fit flags, shared by fit, plan, and quantize:

flagdefaultmeaning
-m, --modelrequiredBF16/F16/F32 source GGUF (already-quantized sources also read, via the in-crate decoders)
-i, --imatrixnoneimatrix file, legacy binary or GGUF-based; omitting it warns and falls back to activation-agnostic weighting
--ctx8192context length the KV budget is computed for
--budgetGPU probetotal memory envelope: 18GiB, 800MB, 4.5G — a bare number like 16 reads as GiB
--targetbudget for a different unified-memory machine by RAM size, e.g. --target 16GB (approximates the macOS working-set limit as 74% of RAM)
--kvf16KV cache type to budget for and run with (f16, q8_0, q4_0); q8_0 roughly halves the KV term, freeing that memory for weights
--reserve512MiB (160MiB with --calibrate)safety margin subtracted from the envelope
--calibrateoffafter writing, load the result in llama.cpp once, measure real KV/compute allocations, re-solve with them, and rewrite — spending the recovered estimate slack on quality (unchanged tensors are reused)
--exact-errorsoffscore every row instead of a 128-row sample per tensor

plan and quantize print the full per-tensor table (shape, chosen type, size, bits/weight), a by-type rollup, budget utilization with the residual slack in bytes, and the projected total VRAM picture at the target context.

run execs llama-server -m <model> -c <ctx> -ngl 99; everything after -- is passed through (--port, --api-key, ...).

vram prints the detected device and its usable GPU memory: Metal's recommended working-set size on macOS, NVML's free VRAM on the first NVIDIA device elsewhere, with a rocm-smi fallback for AMD.

eval wraps llama-perplexity so you can check what a fit actually cost: it measures perplexity on held-out text (by default man pages disjoint from the auto-imatrix calibration set) and, with --baseline, prints the delta against another model — typically the BF16 source:

$ shoehorn eval -m fitted.gguf --baseline model-bf16.gguf
fitted.gguf: PPL 8.0719
model-bf16.gguf: PPL 8.0132
delta: +0.73% vs baseline

discover answers "which model should I even fit?": it scans the most-downloaded GGUF repos on Hugging Face for full-precision sources (the same BF16 > F16 > F32 pick fit uses, so every suggestion is actually fittable), estimates the bits/weight your budget affords each one, and ranks them — biggest model in the best quality tier first. Repos whose "full-precision" file is a speculative-decoding draft companion are detected by cross-checking the file size against the parameter count in the repo name, and flagged instead of recommended. The numbers are estimates; the printed shoehorn fit <repo> --dry-run line gives the exact solve.

ui serves a local web page (and opens it) that drives the whole fit pipeline as a subprocess: pick a model, watch the budget gauge fill, then chat with the result via llama-server. No flags survive contact with it on purpose; the advanced knobs live under "More options".

Imatrix files

Both formats found in the wild are supported and auto-detected:

  • GGUF-based (current llama-imatrix output): tensors named <name>.in_sum2 (per-column sums of squared activations) and <name>.counts; shoehorn divides one by the other.
  • Legacy binary (most published imatrices): n_entries, then per entry name / ncall / nval / f32 values; values are divided by ncall.

Weights are sanitized (non-finite and non-positive entries floored) so a degenerate imatrix can't zero out the fit. For 3D expert tensors (MoE), an imatrix covering ne0 × n_expert is sliced per expert; one covering only ne0 is broadcast.

Calibration text matters less than having an imatrix at all, but it does matter at the margin: auto-generation downloads the community-standard calibration_datav3 mixed corpus (cached), falling back to concatenated man pages offline. Measured on Qwen3-0.6B at an identical 1.75 GiB budget, the mixed corpus cost +1.84% perplexity vs BF16 on neutral prose against +2.81% for man-page calibration — better generalization for one download. The test suite holds out different text from whatever calibrated.

Results

Qwen3-0.6B (596 M params), M4 Pro 24 GB, llama.cpp b10360, 2026-08-13.

Forced into a 1.75 GiB total envelope (weight budget 519.2 MiB, ctx 4096):

BF16shoehorn mix
weights on disk1.13 GiB524.8 MiB (46%)
bits/weight167.306
budget utilizationn/a99.981% (103 KB slack)
held-out perplexity¹14.533 ± 0.55214.623 ± 0.556 (+0.6%)
generation (llama-cli)n/a~210 tok/s

¹ git-rebase man page, text the imatrix never saw.

The solved mix matches what a llama.cpp veteran would hand-tune: ffn_down (known quant-tolerant) at Q4_K, attention projections at Q6_K, Q8_0 and F16 reserved for the tensors whose weighted error per byte is worst. Here it falls out of the optimization, per model, with no hand rules.

Against the machine's real 17.76 GiB budget, the 0.6B model fits at F16 and the solver keeps everything at F16 (6.9% utilization). It degrades nothing when there is no need.

Pushing into IQ territory with tighter budgets (same model, same held-out text):

envelopesizeoverall bpwheld-out PPL
BF16 baseline1.13 GiB1614.53
1.75 GiB525 MB7.3114.62
1.53 GiB300 MB~4.121.53
1.44 GiB207 MB2.84212.7
llama-quantize IQ2_XXS (control)219 MB2.34446.8

The bottom two rows are the differential validation: at comparable size, the solved mix halves the perplexity of llama.cpp's own IQ2_XXS preset, because the knapsack spends bits per tensor instead of uniformly. The absolute numbers also show why sub-3 bpw formats exist for 7B+ models: a 0.6B is severely degraded there no matter who does the quantizing.

Qwen3-14B (the real test)

Source: 29.5 GB BF16 — bigger than this machine's entire GPU working set — with bartowski's published (legacy-format) imatrix.

Against the detected 17.76 GiB budget at ctx 8192, the solver fills 99.998% of the 15.64 GiB weight budget (340 KB slack): a 9.1 bpw mix of Q8_0 (10.7 GiB), F16 (3.6 GiB where the imatrix concentrates importance), and a Q6_K/Q5_K tail. Measure + solve on the 28 GB file: 39 s.

Forced into an 8 GiB total envelope (the "make it fit my friend's 8 GiB M1" case): 100.000% of the 5.88 GiB weight budget used — 28 KB of slack — via a 3.42 bpw mix spanning the entire ladder, IQ2_XXS (87 tensors) through Q6_K. Encode time ~6 min. The result generates correct, fluent text at 23 tok/s and scores PPL 6.85 on the held-out text.

For scale: that is less than half the perplexity of the unquantized 0.6B (14.53) on the same text. Given a fixed memory budget, a big model shoehorned hard beats a small model treated gently — which is exactly the trade the solver exists to make well.

Qwen3-30B-A3B (MoE, split source)

Source: 61 GB BF16 in two llama.cpp-style shards — read as one model — with bartowski's published imatrix, targeting the real 17.76 GiB budget at ctx 8192 with --kv q8_0.

The solver filled the 16.88 GiB weight budget to 0 bytes of slack at 4.75 bpw overall, and the mix it found is the MoE structure experts hand-tune, discovered from the imatrix alone: expert tensors (ffn_*_exps, each expert rarely active) take IQ2/IQ3, while the always-hot attention paths, router, and embedding hold Q5_K/Q6_K.

Result: PPL 6.91 ± 0.23 held-out — statistically tied with the dense 14B fit (6.85) — at 50.6 tok/s, more than twice the dense model's speed, because only ~3B parameters activate per token. On a 24 GB Mac, this is the strongest configuration we measured: 30B-class quality at interactive speed, filled to the byte.

Project layout

src/gguf.rs      GGUF v3 reader/writer (arbitrary KVs preserved, aligned offsets)
src/fetch.rs     HF repo / URL resolution, shard-aware downloads, auto-imatrix
src/quant.rs     the 8 scale+round encoders + decoders, weighted scale search
src/quant_iq.rs  the 7 IQ codebook encoders + decoders, lattice neighbour search
src/iq_tables.rs script-extracted lattice/codebook tables (generated file)
src/imatrix.rs   legacy + GGUF imatrix parsing, weight sanitization
src/solver.rs    Lagrangian knapsack + greedy top-up
src/vram.rs      GPU probe: Metal working set (macOS), NVML free VRAM (elsewhere)
src/main.rs      CLI, budget model, measurement orchestration (rayon)
src/ui.rs        `shoehorn ui` local web server driving fit as a subprocess
src/ui.html      the page it serves (embedded at compile time)
src/e2e_tests.rs synthetic-GGUF pipeline test + metadata roundtrip
DESIGN.md        the how/why of every decision, in order (D1-D14), gotchas
docs/            decision-graph export (deciduous)

Testing

cargo test covers round-trip encode/decode for every format against RMSE tolerance, a check that imatrix weighting actually shifts the fit toward important columns, solver unit tests (max quality when everything fits, budget respected, infeasibility detected), HF file/shard selection, the rocm-smi parser, a metadata roundtrip over every GGUF value type, and an end-to-end pipeline test that synthesizes a small BF16 GGUF, fits it into a deliberately tight budget, and re-reads the output. All of it runs in CI on macOS, Linux, and Windows.

For measuring what a specific fit cost on your machine, shoehorn eval -m fitted.gguf --baseline source.gguf prints held-out perplexity for both and the delta.

The end-to-end oracle is llama.cpp itself: an independent implementation loads the quantized file on Metal. A single mispacked bit plane produces garbage text, so coherent greedy-decoded output plus near-baseline held-out perplexity is strong evidence the encoders are bit-compatible.

Limitations and roadmap

  • The floor is now IQ2_XXS at ~2.06 bits/weight. IQ1_S/IQ1_M (~1.6 bpw) are not implemented; below IQ2 quality collapses for all but very large models anyway. Expect IQ-heavy encodes to be slower than K-quant ones (lattice neighbour search): the 0.6B test model takes ~25 s instead of ~6 s.
  • The compute buffer is an estimate (real allocation varies by llama.cpp version and flash-attention path; --reserve absorbs the difference) unless you pass --calibrate, which replaces the guess with a measured one at the cost of one extra model load and a partial rewrite.
  • The probe covers Metal (macOS), NVML (NVIDIA on Linux/Windows), and rocm-smi (AMD on Linux, parse-based and not yet exercised on real hardware). Intel GPUs need an explicit --budget. Off-macOS the probe reports the first device's free VRAM, so close the big things before fitting. Multi-GPU boxes are budgeted for one device.
  • Embedded imatrix stats, attention-sink tensors, and other exotic GGUF extras are passed through untouched but not exploited.

Design history

Every significant decision, and several dead ends (a llama-cli interactive-mode hang, a 15-byte "Entry not found" model download that exits 0), is written up as it happened in DESIGN.md.

Troubleshooting

  • no room for weights: your --ctx KV cache plus reserve exceeds the envelope. Lower --ctx, or lower --reserve if you've measured real usage.
  • even the smallest mix exceeds the weight budget: the model can't fit even at ~2 bits/weight. Use a smaller model.
  • Scripted smoke tests hang in llama-cli: recent llama.cpp defaults into conversation mode and spins on a closed stdin even with -no-cnv; use -st (single-turn), or llama-perplexity, which always exits.
  • Output loads but talks nonsense: that is the oracle failing. File a bug with the tensor table from shoehorn plan; a specific format's packing is suspect.
  • no probeable GPU found on Linux: NVML comes with the NVIDIA driver — if nvidia-smi works, the probe should too. On AMD, the probe shells out to rocm-smi; without either, pass --budget.
  • Fits work but shoehorn run can't find llama-server on Windows: PATH changes only apply to terminals opened after editing it; open a new one.
  • Quantized output seems too conservative on a busy GPU: off-macOS the probe uses free VRAM at fit time. Close whatever is holding memory (games, browsers with GPU accel) and re-run, or pass --budget with the card's full size.
  • The web UI shows "a fit is already running": one fit at a time per shoehorn ui process — Stop the current one or wait it out.

Contributors

notactuallytreyanastasio/shoehorn

88

stars

57

commits

Rust

primary language

Aug 19, 2026

updated

notactuallytreyanastasio.github.io/shoehorn/

README

shoehorn

Quantize a BF16 GGUF model with an imatrix so it fits exactly into your available VRAM, then run it with llama.cpp.

shoehorn's site has downloads and a browser-side "what fits your machine?" calculator.

Preset quantizations (Q4_K_M, Q5_K_S, ...) ignore your hardware. Pick one that fits your machine and you either leave hundreds of megabytes of quality unused or find out at load time that it didn't fit after all. shoehorn starts from the memory you actually have, subtracts what inference itself will need (KV cache, compute buffers), and solves a per-tensor mixed-precision assignment whose total size lands within a rounding error of the remainder. Every spare megabyte goes where the importance matrix says it buys the most model quality.

$ shoehorn fit unsloth/Qwen3-4B-GGUF --serve

That one command finds and downloads the BF16 GGUF from Hugging Face, picks up the repo's imatrix (or generates one locally), solves the quant mix for your machine, writes the file, and launches llama-server on it. The pieces are also available separately:

$ shoehorn vram
Apple M4 Pro: 17.76 GiB usable for GPU working set

$ shoehorn quantize -m Qwen3-0.6B-BF16.gguf -i qwen3.imatrix \
    --ctx 4096 --budget 1.75GiB -o fitted.gguf
...
weights: 519.2 MiB of 519.2 MiB budget (99.981% used, 103424 B slack) | overall 7.306 bpw

$ shoehorn run -m fitted.gguf --ctx 4096

The quantizer is implemented from scratch in this repo (Rust, no llama.cpp code linked). The output is standard GGUF v3 that any llama.cpp build, or anything downstream of it, loads directly. llama.cpp handles inference and doubles as an independent correctness oracle.


Contents


Quick start

shoehorn needs llama.cpp installed (inference backend + imatrix generation). For shoehorn itself, grab a prebuilt binary from the releases page or build it with a Rust toolchain as shown below.

macOS (Apple Silicon)

brew install notactuallytreyanastasio/shoehorn/shoehorn   # prebuilt binary + llama.cpp
shoehorn fit unsloth/Qwen3-4B-GGUF --serve

Or from source: brew install llama.cpp then cargo install --path .

Linux (NVIDIA)

# llama.cpp with CUDA — build from source (or grab a -cuda release binary
# from https://github.com/ggml-org/llama.cpp/releases and put it on PATH):
git clone https://github.com/ggml-org/llama.cpp
cmake -S llama.cpp -B llama.cpp/build -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j
sudo cmake --install llama.cpp/build

cargo install --path .
shoehorn fit unsloth/Qwen3-4B-GGUF --serve

The VRAM probe reads the first NVIDIA device's free memory through NVML (ships with the regular driver — nothing extra to install), falling back to rocm-smi for AMD. Intel GPUs aren't probed yet; pass --budget. On AMD, use a Vulkan or ROCm build of llama.cpp.

Windows (NVIDIA)

  1. Install Rust via rustup, then cargo install --path .
  2. Download a -cuda llama.cpp zip from the releases page, unzip it, and add the folder to PATH.
  3. shoehorn fit unsloth/Qwen3-4B-GGUF --serve

Windows support is compile-tested but not yet field-tested. One known gap: auto-generating an imatrix is skipped there (the calibration text is built from man pages), so fit a repo that publishes one — it's picked up automatically — or pass -i yourself.

fit does the whole pipeline: finds the BF16 GGUF in the repo, downloads it to ~/.cache/shoehorn (resumable), picks up or generates an imatrix, solves the mix for your machine, writes <model>-fit.gguf, and serves it. macOS on Apple Silicon and Linux/Windows with an NVIDIA GPU are probed automatically; --budget/--target work anywhere.

Prefer buttons to flags? See the web UI.

Doing the steps by hand instead:

# 1. Get a BF16 (or F16/F32) GGUF of your model.
#    Most HF quant repos (unsloth, bartowski, ggml-org) publish one.

# 2. Generate an importance matrix over calibration text:
llama-imatrix -m model-bf16.gguf -f calibration.txt -o model.imatrix -ngl 99

# 3. Solve and quantize to your machine's actual capacity:
shoehorn quantize -m model-bf16.gguf -i model.imatrix --ctx 8192 -o fitted.gguf

# 4. Serve it (execs llama-server with full GPU offload):
shoehorn run -m fitted.gguf --ctx 8192

shoehorn plan takes the same flags as quantize without -o and prints the solved per-tensor mix without writing anything, so you can preview what a budget implies before spending the encode time. ./demo/run.sh reproduces the full size/quality ladder on a small model in a few minutes.

The web UI

shoehorn ui

opens a local page (default http://127.0.0.1:7788) that drives the same pipeline with no flags: it shows the measured budget for your machine, takes a model name, and runs the fit with the log streamed into the page.

the fit form: your machine's measurement, a model field, and one button

Don't know which model to pick? What fits this machine? ranks Hugging Face's most-downloaded fittable models by what your budget affords:

discovery: ranked fittable models with Use buttons

While the fit runs, the budget renders as a tape measure — weights grow from the left while the fixed costs (KV cache, compute buffers, safety margin) hold the right edge. When it lands, Chat with it serves the result with llama-server and opens its chat page once the model is warm.

a finished fit: the tape-measure gauge at 99.998% used, and a Chat with it button

Step by step:

  1. shoehorn ui from the directory you want the fitted model written to — the output <model>-fit.gguf lands in the current directory, like the CLI. The browser opens by itself (--no-open to suppress, --port to move it off 7788).
  2. Type a model — or click What fits this machine? in the machine card to have discover rank fittable models for your budget, each with a Use button (also reachable directly at /#discover). The field suggests a few known-good repos and live-searches Hugging Face as you type. First fit of a repo downloads the BF16, which for big models is tens of GB — the download resumes if interrupted and is cached in ~/.cache/shoehorn.
  3. Pick how much conversation room you want. More context means a bigger KV cache, which comes straight out of the weight budget.
  4. Fit it to my machine — or Preview the mix first, which solves and shows the tape and the per-type breakdown without writing anything (fit --dry-run underneath), then offers Fit it for real. Downloads render as a progress bar; Stop aborts cleanly, and reloading the page mid-fit picks the run back up.
  5. Measure the cost (optional) runs shoehorn eval on the fit against the original model and reports the perplexity delta right in the card — the honest answer to "how much quality did I give up?"
  6. Chat with it starts llama-server on port 8093 and opens its chat page once the model is loaded. The fitted GGUF is a normal file — anything that speaks GGUF (LM Studio, ollama, llama.cpp on another machine) can use it afterward.
  7. Ctrl-C in the terminal running shoehorn ui shuts everything down.

Finished fits are remembered (in ~/.cache/shoehorn/fits.json) and listed in a "Previous fits" card with one-click Chat buttons, so a model fitted last week doesn't need re-fitting to be served again.

The advanced knobs (KV cache type, budget override, --calibrate) are under "More options". The page runs shoehorn fit as a subprocess and streams its output, so it can't drift from what the CLI does — everything the CLI would have printed is under "Show the work".

How it works

The pipeline runs in five stages.

Probing comes first. On Apple Silicon, "VRAM" is not RAM: Metal will only wire a fraction of unified memory for the GPU. shoehorn asks the Metal device for recommendedMaxWorkingSetSize (17.76 GiB on a 24 GB M4 Pro, about 75%). On other machines it asks NVML — or rocm-smi on AMD — for the first device's free VRAM. --budget overrides the probe, which also lets you quantize for a different machine ("make this fit my friend's 8 GiB M1") or for an artificial envelope.

Next it computes the budget. From the target context length and the model's own GGUF hyperparameters, shoehorn computes the KV cache size exactly and estimates the compute buffer, then subtracts both plus a safety --reserve. What remains is the weight budget. See The budget model.

Then it measures. Every quantizable tensor gets a candidate ladder — the IQ codebook formats (IQ2_XXS up through IQ4_XS), the K-quants, Q8_0, and F16, or the legacy 32-block formats when the row length isn't divisible by 256. Each (tensor, candidate) pair is scored by actually encoding and decoding a sample of rows and accumulating the imatrix-weighted squared error. That is the true end-to-end distortion under the decoder llama.cpp will use. The work parallelizes across all cores; on Qwen3-0.6B the whole measure and solve pass takes 0.7 s.

The solve itself is a multiple-choice knapsack: pick one type per tensor, minimize total weighted error subject to total bytes staying at or under budget. shoehorn uses Lagrangian relaxation (bisect the shadow price of a byte; each tensor independently picks the candidate minimizing err + λ·bytes), then a greedy pass that spends the slack the relaxation leaves behind: repeatedly apply the single-tensor upgrade with the best error reduction per byte that still fits. Utilization in practice exceeds 99.9% of the weight budget.

Finally it writes. Chosen types are re-encoded row-parallel and streamed out as a GGUF v3 with all source metadata preserved and general.file_type set to the dominant type for display purposes. Norms, biases, and anything else 1D stay F32, matching llama.cpp convention: they are tiny and numerically sensitive.

The BF16 source is mmap'd and never fully materialized. Peak memory is roughly one tensor per worker thread, so 30B-class models are fine on a laptop.

The math

Quantized formats represent a block of weights as one or two f16 scale factors plus low-bit integers: symmetric formats decode as x̂ = d·q, asymmetric ones as x̂ = d·q − m, with sub-block scales in the K-quants. The decoder is fixed (it's llama.cpp's dequantization), so the encoder's entire job is choosing d, m, and the integers to minimize error under a weighting that reflects how much each weight matters.

llama-imatrix supplies that weighting. It runs the model over calibration text and accumulates, for each matmul weight, the mean squared activation of each input column. Columns that see large activations amplify their weights' quantization error in the layer's output, so they deserve more of the bit budget. shoehorn uses the element weight

w[j] = imatrix[j] · sqrt(σ² + x[j]²)      σ² = mean square of the row

which is the same shaping ggml's imatrix-aware quantizers use. Tensors without an imatrix entry (typically token_embd, which is never a matmul input) fall back to the activation-agnostic sqrt(σ² + x²).

For symmetric formats the encoder mirrors ggml's make_qx_quants: try 19 candidate grids iscale = −(nmax + 0.1·is)/max for is in [−9, 9]; for each, round every element and evaluate the weighted least-squares objective; keep the grid maximizing (Σ w·x·l)² / Σ w·l², whose optimal scale is Σ w·x·l / Σ w·l². For asymmetric formats it mirrors make_qkx3_quants: over 37 candidate grids, solve the two-parameter weighted regression for scale and offset jointly, clamping the offset so the K-quant d·q − dmin·m convention keeps its positive-min invariant. K-quant super-blocks then quantize the 8 or 16 sub-block scales themselves to 6 or 8 bits and re-round every element against the quantized scales.

Why copy ggml's objective instead of inventing one? The reference quantizers have years of use against this exact decoder, and matching their semantics makes shoehorn's output directly comparable to llama-quantize's. The new work went into the solver.

The solver minimizes the sum over tensors of measured weighted error, with no per-tensor normalization: the imatrix magnitudes already encode relative importance across tensors, and total weighted distortion is exactly what the knapsack should minimize.

Supported formats

formatbits/weightblockencode notes
IQ2_XXS / IQ2_XS / IQ2_S2.06 / 2.31 / 2.56256E8-lattice codebook, neighbour search, 7-bit parity-packed or verbatim signs
IQ3_XXS / IQ3_S3.06 / 3.44256D4-lattice codebook, same search machinery
IQ4_XS4.25256nonlinear 16-value codebook, 6-bit sub-scales
Q4_K4.502568×32 sub-blocks, 6-bit scales+mins, weighted 2-param regression
Q5_K5.50256as Q4_K plus a high-bit plane
Q6_K6.5625616×16 sub-blocks, 8-bit signed scales, weighted grid search
Q8_08.5032absmax scaling (error is negligible; search unnecessary)
IQ4_NL4.5032nonlinear codebook, fallback for rows not divisible by 256
Q4_0 / Q4_14.50 / 5.0032legacy fallback for rows not divisible by 256
Q5_0 / Q5_15.50 / 6.0032legacy fallback, high-bit plane
F16 / BF16 / F3216 / 16 / 32n/apassthrough / conversion

Rows divisible by 256 get the full ladder from IQ2_XXS up to F16; rows divisible only by 32 get {IQ4_NL, Q4_0, Q4_1, Q5_0, Q5_1, Q8_0, F16}. token_embd.weight and output.weight are floored at 4-bit (IQ4_XS): the embedding has no imatrix data, weighted MSE understates LM-head sensitivity, and llama.cpp's own IQ2 presets apply the same guard. Every format has a matching in-crate decoder used for error measurement and tests.

The IQ ports follow ggml-quants.c at the exact commit the installed llama.cpp was built from, including its fudge constants and sign-parity rules; the lattice tables in src/iq_tables.rs are script-extracted from the reference, never hand-typed. ggml keeps two grids per IQ format (a true lattice for the encoder, tuned magnitudes for the decoder) and shoehorn reproduces that asymmetry — see DESIGN.md D11 for the details.

The budget model

weight_budget = usable_vram − kv_cache − compute_est − reserve

kv_cache    = n_layer · ctx · n_kv_heads · (key_len + value_len) · 2   (f16 K+V, exact)
compute_est = ubatch·n_vocab·4 + ubatch·n_embd·32                      (heuristic)
ubatch      = min(512, ctx)

All hyperparameters come from the model's own GGUF metadata (block_count, attention.head_count_kv, attention.key_length, ...), so grouped-query attention and unusual head sizes are handled per model rather than assumed.

Worked example, Qwen3-0.6B at --budget 1.75GiB --ctx 4096:

1.75 GiB − 448 MiB KV (28 layers · 4096 · 8 kv-heads · 256 · 2 B)
         − 313 MiB compute est (512·151936·4 logits dominate)
         − 512 MiB reserve
         = 519.2 MiB for weights → solver fills 519.1 MiB of it

The KV term is exact. The compute term is deliberately rough, since llama.cpp's graph allocation depends on flash-attention availability, batch shape, and version; the --reserve margin absorbs its error plus Metal shader buffers and the host process. If you have a measured number for your setup, --budget and --reserve let you dial it in precisely.

CLI reference

shoehorn fit       <path | owner/repo | url> [-i <imatrix>] [fit flags] [-o out.gguf] [--serve] [--dry-run]
shoehorn plan      -m <bf16.gguf> [-i <imatrix>] [fit flags]
shoehorn quantize  -m <bf16.gguf> [-i <imatrix>] [fit flags] -o <out.gguf>
shoehorn run       -m <model.gguf> [--ctx N] [--kv q8_0] [-- <llama-server args...>]
shoehorn vram
shoehorn ui        [--port 7788] [--no-open]
shoehorn eval      -m <model.gguf> [-f <text>] [--baseline <other.gguf>] [--ctx N]
shoehorn discover  [--ctx N] [--budget <size>] [--scan N]

fit accepts a local path, a Hugging Face repo id (it picks the largest BF16 file — falling back to F16, then F32 — downloads every shard of a split model, and grabs any imatrix in the repo), or a direct URL. Downloads land in ~/.cache/shoehorn and resume if interrupted. Without an imatrix, it generates one with llama-imatrix when the model fits the GPU. --dry-run stops after printing the solved plan — like plan, but with fit's fetching, so it works on repo ids and URLs too.

Fit flags, shared by fit, plan, and quantize:

flagdefaultmeaning
-m, --modelrequiredBF16/F16/F32 source GGUF (already-quantized sources also read, via the in-crate decoders)
-i, --imatrixnoneimatrix file, legacy binary or GGUF-based; omitting it warns and falls back to activation-agnostic weighting
--ctx8192context length the KV budget is computed for
--budgetGPU probetotal memory envelope: 18GiB, 800MB, 4.5G — a bare number like 16 reads as GiB
--targetbudget for a different unified-memory machine by RAM size, e.g. --target 16GB (approximates the macOS working-set limit as 74% of RAM)
--kvf16KV cache type to budget for and run with (f16, q8_0, q4_0); q8_0 roughly halves the KV term, freeing that memory for weights
--reserve512MiB (160MiB with --calibrate)safety margin subtracted from the envelope
--calibrateoffafter writing, load the result in llama.cpp once, measure real KV/compute allocations, re-solve with them, and rewrite — spending the recovered estimate slack on quality (unchanged tensors are reused)
--exact-errorsoffscore every row instead of a 128-row sample per tensor

plan and quantize print the full per-tensor table (shape, chosen type, size, bits/weight), a by-type rollup, budget utilization with the residual slack in bytes, and the projected total VRAM picture at the target context.

run execs llama-server -m <model> -c <ctx> -ngl 99; everything after -- is passed through (--port, --api-key, ...).

vram prints the detected device and its usable GPU memory: Metal's recommended working-set size on macOS, NVML's free VRAM on the first NVIDIA device elsewhere, with a rocm-smi fallback for AMD.

eval wraps llama-perplexity so you can check what a fit actually cost: it measures perplexity on held-out text (by default man pages disjoint from the auto-imatrix calibration set) and, with --baseline, prints the delta against another model — typically the BF16 source:

$ shoehorn eval -m fitted.gguf --baseline model-bf16.gguf
fitted.gguf: PPL 8.0719
model-bf16.gguf: PPL 8.0132
delta: +0.73% vs baseline

discover answers "which model should I even fit?": it scans the most-downloaded GGUF repos on Hugging Face for full-precision sources (the same BF16 > F16 > F32 pick fit uses, so every suggestion is actually fittable), estimates the bits/weight your budget affords each one, and ranks them — biggest model in the best quality tier first. Repos whose "full-precision" file is a speculative-decoding draft companion are detected by cross-checking the file size against the parameter count in the repo name, and flagged instead of recommended. The numbers are estimates; the printed shoehorn fit <repo> --dry-run line gives the exact solve.

ui serves a local web page (and opens it) that drives the whole fit pipeline as a subprocess: pick a model, watch the budget gauge fill, then chat with the result via llama-server. No flags survive contact with it on purpose; the advanced knobs live under "More options".

Imatrix files

Both formats found in the wild are supported and auto-detected:

  • GGUF-based (current llama-imatrix output): tensors named <name>.in_sum2 (per-column sums of squared activations) and <name>.counts; shoehorn divides one by the other.
  • Legacy binary (most published imatrices): n_entries, then per entry name / ncall / nval / f32 values; values are divided by ncall.

Weights are sanitized (non-finite and non-positive entries floored) so a degenerate imatrix can't zero out the fit. For 3D expert tensors (MoE), an imatrix covering ne0 × n_expert is sliced per expert; one covering only ne0 is broadcast.

Calibration text matters less than having an imatrix at all, but it does matter at the margin: auto-generation downloads the community-standard calibration_datav3 mixed corpus (cached), falling back to concatenated man pages offline. Measured on Qwen3-0.6B at an identical 1.75 GiB budget, the mixed corpus cost +1.84% perplexity vs BF16 on neutral prose against +2.81% for man-page calibration — better generalization for one download. The test suite holds out different text from whatever calibrated.

Results

Qwen3-0.6B (596 M params), M4 Pro 24 GB, llama.cpp b10360, 2026-08-13.

Forced into a 1.75 GiB total envelope (weight budget 519.2 MiB, ctx 4096):

BF16shoehorn mix
weights on disk1.13 GiB524.8 MiB (46%)
bits/weight167.306
budget utilizationn/a99.981% (103 KB slack)
held-out perplexity¹14.533 ± 0.55214.623 ± 0.556 (+0.6%)
generation (llama-cli)n/a~210 tok/s

¹ git-rebase man page, text the imatrix never saw.

The solved mix matches what a llama.cpp veteran would hand-tune: ffn_down (known quant-tolerant) at Q4_K, attention projections at Q6_K, Q8_0 and F16 reserved for the tensors whose weighted error per byte is worst. Here it falls out of the optimization, per model, with no hand rules.

Against the machine's real 17.76 GiB budget, the 0.6B model fits at F16 and the solver keeps everything at F16 (6.9% utilization). It degrades nothing when there is no need.

Pushing into IQ territory with tighter budgets (same model, same held-out text):

envelopesizeoverall bpwheld-out PPL
BF16 baseline1.13 GiB1614.53
1.75 GiB525 MB7.3114.62
1.53 GiB300 MB~4.121.53
1.44 GiB207 MB2.84212.7
llama-quantize IQ2_XXS (control)219 MB2.34446.8

The bottom two rows are the differential validation: at comparable size, the solved mix halves the perplexity of llama.cpp's own IQ2_XXS preset, because the knapsack spends bits per tensor instead of uniformly. The absolute numbers also show why sub-3 bpw formats exist for 7B+ models: a 0.6B is severely degraded there no matter who does the quantizing.

Qwen3-14B (the real test)

Source: 29.5 GB BF16 — bigger than this machine's entire GPU working set — with bartowski's published (legacy-format) imatrix.

Against the detected 17.76 GiB budget at ctx 8192, the solver fills 99.998% of the 15.64 GiB weight budget (340 KB slack): a 9.1 bpw mix of Q8_0 (10.7 GiB), F16 (3.6 GiB where the imatrix concentrates importance), and a Q6_K/Q5_K tail. Measure + solve on the 28 GB file: 39 s.

Forced into an 8 GiB total envelope (the "make it fit my friend's 8 GiB M1" case): 100.000% of the 5.88 GiB weight budget used — 28 KB of slack — via a 3.42 bpw mix spanning the entire ladder, IQ2_XXS (87 tensors) through Q6_K. Encode time ~6 min. The result generates correct, fluent text at 23 tok/s and scores PPL 6.85 on the held-out text.

For scale: that is less than half the perplexity of the unquantized 0.6B (14.53) on the same text. Given a fixed memory budget, a big model shoehorned hard beats a small model treated gently — which is exactly the trade the solver exists to make well.

Qwen3-30B-A3B (MoE, split source)

Source: 61 GB BF16 in two llama.cpp-style shards — read as one model — with bartowski's published imatrix, targeting the real 17.76 GiB budget at ctx 8192 with --kv q8_0.

The solver filled the 16.88 GiB weight budget to 0 bytes of slack at 4.75 bpw overall, and the mix it found is the MoE structure experts hand-tune, discovered from the imatrix alone: expert tensors (ffn_*_exps, each expert rarely active) take IQ2/IQ3, while the always-hot attention paths, router, and embedding hold Q5_K/Q6_K.

Result: PPL 6.91 ± 0.23 held-out — statistically tied with the dense 14B fit (6.85) — at 50.6 tok/s, more than twice the dense model's speed, because only ~3B parameters activate per token. On a 24 GB Mac, this is the strongest configuration we measured: 30B-class quality at interactive speed, filled to the byte.

Project layout

src/gguf.rs      GGUF v3 reader/writer (arbitrary KVs preserved, aligned offsets)
src/fetch.rs     HF repo / URL resolution, shard-aware downloads, auto-imatrix
src/quant.rs     the 8 scale+round encoders + decoders, weighted scale search
src/quant_iq.rs  the 7 IQ codebook encoders + decoders, lattice neighbour search
src/iq_tables.rs script-extracted lattice/codebook tables (generated file)
src/imatrix.rs   legacy + GGUF imatrix parsing, weight sanitization
src/solver.rs    Lagrangian knapsack + greedy top-up
src/vram.rs      GPU probe: Metal working set (macOS), NVML free VRAM (elsewhere)
src/main.rs      CLI, budget model, measurement orchestration (rayon)
src/ui.rs        `shoehorn ui` local web server driving fit as a subprocess
src/ui.html      the page it serves (embedded at compile time)
src/e2e_tests.rs synthetic-GGUF pipeline test + metadata roundtrip
DESIGN.md        the how/why of every decision, in order (D1-D14), gotchas
docs/            decision-graph export (deciduous)

Testing

cargo test covers round-trip encode/decode for every format against RMSE tolerance, a check that imatrix weighting actually shifts the fit toward important columns, solver unit tests (max quality when everything fits, budget respected, infeasibility detected), HF file/shard selection, the rocm-smi parser, a metadata roundtrip over every GGUF value type, and an end-to-end pipeline test that synthesizes a small BF16 GGUF, fits it into a deliberately tight budget, and re-reads the output. All of it runs in CI on macOS, Linux, and Windows.

For measuring what a specific fit cost on your machine, shoehorn eval -m fitted.gguf --baseline source.gguf prints held-out perplexity for both and the delta.

The end-to-end oracle is llama.cpp itself: an independent implementation loads the quantized file on Metal. A single mispacked bit plane produces garbage text, so coherent greedy-decoded output plus near-baseline held-out perplexity is strong evidence the encoders are bit-compatible.

Limitations and roadmap

  • The floor is now IQ2_XXS at ~2.06 bits/weight. IQ1_S/IQ1_M (~1.6 bpw) are not implemented; below IQ2 quality collapses for all but very large models anyway. Expect IQ-heavy encodes to be slower than K-quant ones (lattice neighbour search): the 0.6B test model takes ~25 s instead of ~6 s.
  • The compute buffer is an estimate (real allocation varies by llama.cpp version and flash-attention path; --reserve absorbs the difference) unless you pass --calibrate, which replaces the guess with a measured one at the cost of one extra model load and a partial rewrite.
  • The probe covers Metal (macOS), NVML (NVIDIA on Linux/Windows), and rocm-smi (AMD on Linux, parse-based and not yet exercised on real hardware). Intel GPUs need an explicit --budget. Off-macOS the probe reports the first device's free VRAM, so close the big things before fitting. Multi-GPU boxes are budgeted for one device.
  • Embedded imatrix stats, attention-sink tensors, and other exotic GGUF extras are passed through untouched but not exploited.

Design history

Every significant decision, and several dead ends (a llama-cli interactive-mode hang, a 15-byte "Entry not found" model download that exits 0), is written up as it happened in DESIGN.md.

Troubleshooting

  • no room for weights: your --ctx KV cache plus reserve exceeds the envelope. Lower --ctx, or lower --reserve if you've measured real usage.
  • even the smallest mix exceeds the weight budget: the model can't fit even at ~2 bits/weight. Use a smaller model.
  • Scripted smoke tests hang in llama-cli: recent llama.cpp defaults into conversation mode and spins on a closed stdin even with -no-cnv; use -st (single-turn), or llama-perplexity, which always exits.
  • Output loads but talks nonsense: that is the oracle failing. File a bug with the tensor table from shoehorn plan; a specific format's packing is suspect.
  • no probeable GPU found on Linux: NVML comes with the NVIDIA driver — if nvidia-smi works, the probe should too. On AMD, the probe shells out to rocm-smi; without either, pass --budget.
  • Fits work but shoehorn run can't find llama-server on Windows: PATH changes only apply to terminals opened after editing it; open a new one.
  • Quantized output seems too conservative on a busy GPU: off-macOS the probe uses free VRAM at fit time. Close whatever is holding memory (games, browsers with GPU accel) and re-run, or pass --budget with the card's full size.
  • The web UI shows "a fit is already running": one fit at a time per shoehorn ui process — Stop the current one or wait it out.

Contributors

Languages

Rust

90.9%

HTML

9.1%