Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable
Objective-C++
12
7 commits
updated Oct 2, 2026
Plug your iPhone into your MacBook with a 10 Gb/s USB-C cable and it helps run Qwen3.8-27B locally:
The engine is a llama.cpp fork (llama.cpp/, StayLameBro/backburner-llama.cpp)
with its own Mac kernels (SME2, Metal fusions, DFlash2 speculative decoding). Those speed things up on the Mac alone too; the
numbers below keep the two apart.
Tested on a MacBook Pro M4 Pro (24 GB) with iPhone 17 Pro Max (A19 Pro) and iPhone 16 Pro Max (A18 Pro) phones.

Measured on 2026-10-01: Qwen3.8-27B IQ4_XS, MacBook Pro M4 Pro 24 GB, iPhone 17 Pro Max over USB-C. Raw rows are in
bench/results/*.jsonl; the scripts that produced them are in bench/.
A 2,000-token file or tool result read into a saved agent session (bench/turn-bench.py, two reads per depth):
| Context already in the session | Mac alone | Mac + iPhone | |
|---|---|---|---|
| 16k | 109 tok/s (18.8 s) | 157 tok/s (13.1 s) | +44%, 31% less waiting |
| 32k | 101 tok/s (20.3 s) | 130 tok/s (15.8 s) | +29%, 22% less waiting |
| 48k | 87 tok/s (23.5 s) | 113 tok/s (18.1 s) | +30%, 23% less waiting |
A new omp session (omp's system prompt, project notes and 12 tools, 26,849 tokens, read cold; bench/session-bench.py):
| stock llama.cpp | this fork, Mac alone | this fork + iPhone | |
|---|---|---|---|
| first answer | 245 s | 228 s | 168 s |
| later turns (1.3-1.9k-token tool results) | 17.9 s | 19.2 s | 14.5 s |
After the first time, the SSD prompt cache (scripts/proxy.py) restores that 27k-token start in 0.3-5 s.
Past 64k the Mac-alone config switches to 4-bit context: 128k at 8-bit measured ~0.3 GB over the GPU's 20 GB memory limit with the draft model loaded (2026-09-23); 8-bit between 64k and 128k was not tested. With the iPhone it stays 8-bit. There the phone changes jobs: instead of running layers 41-64 it computes attention over the old keys it holds, while the Mac runs all 64 layers (see "Who does what"). Prefill past 64k is a little faster with the phone and at higher precision: 67-73 tok/s with the iPhone at 8-bit vs 59-68 tok/s Mac alone at 4-bit (64k-96k). At 128k with the iPhone: 3 of 3 planted facts recalled (positions 1.5k, 40k, 100k), phone thermal state nominal.
The phone does not change writing speed below 64k; the fork's kernels and draft model do. Past 64k the phone's GPU and Neural Engine compute attention over the old keys for every generated token:
| context | tok/s | |
|---|---|---|
| stock llama.cpp (Homebrew), Mac | 27-33k | 11.3 |
| this fork, Mac alone | 27-33k | 25.0 |
| this fork + iPhone | 27-33k | 25.1 |
| this fork + iPhone, a real omp session (36 requests) | under 16k / 16-32k / 32-49k | 29.8 / 27.5 / 24.3 (medians) |
| this fork + iPhone | 128k | 12.6 (greedy, 256 tokens) |
Medium thinking, omp's request fields, the server's default sampling. Mac-alone writing speed at 128k (4-bit) was not measured on the same day, so there is no head-to-head number for it here.
| 8-bit context | |
|---|---|
| Mac alone (24 GB) | 64k measured (128k only fits with 4-bit) |
| Mac + iPhone 17 Pro Max | 196k-229k by the phone's free memory (sized at startup); tested to 128k (140k at 4-bit) |
With the defaults and one iPhone 17 Pro Max:
| Mac GPU | Mac CPU (SME2) | iPhone GPU (matrix units) | iPhone Neural Engine | |
|---|---|---|---|---|
| Prefill, context up to 64k | layers 1-40 | ~30% of each big matmul | layers 41-64 (split prefill) | - |
| Prefill past 64k | all 64 layers | ~30% of each big matmul | attention over the old keys it holds | - (builds its pages in the background) |
| Writing, up to 64k | everything, plus the draft model | oldest keys of each attention layer past 40k | - | - |
| Writing past 64k | everything else | same | attention over the old keys | part of that attention |
The Mac's own Neural Engine is not used: it shares the Mac's memory bandwidth and slowed decode by 26% when running
(docs/ANE.md).
llama.cpp/src/llama-split.cpp; the phone's tail server is in ios/Backburner). The Mac runs layers
1-40 of each 256-token ubatch and streams the residual to the phone, which runs layers 41-64 on its GPU while the Mac starts
the next ubatch. The phone keeps a mirror of its layers' KV rows and recurrent state; only new rows cross the cable. The last
ubatch of each batch runs on the Mac so outputs stay local. On the A19 Pro the phone's layers use the GPU's matrix units
(Metal 4 tensor ops): 2.4x faster than the same phone with them off.phone-attn/, protocol in phone-attn/phone-attn.h). Past the Mac's 64k cells, the oldest KV pages
(4,096 keys each) move to the phone. Each attention step sends Q to the phone and merges its partial result (O, max, sum)
with the Mac's. The phone computes it with a matrix-unit kernel on its GPU (phone-attn/pa-metal.mm). While the phone
holds keys, 512-token ubatches run as two staggered halves so Mac and phone overlap (140k: 58 -> 68 tok/s prefill). Two
phones can share the old pages (docs/TWO-PHONES.md).phone-attn/pa-ane.mm). Old keys never change, so each 16,384-key page of a
layer is compiled into a Neural Engine model with the keys and values as its weights. While writing, the Neural Engine
takes part of each old-key attention call and the GPU the rest: at 140k, 279 -> 176 ms per generated token. scripts/serve.sh
puts the page template on the phone the first time it sees it (scripts/phone-ane.sh) and prints "ANE pages on".scripts/proxy.py): a known system prompt is restored from disk instead of re-read.--load-mode none) so macOS can't page it out; the token-embedding table is read
from the mapped file, the scheduler's worst-case buffer is mapped on demand, and freed heap goes back to macOS. Server
footprint after a read: 19.7 -> 18.7 GB (2026-10-01).docs/ANE.md.-np 1).You need an Apple Silicon Mac (tested: M4 Pro, 24 GB), an iPhone 15 Pro or newer (tested: 17 Pro Max, 16 Pro Max), a 10 Gb/s USB-C cable (the cable in the iPhone box is USB 2 and too slow), and Xcode with an Apple developer team id for the app.
git clone --recursive https://github.com/StayLameBro/backburner && cd backburner
# 1. the Mac engine
cmake -S llama.cpp -B llama.cpp/build-metal -DCMAKE_BUILD_TYPE=Release
cmake --build llama.cpp/build-metal --target llama-server llama-quantize -j
# 2. the models: a Qwen3.8-27B IQ4_XS GGUF at ~/Models/Qwen3.8-27B-IQ4_XS.gguf, then the draft model
huggingface-cli download z-lab/Qwen3.8-27B-DFlash2 --local-dir ~/Models/qwen38-27b-dflash2
scripts/make-drafter.sh ~/Models/qwen38-27b-dflash2 ~/Models/dflash2-v2-q4km-self16.gguf
# 3. the iPhone app (phone plugged in and unlocked). Keep DEVELOPMENT_TEAM exported: serve.sh uses it to relaunch the app.
export DEVELOPMENT_TEAM=<your team id>
UDID=<your iPhone's UDID> scripts/build-iphone.sh
pip3 install coremltools # serve.sh builds the phone's Neural Engine page model with it, once
# 4. the phone's half of the model (layers 41-64, ~5.1 GB), copied over the cable
python3 scripts/split-gguf.py ~/Models/Qwen3.8-27B-IQ4_XS.gguf ~/Models/tail-iq4xs-L40-nohead.gguf -L 40
scripts/phone-tail.sh L40
# 5. after every reboot: let the GPU keep the model wired (macOS resets this limit)
sudo sysctl iogpu.wired_limit_mb=20480
# 6. run: OpenAI-compatible on :8080, uses the iPhone when it's plugged in with Backburner open
scripts/serve.sh
PHONE=0 scripts/serve.sh # the Mac alone
The first serve.sh with the phone pushes the Neural Engine page model to it and relaunches the app (about a minute). After
that the startup line should read split prefill on, remote KV on and ANE pages on.
scripts/serve.sh documents each setting next to the measurement that chose it.
bench/turn-bench.py --build # once: a saved session at 16k / 32k / 48k (Mac alone)
bench/turn-bench.py --config mac # read 2,000-token files into each saved session
bench/turn-bench.py --config phone
bench/session-bench.py --config stock|fork-mac|fork-phone # an omp-shaped session, ~5 min each
bench/long-bench.py # past 64k (long: cold reads to 128k)
Pre-release. Next: the phone's layers seeing the keys it holds (split prefill past 64k), a second phone in the prefill chain, an App Store build, and upstreaming what makes sense to llama.cpp. Built with a lot of help from Claude Opus 5.5.
MIT license (llama.cpp keeps its own MIT license).
Objective-C++
28.6%
Python
26.9%
C++
13.7%
Swift
12.4%
C
10.2%
Shell
7.7%
Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable
Objective-C++
12
7 commits
updated Oct 2, 2026
Plug your iPhone into your MacBook with a 10 Gb/s USB-C cable and it helps run Qwen3.8-27B locally:
The engine is a llama.cpp fork (llama.cpp/, StayLameBro/backburner-llama.cpp)
with its own Mac kernels (SME2, Metal fusions, DFlash2 speculative decoding). Those speed things up on the Mac alone too; the
numbers below keep the two apart.
Tested on a MacBook Pro M4 Pro (24 GB) with iPhone 17 Pro Max (A19 Pro) and iPhone 16 Pro Max (A18 Pro) phones.

Measured on 2026-10-01: Qwen3.8-27B IQ4_XS, MacBook Pro M4 Pro 24 GB, iPhone 17 Pro Max over USB-C. Raw rows are in
bench/results/*.jsonl; the scripts that produced them are in bench/.
A 2,000-token file or tool result read into a saved agent session (bench/turn-bench.py, two reads per depth):
| Context already in the session | Mac alone | Mac + iPhone | |
|---|---|---|---|
| 16k | 109 tok/s (18.8 s) | 157 tok/s (13.1 s) | +44%, 31% less waiting |
| 32k | 101 tok/s (20.3 s) | 130 tok/s (15.8 s) | +29%, 22% less waiting |
| 48k | 87 tok/s (23.5 s) | 113 tok/s (18.1 s) | +30%, 23% less waiting |
A new omp session (omp's system prompt, project notes and 12 tools, 26,849 tokens, read cold; bench/session-bench.py):
| stock llama.cpp | this fork, Mac alone | this fork + iPhone | |
|---|---|---|---|
| first answer | 245 s | 228 s | 168 s |
| later turns (1.3-1.9k-token tool results) | 17.9 s | 19.2 s | 14.5 s |
After the first time, the SSD prompt cache (scripts/proxy.py) restores that 27k-token start in 0.3-5 s.
Past 64k the Mac-alone config switches to 4-bit context: 128k at 8-bit measured ~0.3 GB over the GPU's 20 GB memory limit with the draft model loaded (2026-09-23); 8-bit between 64k and 128k was not tested. With the iPhone it stays 8-bit. There the phone changes jobs: instead of running layers 41-64 it computes attention over the old keys it holds, while the Mac runs all 64 layers (see "Who does what"). Prefill past 64k is a little faster with the phone and at higher precision: 67-73 tok/s with the iPhone at 8-bit vs 59-68 tok/s Mac alone at 4-bit (64k-96k). At 128k with the iPhone: 3 of 3 planted facts recalled (positions 1.5k, 40k, 100k), phone thermal state nominal.
The phone does not change writing speed below 64k; the fork's kernels and draft model do. Past 64k the phone's GPU and Neural Engine compute attention over the old keys for every generated token:
| context | tok/s | |
|---|---|---|
| stock llama.cpp (Homebrew), Mac | 27-33k | 11.3 |
| this fork, Mac alone | 27-33k | 25.0 |
| this fork + iPhone | 27-33k | 25.1 |
| this fork + iPhone, a real omp session (36 requests) | under 16k / 16-32k / 32-49k | 29.8 / 27.5 / 24.3 (medians) |
| this fork + iPhone | 128k | 12.6 (greedy, 256 tokens) |
Medium thinking, omp's request fields, the server's default sampling. Mac-alone writing speed at 128k (4-bit) was not measured on the same day, so there is no head-to-head number for it here.
| 8-bit context | |
|---|---|
| Mac alone (24 GB) | 64k measured (128k only fits with 4-bit) |
| Mac + iPhone 17 Pro Max | 196k-229k by the phone's free memory (sized at startup); tested to 128k (140k at 4-bit) |
With the defaults and one iPhone 17 Pro Max:
| Mac GPU | Mac CPU (SME2) | iPhone GPU (matrix units) | iPhone Neural Engine | |
|---|---|---|---|---|
| Prefill, context up to 64k | layers 1-40 | ~30% of each big matmul | layers 41-64 (split prefill) | - |
| Prefill past 64k | all 64 layers | ~30% of each big matmul | attention over the old keys it holds | - (builds its pages in the background) |
| Writing, up to 64k | everything, plus the draft model | oldest keys of each attention layer past 40k | - | - |
| Writing past 64k | everything else | same | attention over the old keys | part of that attention |
The Mac's own Neural Engine is not used: it shares the Mac's memory bandwidth and slowed decode by 26% when running
(docs/ANE.md).
llama.cpp/src/llama-split.cpp; the phone's tail server is in ios/Backburner). The Mac runs layers
1-40 of each 256-token ubatch and streams the residual to the phone, which runs layers 41-64 on its GPU while the Mac starts
the next ubatch. The phone keeps a mirror of its layers' KV rows and recurrent state; only new rows cross the cable. The last
ubatch of each batch runs on the Mac so outputs stay local. On the A19 Pro the phone's layers use the GPU's matrix units
(Metal 4 tensor ops): 2.4x faster than the same phone with them off.phone-attn/, protocol in phone-attn/phone-attn.h). Past the Mac's 64k cells, the oldest KV pages
(4,096 keys each) move to the phone. Each attention step sends Q to the phone and merges its partial result (O, max, sum)
with the Mac's. The phone computes it with a matrix-unit kernel on its GPU (phone-attn/pa-metal.mm). While the phone
holds keys, 512-token ubatches run as two staggered halves so Mac and phone overlap (140k: 58 -> 68 tok/s prefill). Two
phones can share the old pages (docs/TWO-PHONES.md).phone-attn/pa-ane.mm). Old keys never change, so each 16,384-key page of a
layer is compiled into a Neural Engine model with the keys and values as its weights. While writing, the Neural Engine
takes part of each old-key attention call and the GPU the rest: at 140k, 279 -> 176 ms per generated token. scripts/serve.sh
puts the page template on the phone the first time it sees it (scripts/phone-ane.sh) and prints "ANE pages on".scripts/proxy.py): a known system prompt is restored from disk instead of re-read.--load-mode none) so macOS can't page it out; the token-embedding table is read
from the mapped file, the scheduler's worst-case buffer is mapped on demand, and freed heap goes back to macOS. Server
footprint after a read: 19.7 -> 18.7 GB (2026-10-01).docs/ANE.md.-np 1).You need an Apple Silicon Mac (tested: M4 Pro, 24 GB), an iPhone 15 Pro or newer (tested: 17 Pro Max, 16 Pro Max), a 10 Gb/s USB-C cable (the cable in the iPhone box is USB 2 and too slow), and Xcode with an Apple developer team id for the app.
git clone --recursive https://github.com/StayLameBro/backburner && cd backburner
# 1. the Mac engine
cmake -S llama.cpp -B llama.cpp/build-metal -DCMAKE_BUILD_TYPE=Release
cmake --build llama.cpp/build-metal --target llama-server llama-quantize -j
# 2. the models: a Qwen3.8-27B IQ4_XS GGUF at ~/Models/Qwen3.8-27B-IQ4_XS.gguf, then the draft model
huggingface-cli download z-lab/Qwen3.8-27B-DFlash2 --local-dir ~/Models/qwen38-27b-dflash2
scripts/make-drafter.sh ~/Models/qwen38-27b-dflash2 ~/Models/dflash2-v2-q4km-self16.gguf
# 3. the iPhone app (phone plugged in and unlocked). Keep DEVELOPMENT_TEAM exported: serve.sh uses it to relaunch the app.
export DEVELOPMENT_TEAM=<your team id>
UDID=<your iPhone's UDID> scripts/build-iphone.sh
pip3 install coremltools # serve.sh builds the phone's Neural Engine page model with it, once
# 4. the phone's half of the model (layers 41-64, ~5.1 GB), copied over the cable
python3 scripts/split-gguf.py ~/Models/Qwen3.8-27B-IQ4_XS.gguf ~/Models/tail-iq4xs-L40-nohead.gguf -L 40
scripts/phone-tail.sh L40
# 5. after every reboot: let the GPU keep the model wired (macOS resets this limit)
sudo sysctl iogpu.wired_limit_mb=20480
# 6. run: OpenAI-compatible on :8080, uses the iPhone when it's plugged in with Backburner open
scripts/serve.sh
PHONE=0 scripts/serve.sh # the Mac alone
The first serve.sh with the phone pushes the Neural Engine page model to it and relaunches the app (about a minute). After
that the startup line should read split prefill on, remote KV on and ANE pages on.
scripts/serve.sh documents each setting next to the measurement that chose it.
bench/turn-bench.py --build # once: a saved session at 16k / 32k / 48k (Mac alone)
bench/turn-bench.py --config mac # read 2,000-token files into each saved session
bench/turn-bench.py --config phone
bench/session-bench.py --config stock|fork-mac|fork-phone # an omp-shaped session, ~5 min each
bench/long-bench.py # past 64k (long: cold reads to 128k)
Pre-release. Next: the phone's layers seeing the keys it holds (split prefill past 64k), a second phone in the prefill chain, an App Store build, and upstreaming what makes sense to llama.cpp. Built with a lot of help from Claude Opus 5.5.
MIT license (llama.cpp keeps its own MIT license).
Objective-C++
28.6%
Python
26.9%
C++
13.7%
Swift
12.4%
C
10.2%
Shell
7.7%