Plug an iPhone into your Mac to help run a 27B model

Plug an iPhone into your Mac to help run a 27B model

Your plugged-in iPhone can now act as a co-processor to speed up local AI models on your Mac.

Open-source project Backburner offloads prompt reading and context memory to an iPhone connected over a 10 Gb/s USB-C cable. In benchmarks running Qwen3.8-27B on a 24 GB M4 Pro MacBook, pairing an iPhone cut prompt reading wait times by up to 31% and boosted prefill speeds by 29% to 44% at 16k–48k context depths.

Why it matters: Running a 27-billion-parameter model on a 24 GB Mac usually forces a choice between slow prompt reading and shrinking your context window. By offloading the model's upper layers and older attention keys to the phone's GPU, you get higher precision and larger context sizes without upgrading your Mac's hardware.

Here's the gist: the Mac runs layers 1 through 40, streams data over USB-C, and the phone computes layers 41 through 64 on its GPU. Output answers stay token-identical to running on the Mac alone, and the app builds on an open-source llama.cpp fork that also runs on M-series iPads.

Turns out that phone sitting on your desk makes a pretty decent accelerator card.