A Reddit user turned an iPhone 17 Pro Max into a second processor to run the Qwen 3.8 27B AI model on a MacBook Pro with an M4 Pro chip and 24 GB of unified memory. According to the account republished by Wccftech and IT Home on October 4, 2026, prefill — the stage where the model reads the prompt before answering — was up to 44% faster at a 16,000-token context.
What happened?
The experiment comes from u/StayLameBro in the LocalLLaMA community. The MacBook alone has enough memory to load the 27-billion-parameter model, but prefill is limited. The homemade setup connects the iPhone 17 Pro Max over USB-C and splits the work with custom software.
For each 256-token batch, the MacBook handles layers 1 to 40 and streams activations to the phone. The iPhone runs layers 41 to 64 on the A19 Pro GPU while the Mac starts the next batch. The tester says the phone GPU makes that slice about 2.4 times faster than it would be without it.
The published figures, for prefilling a roughly 2,000-token file and saving the session, are:
- 8K context: from 132 to 177 tokens per second, up 35%.
- 16K context: from 109 to 157 tokens per second, up 44%.
- 32K context: from 101 to 130 tokens per second, up 29%.
Why it matters
Running large models on-device avoids sending text to the cloud, but memory is the bottleneck. An entry M4 Pro MacBook Pro with 24 GB already feels the limit on Qwen 3.8 27B. The iPhone 17 Pro Max test shows a recent phone can act as an auxiliary accelerator, not only as a client.
This is not an official Apple feature. It is a user experiment, measured by the author and picked up by specialist outlets. The gain shows up in prefill, not necessarily in token-by-token generation after the answer starts. Prefill is the heavy part when the prompt is long: the model must read everything before writing the first word.
What changes in practice
For local AI users, the most useful number is 16K: nearly 50 extra tokens per second just by plugging in the phone. That shortens the wait when opening a long document. The cost is custom software, a cable, and a phone tied up during inference.
There is no independent lab confirmation and no ready-made App Store tool. The path shown is an enthusiast one: split layers between the Mac’s unified memory and the A19 Pro GPU. Sources: IT Home and Wccftech.
Image credit: 茅野ふたば — Source: Wikimedia Commons
By GeekikiBot