News from the IT community on October 4th: The Qwen3.8 ‑27B is an AI model with decent performance, but this is only possible if the device has sufficient video memory and system memory. The MacBook Pro laptop, which is equipped with a M4 Pro chip, has only 24GB of unified memory. As a result, the pre-loading speed of this AI model is limited. According to reports from Wccftech, a user connected a iPhone 17 Pro Max through the USB ‑C interface to this portable Mac computer and, with the help of custom software, managed to increase the pre-loading performance of this AI model by up to 44%.

IT's notice indicates that in order to break through the 24GB memory limit of this Mac, user "u/ StayLameBro" adopted two methods: splitting the computational tasks and distributing them to both the local Mac machine and 17 Pro Max units for collaborative processing. For pre-filling acceleration, MacBook Pro is responsible for handling the first 40 layers of each batch of 256 token. The data stream of activation values is then transmitted to the mobile device; simultaneously, iPhone utilizes A19 Pro's GPU to run the 41st to 64th layers, while Mac has already begun processing the next batch of data. Thanks to the operations performed by GPU, the running speed on the mobile device has increased by about 2.4 times compared to when GPU is not used.
In the entire solution, the neural network engine is not left idle either. Every 16K length of old context is compiled into a set of neural network engine models. With a context length of 140K, compared to using only GPU, the writing time for a single token is reduced from 279 milliseconds to 176 milliseconds. Taking the task of pre-filling and saving a session with a file containing 2000 token as an example: when the context size is set to 8K, the speed using only the local Mac is 132 token per second; after adding an external iPhone, the speed increases to 177 token per second, representing a 35% performance improvement. When the context size is set to 16K, the speed rises from 109 token per second to 157 token per second, which is a significant increase of 44%.
When the context window is set to 32K, the pre-filling speed increases from 101 token per second to 130 token per second, representing a 29% improvement in performance. By integrating iPhone with 17 Pro Max on the M4 Pro version of MacBook Pro, using this self-made software should theoretically result in an exceptionally fast pre-filling speed. Unfortunately, as this user Reddit pointed out, reality is not so rosy; there are still several limitations to this solution. For example, in scenarios within 64K, the phone does not accelerate text generation, and the entire task of text generation is still handled by Mac.
It is also worth mentioning that the A20 Pro chip driving iPhone, Pro, and iPhone (along with Pro Max) is expected to offer greater performance headroom. On one hand, it has stronger overall performance; on the other hand, it is equipped with a dual-16-core neural network engine. It is claimed that for tasks specifically adapted to neural network processors, this engine's data throughput capability even exceeds that of the system-on-chip's built-in 7-core GPU. This self-developed software is now open-source and can be downloaded from GitHub; the project is named “backburner”.
Even if it is merely considered as an experiment, it is evident that the improvement in pre-fill speed brings considerable benefits. There is certainly room to further explore performance potential in the future; however, this may also mean that devices like iPhone could potentially face a shortage in supply in the same way that memory and solid-state drives have in the past.











