DFlash Team Advances into Mac; 27B Model Achieves 144 Tokens per Second
2026-09-19 16:25:47
According to CoinMeta, the DFlash team announced their entry into the Mac platform, where the 27B model achieved a peak performance of 144 token /s on M5 Max MacBook Pro. The team's open-source local inference engine, Splash, has been integrated into LM Studio version 0.4.25 and supports Qwen3.8-27B and Qwen3.6-35B-a3b models. By processing multiple token in parallel, the overall performance has been improved by reducing the computation required for individual generation. During horizontal testing with 48GB of M5 Pro, the single-threaded short context speed reached 74 token /s, and with four-way concurrency, the total throughput reached 170 token /s, which is 3.9 times that of the second-place competitor. Splash is open-source and not tied to any particular LM Studio; it requires a M3 or newer Mac, macOS with a version above 26.4, and at least 36GB of unified memory.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.