LM Studio announced a day-0 partnership with Inco to bring the Splash inference engine into LM Studio. They cite up to about 144 tok/s for Qwen3.8-27B on M5 Max, with setup docs at lmstudio.ai/blog/splash-engine. Splash was already open-sourced by Inco; this puts the Apple-silicon-optimized stack inside LM Studio so local desktop users do not have to assemble the engine separately.
Key Takeaways
- ✓LM Studio integrates Inco Splash on day 0 for one-click local desktop use.
- ✓Claims ~144 tok/s for Qwen3.8-27B on M5 Max, with an official setup post.
- ✓Moves Apple-silicon-optimized inference into a mainstream local client distribution path.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.