M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents - MacStories
Concrete tokens-per-second numbers landed alongside a browser benchmark its own author questioned.
TL;DR
- An r/LocalLLaMA review called the M5 Ultra Mac Studio the dream Mac for local AI agents, while another thread argued 16GB — and often 12GB — is the maximum VRAM most people will reasonably have.
- A thread reported running Qwen3.8-27B in native 8-bit at 37-55 tokens per second on Apple Silicon with the Splash engine, extended to Q8 and 256k context.
- On Hacker News, a vLLM deep dive and a 3,000-tokens-per-second browser benchmark ran alongside a claim of 94% on AIME with a 1B-parameter model.
An r/LocalLLaMA review called the M5 Ultra Mac Studio the dream Mac for local AI agents, while another thread argued that 16GB — and often 12GB — is the maximum VRAM most people will ever reasonably have. [1] [2]
Engine work kept pace: a thread reported running Qwen3.8-27B in native 8-bit at 37-55 tokens per second on Apple Silicon with the Splash engine, extended to Q8 and 256k context. [3]
On Hacker News, a deep dive covered vLLM architecture, memory and benchmarks; a separate post asked whether a 3,000-tokens-per-second small-model benchmark running in a browser is real or a blooper; and another claimed 94% on AIME with a 1B-parameter model. [4] [5] [6]
Why it matters
Every credible single-device result raises the floor for what runs outside a data center, which changes inference economics and, over time, demand for memory and accelerators.
Editor's note
Throughput and benchmark claims are community reports and vendor posts, not independent measurements; the 3,000 tok/s result is explicitly questioned by its own author.