Skip to content
AI DEEP 4 sources · 4 min · cluster 3 · updated 22:03 UTC

M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents - MacStories

Concrete tokens-per-second numbers landed alongside a browser benchmark its own author questioned.

TL;DR

  1. An r/LocalLLaMA review called the M5 Ultra Mac Studio the dream Mac for local AI agents, while another thread argued 16GB — and often 12GB — is the maximum VRAM most people will reasonably have.
  2. A thread reported running Qwen3.8-27B in native 8-bit at 37-55 tokens per second on Apple Silicon with the Splash engine, extended to Q8 and 256k context.
  3. On Hacker News, a vLLM deep dive and a 3,000-tokens-per-second browser benchmark ran alongside a claim of 94% on AIME with a 1B-parameter model.

An r/LocalLLaMA review called the M5 Ultra Mac Studio the dream Mac for local AI agents, while another thread argued that 16GB — and often 12GB — is the maximum VRAM most people will ever reasonably have. [1] [2]

Engine work kept pace: a thread reported running Qwen3.8-27B in native 8-bit at 37-55 tokens per second on Apple Silicon with the Splash engine, extended to Q8 and 256k context. [3]

On Hacker News, a deep dive covered vLLM architecture, memory and benchmarks; a separate post asked whether a 3,000-tokens-per-second small-model benchmark running in a browser is real or a blooper; and another claimed 94% on AIME with a 1B-parameter model. [4] [5] [6]

Why it matters

Every credible single-device result raises the floor for what runs outside a data center, which changes inference economics and, over time, demand for memory and accelerators.

Editor's note

Throughput and benchmark claims are community reports and vendor posts, not independent measurements; the 3,000 tok/s result is explicitly questioned by its own author.

Type to search

↑↓ navigate ↵ open esc close