/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Mac Studio (M5 Ultra) with 256 GB of RAM review: a dream machine to run local AI agents and a massive leap over M3 Ultra for prompt processing and generation

Two models (Flash-Next and GLM-5.3-Flash)  — 4k, 8k, 16k, 64k, 128k, and 256k prompts  — M5 Ultra vs M3 Ultra vs 5090  — 4/5/6/8-bit quants  —  I created all the possible charts and visualizations I could think of. …

MacStories Federico Viticci

Context & Ripple Effects

Apple positioned the M5 Ultra Mac Studio as a high-memory AI workstation in August, advertising up to 512GB of unified memory and substantially faster AI performance. A 2025 [[a:883438|M3 Ultra review demonstrated the appeal of running a quantized 671B-parameter model locally]], establishing memory capacity as the key enabler for workloads that do not fit on typical desktops.

This review tests the 256GB configuration across two models, prompt lengths and quantization levels, comparing it with the M3 Ultra and Nvidia’s RTX 5090. The comparison gives prospective local-AI buyers evidence about prompt-processing and generation performance rather than treating peak benchmark scores alone as the buying criterion.

First-order effects

  • Developers and power users running local agents gain a tested 256GB Mac Studio option for larger context windows and quantized models, with the review finding a major prompt-processing and generation improvement over M3 Ultra.
  • Apple’s M5 Ultra configuration becomes a more clearly differentiated local-inference product, while buyers must weigh its performance case against the lower-cost M5 Max configuration highlighted by a separate same-day review.

Second-order effects

  • Nvidia RTX 5090 systems become a more explicit alternative for local-model buyers, since the review places Apple’s unified-memory workstation and a discrete-GPU setup in the same model, prompt-length and quantization comparisons.
  • The decision shifts toward memory capacity and sustained inference behavior: model users choosing among M5 Max, M5 Ultra and GPU-based PCs have more reason to price hardware around the models and context sizes they intend to run.

Third-order effects

  • If local agents continue to push model size and context requirements upward, high-capacity unified memory becomes a primary differentiator in desktop AI hardware rather than a workstation add-on.
  • The market is moving toward heterogeneous local inference, where integrated high-memory systems and discrete-GPU PCs compete on usable model capacity, prompt throughput and total system cost rather than a single benchmark.

The trend: AI demand is turning desktop memory capacity and local-inference throughput into core purchasing criteria for edge compute devices.

Discussion

  • @ivanfioravanti Ivan Fioravanti on x
    @MiaAI_lab @plotarmordev Same thinking! Wondering why all testers went for Qwen 3.8 Flash Next, but tomorrow we'll see other models in action!
  • @thekitze @thekitze on x
    @viticci This convinced me to get a 5090 instead of m5 studio lol
  • @viticci Federico Viticci on x
    @MKBHD this is 100% spot-on. And if there's an area where 256 GB RAM (or even 512 GB) still ISN'T enough, that's local AI. give us more 😜
  • @ekurutepe Engin Kurutepe on x
    @viticci This is where the puck is headed. In a few years I think we'll all have an ultra-like machine powering our local models which we access from our mobile devices. I wonder if we'll really use MacBooks anymore.
  • @viticci Federico Viticci on x
    finally, shout out to GPT-6 Astra, which built a testing harness for local AI, then coordinated 2 Macs + 1 PC with remote connections and Computer Use over 3 DAYS Astra orchestrated ~140 local model runs across 3 machines. Wild thing to watch. https://www.macstories.net/...
  • @viticci Federico Viticci on x
    oh, and Apple also sent me an M5 Pro Mac mini, 64GB of RAM I spent more time with the M5 Ultra, but for the price and size, this thing also rips! will do more testing on this soon
  • @viticci Federico Viticci on x
    @TheAhmadOsman 100% true and crazy that I can run Flash-Next on my desk with this performance right now. That model is so good. https://www.macstories.net/...
  • @viticci.macstories.net Federico Viticci on bluesky
    I ran all the M5 Ultra tests I could think of:  — Two models (Flash-Next and GLM-5.3-Flash)  — 4k, 8k, 16k, 64k, 128k, and 256k prompts  — M5 Ultra vs M3 Ultra vs 5090  — 4/5/6/8-bit quants  —  I created all the possible charts and visualizations I could think of. …
  • @miaai_lab Mia on x
    M5 Ultra's prefill speeds are up around 150% vs M3 Ultra when running Qwen3.8 Next Flash 🤯 Strong numbers! I wonder what the prefill speeds would look like with a model that has more active parameters. Qwen3.8 Flash Next has only 6B active.
  • @mkbhd Marques Brownlee on x
    Been using the M5 Ultra Mac Studio for the past week and predictably it is the most powerful computer I've ever used. At this point, the improvements are not even targeted at me (a video creator) anymore. They're targeted at being the best machines for local AI work. There's noth…
  • @ivanfioravanti Ivan Fioravanti on x
    Great review of the M5 Ultra by Federico and let's keep in mind these numbers are just the starting point, everything will just get better in the upcoming weeks! Can't wait to try one and buy the 512GB version!
  • @mweinbach Max Weinbach on x
    I have been using the Mac Studio with M5 Ultra for the past week, and it's the most powerful computer I've used. Been a bit busy over the past week so my report will be slightly delayed, but just know... Qwen 3.8 Flash Next is the best model for the 256GB version, it rips
  • @viticci Federico Viticci on x
    I ran all the M5 Ultra tests I could think of: - Two models (Flash-Next and GLM-5.3-Flash) - 4k, 8k, 16k, 64k, 128k, and 256k prompts - M5 Ultra vs M3 Ultra vs RTX 5090 - 4/5/6/8-bit Flash-Next quants I created all the possible charts and visualizations I could think of. Hope the…
  • @zoneoftech Daniel on x
    The M5 Ultra Mac Studio is INSANE! Apple sent me the maxed out version, with a 36-core CPU, 80-core GPU and 256GB of RAM. I used it to run Deepseek v4 Flash (156GB model) fully on device to build a mini-game, which it did in just a few minutes! #MacStudio #M5Ultra @apple
  • @pcmag @pcmag on x
    We got an early unboxing of the 2026 Mac Studio, and Apple packed serious internal upgrades into this compact workstation. 🤖🚨 Powered by either the M5 Max or the new M5 Ultra chip with up to a 36-core CPU and 80-core GPU, it delivers up to 4.3x faster local AI performance and sig…
  • @jundotkim Jun Kim on x
    Great review from @viticci on the M5 Ultra Mac Studio. Glad oMLX could be part of it. Stay tuned for the next oMLX release!
  • @mweinbach Max Weinbach on x
    The other thing I noticed was this was the first time where I could run a ton of subagents, with cloud models, and each had access to enough CPU resources that it wasn't really limited. It also compiled rust like 2x faster than M3 Ultra
  • r/LocalLLaMA r on reddit
    M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents - MacStories
  • @stevemoser @stevemoser on x
    “One small wrinkle in the M6 variants the Mac mini uses: The basic 16GB Mac mini offers 153 GB/s of memory bandwidth, the same amount as the M5. The 24GB and 32GB versions offer 170 GB/s, an 11 percent increase that should marginally improve graphics performance and the speed of …
  • @chancehmiller Chance Miller on x
    My review of the M6 Mac mini is now live. It keeps the same adorable design we've come to love, and the M6 chip is impressive. A few things have changed since the last Mac mini update, though. My full thoughts: https://9to5mac.com/...
  • r/LocalLLaMA r on reddit
    M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents
  • r/apple r on reddit
    Mac Studio (M5 Max) 2026 review: Shifting into overdrive - MacWorld