214 comments

[ 0.21 ms ] story [ 44.1 ms ] thread
$12,299
5 years of (200/month) tokens at that price, meanwhile an rtx 5090 pc is about half that… hmm

but i wonder how much these token costs are sustainable or not, it may be in the long term cheaper to have your own hardware if token costs go up (and hopefully hardware gets cheaper again)

The token price isn't the only reason to run a model locally though. You can do additional training to specialize or remove censorship that may be a no-no per TOS with cloud GPUs.
Cloud GPUs have ToS about purposes you’re allowed to crank numbers for?

I thought this only applies to LLM inference providers, but not raw GPU rentals.

some GPU-rental providers have restrictions that aren't really enforcable, like bans on crypto mining
Those tokens aren’t guaranteed (esp. with RE and security tasks - rooted my own TV last week, Claude crapped out on “cyber safety” grounds; but also no guarantees about the model served - providers can pull a switcheroo on weights or quantization at any moment, and new options may not work for you), and you’re throwing money at entities that aren’t aligned with your interests instead of entities who are interested in actually empowering you.
> Claude crapped out on “cyber safety” grounds.

A $10/mo subscription to OpenCode Go would have done the job for you.

They have models like Kimi K3, Grok 4.6 , GLM-5.3, Mimo 2.6 Pro (launched today, already available) which are happy to follow your orders without accusing you of being a terrorist.

https://models.dev/providers/opencode/

The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

  Qwen3.8 27B tokens/sec generation speed

  Prompt size    8K    64K   128K   256K
  RTX 5090 PC    59    51    44     n/a
  M5 Ultra       48    39    32     24
  M3 Ultra       31    23.5  20     15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
Those are some incredible graphs, that improvement in prompt processing going from M3 to M5.

Also: ~30 token/s on GLM 5.3-flash, locally.

/meta Here's a CSS filter that stops those nuisance chart animations,

    macstories.net##*:style(animation: none !important; transition: none !important)
A dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures. But when you hit the limit of what you can hold in memory, you reach the limitation of the platform.

Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.

> A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures

Inference time is going to be dominated by the low memory bandwidth on these Macs, so a dense model will suffer most. It’s more of an opportunity for large MoE models with a low number of active experts since you can keep all experts in VRAM but not pay the bandwidth cost until they are used.

> you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers

This is an interesting direction that I expect to see more of. But for most models currently you need basically all experts loaded since they are chosen per token.

Apple seems to be researching longer horizon expert caching, where they keep experts swapped in for longer runs of tokens [1]. Other labs are offloading ngram caches but not sure if they’re pursuing anything like this?

1. https://machinelearning.apple.com/research/introducing-third...

1.2TB/s is already considered slow? Things are moving quickly!
1.2 T/s is not that modest is it? That's very close to an RTX pro 5000
Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
can confirm.

I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.

A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.
> for me at least a GPU is completely useless for anything but being a token generator.

No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.

> No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.

Crossover works on macos, too. You can run most games without a hitch these days (allegedly, according to /r/macgaming). But I don't play video games, so a GPU would probably be better off in some kid's computer.

A GPU would be better-off attached to your Mac in an eGPU enclosure. There is not a single Apple Silicon GPU on the market that leads the industry in prefill, decode or power efficiency.

But of course, Apple doesn't allow that as part of their ecosystem. It's really a privilege to have MoltenVK perform worse than the fanmade HoneyKrisp driver. It's valuable when Apple refuses to sign AArch64 CUDA drivers for macOS. It's exciting to pay Crossover to support half of the library Proton offers for free.

Clearly, I'm some sort of ingrate that selfishly demands the best things, without considering how to accommodate the poor trillion-dollar megacorporation.

A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.
Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
I think you’re missing that MTP can predict more than 1 token in advance.
In fact, isn’t that the “M” in “MTP”?
It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch
You are correct that I only specified for a single predicted token. Without knowing the workload though, claiming a consistent 2.5+ token prediction is not useful because it can't be reproduced. It is a software configuration and a specific workload, not a benchmark.
Have a 5090, and yes it's very fast. But it's like the worst ADHD team member and requires constant supervision and review from larger models. It's context size on-card is good for super, suuuuuper shallow precision work. The gb10/spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.
> it's like the worst ADHD team member and requires constant supervision

Perhaps consider some non-offensive language for your comparison?

could you not be like that? provide alternative language or go away. If youre offended, say so and be real. Noncommittal posits of personal preference are linguistic mosquitos of communication. on the flip side, how dare you disenfranchise a legitimate adhd perspective. one that i would say is entirely valid as someone functionally crippled by such plight. If you truly are offended, perhaps there is some truth you are reacting to preventing you from truly responding in good faith. words are lame like that ya? mine are as nauseating as your flyby ego droppings.
Ok, as somebody with ADHD I find it offensive because I don't need constant supervision, implying people with ADHD need constant supervision is belittling and just plain wrong. So, I will do that, thank you.
Is a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?
Personally I think the price is way too high right now. It’s a power hungry gaming GPU. The efficient single card equivalent would be a 4500 Blackwell which launched at about $3500. Or you could get a 9700 32GB or an Arc B70 for well under $2k, today. You only buy a 5090 if you want absolute speed.

32GB is still not that much. I would rather get a Spark and have the RAM to experiment with larger LLMs, even if it was slow.

You can lower the wattage and it doesn't lose much perforamnce.
Being fast and having a power target doesn’t mean it’s cost efficient though. I would pay the launch cost for one, but not 3-4x inflated.
How does this relate to 4500 vs 5090? I'm just pointing out that 5090 likely has twice the performance of the 4500 and likely maintains that at 2x watts if you want.
You didn't specify in your earlier post, so I wasn't sure exactly which comparison you were making. But yeah, the perf/watt actually looks the same for those, so the cost per token evens out. It is nice not having to manage 400+W though. I like the 4000 for that reason, it's effectively a 3090 that runs at half the TDP.
Qwen3.8-Flash-Next is pretty damn worth the extra ram you need.
How are you deciding which work to send to the 5090 vs a frontier model, or making the two work together nicely?

Correct is much more important than fast for me, but if I could get correct and fast, that would obviously be amazing.

Because you can run Qwen 3.8 Flash Next, Laguna S 2.1 and other medium-sized models that simply don't fit on a 5090?
Both are probably single-token decode performance, which is reasonable to show. Otherwise agree RTX 5090 should shinebetter with NVFP4.
When I say that Apple is astroturfing HN, this is what I mean. Anyone can easily verify that even a 3090 will blow any Mac out of the water in terms of Tok/sec. Somehow its ok to just post outright lies when it comes to Apple product comparison.
I guess it is possible, but Apple has had very vocal fans for decades. I suspect, rather than astroturfing, it is just people who are in their ecosystem.
Tok/sec is 0 on a 3090 for most of the models that the mac can run
Running very large models on Mac is unusable at 10 tok/sec. You get more average inference over the day using free Google Gemini.

And for the price of a Mac that can run a large model, you can get 2 3090s humming along running a small model so fast that it can simulate a lot of the behavior in large models just through sheer number of context it generates. For example, editing code means that by the time your large model on your Mac is finished writing a file, the smaller models have generated the code, written the code to file, ran it, and debugged any issues.

So given that, which one of these is true about you?

1. You are paid by Apple to push marketing on HN

2. You are a hardcore Apple fanboy and just think that owning a Mac studio is a flex

> So given that, which one of these is true about you?

Well, if those are the only two options you can come up with it's pretty clear that this isn't about me or what I am, you have a false model of reality.

> Running very large models on Mac is unusable at 10 tok/sec.

There are plenty of examples of models running at well over 10 tok/sec that aren't viable on the 3090. In fact such examples are found in the review in the OP. Did you not read the article?

I think you're projecting pretty hard with the two options you've listed. Go touch some grass, you seem overly frustrated that reality doesn't meet your expectations.

Since you clearly don't use local llms, allow me to educate you - anything under 100 tok/sec is USELESS. When you are coding, the idea is that you want to have a system that can generate files fast, hopefully correct on the first try. Cloud models do this. Local models, by nature of having less parameters and more quantization, often require more guidance and repeated inference to get it right. The antigenic harnesses that people set up around local llms leverage this.

Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.

Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok/sec prefill, not to mention insanely quicker during inference. So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper

The whole thing of "well I can run bigger models that don't fit on a GPU" is either paid Apple advertising, or you are just an igorant fanboy.

So I ask you again, which one are you?

> allow me to educate you

No thanks, you're not in a position to do that clearly.

> Since you clearly don't use local llms

I do, probably a lot longer than you have actually.

> anything under 100 tok/sec is USELESS

Objectively wrong. You sound like you're really behind and you're so myopic that you think coding is the only use case for local LLMs. I'm a professional software dev and that's the least interesting use case of local LLMs.

> Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec.

You clearly didn't read the article or have reading comprehension issues. The model is Qwen3.8-Flash-Next 4 and 5-bit quant, neither of which "conveniently fits on one GPU". Sorry that your hardware doesn't live up to your own delusions and can't even run Qwen3.8-Flash-Next at 4/5 bit quant. You are taking the Quen3.8-27B numbers, something that the article isn't really that concerned with, and trying to make it fit into your narrative.

> So I ask you again, which one are you?

Well I'm someone that suggests that you should touch some grass and reevaluate your personal issues. You seem angry. Perhaps it's best to figure your own issues before trying to figure out why people are excited about Apple hardware for local llms. I am sure the people that need to interact with you in society would be very grateful if you took the time to do this.

Nice try.

A) He literally says "I tested a different Qwen model for the comparisons between Mac and PC." The model he tested has to fit on one GPU, otherwise the inference is dogshit slow as you are offloading results to ram. If you ran any amount of local inference, you would know this. Considering that Qwen3.8-Flash-Next Q4 is still 100gb, there is no realistic way to run this with a 5090. The model that was run was this https://ollama.com/library/qwen3.8:27b. And the speed of that model on a 5090 in terms of tok/sec is not 60 lol.

B) If M5 ultra runs 40 tok/sec on qwen3.8:27b (and lets assume its the mlx version to gain a performance boost: https://ollama.com/library/qwen3.8:27b-mlx), you have to be delusional to believe it can run 100gb models at 100 tok/sec lol.

As a bonus, in terms of use, its pretty well known that Qwen models are RLed to chase benchmarks. Check out https://huggingface.co/Qwen/Qwen3.8-27B versus https://qwen.ai/blog?id=qwen3.8-flash-next, using different benchmarks the 27b outperforms the flash next on agentic coding. But it matches it in other areas pretty well. So tell me again why you need 100gb models?

It is so incredibly sad how hard you try to sound intelligent. But thats on par for the course of any person hyping up apple products, throughout apples history.

Considering that Apple probably doesn't want you to engage in this level of pettiness for their advertising posts, you have outed yourself to be #2. And Im not angry at all lol, you keep doing what you do, people like you in the industry are the reason I can work 8 hours a week and still get get paid a lot while being reviewed highly.

Honest question (as a person outside of AI industry): I understand the title "M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents"... but if those LLM benchmarks are so settings-sensitive, why not to dedicate at least half of the article to some general metrics? Maybe matrix/tensor/N-body calculation performance and maybe some memory bandwidth measurements, which all contribute to LLM performance, but allow to compare more-or-less apples to apples for those different devices (M3 Ultra, M5 Ultra, RTX 5090, et al.)
The issue is that the moment you want to run the more capable models that will no longer fit in a single 5090's memory, performance falls off a cliff.
Thank you for this. I wish Apple focused their silicon design on improving the TTFT metrics but coming from an M3 Pro, it still looks laggard compared to Nvidia's TensorCores in the 5090.

Maybe Apple is an acquisition away from changing that balance.

The rumor on Apple's processor roadmap is that they're skipping other M6 variations (all previous generations had Pro and Max, a few had Ultra) in order to focus on the M7 generation for AI reasons. What exactly the M7 improvements are who knows.

https://www.macrumors.com/2026/06/25/2027-macs-m7-chips/

I think that comes down to TSMC. Nvidia apparently booked out the whole A18 or 16 node. Apple is on 2nm right now and M7 will jump right to A14. According to my quick AI research anyway.
That sounds a lot like AI fantasy slop.

Apple just shifted to N2. They’re not going to be doing another major shift right away.

And TSMCs own roadmap would put your hallucination years away at best for a a product that follows a roughly annual cadence https://www.tomshardware.com/tech-industry/semiconductors/ts...

M6 got another prompt processing boost. Likely no M6 Ultra though because Apple is reportedly going all in on AI performance in M7 generation.
The selling point of the M5 Ultra Mac Studio is that you can run much larger models that the 5090 can't without swapping. NVidia aggressively segments the market on VRAM for this reason. That's why a 5090 has an MSRP of ~$2k (but good luck getting one for less than $4k) while a 6000 Pro, which is basically a 5090 with 96GB of RAM has now soared beyond $15k where 3-6 months ago it was more like $10-11k. A 6000 Pro has the same memory bandwidth but slightly more CUDA units (IIRC ~24k vs ~21k).

This advantage won't be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can't do on a 5090.

I don't think we'll get a successor to the 5090 until late 2028, maybe even 2029. I'm basing this on the launch date of the 5000 series and that we haven't got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards.

Apple should see a Mac Studio major update in 2028. That might even force NVidia's hand. But it's really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however.

The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home/enthusiast solutions.

But what about builds that combine 8 of the 5090 with infiniband between boxes? Wouldn't that be comparable to the mac in terms of price and potentially beat it by a lot in terms of performance for the large MoE? I understand the space/heat/noise considerations, but price wise it may still not make as much sense as people think. (Agreed that it is hard to get the NVIDIA hardware and the 6000 pro are priced less competitively).
While that sounds super awesome, How many people are actually going to build and maintain that vs a box you can grab at the mall that fits in a lunchbox?
No, $40K is not comparable to $10K.
When I compare those two numbers, it seems there's $30k of difference
Sounds like nice utility bill in the making.
I can't speak to Infiniband pricing for something like that. It seems like the cheap option is 56/100Gbps with used Enterprise equipment. You'd need 8 HCAs, DAC cabling and a switch but even then you're into thousands of dollars. If you want 200Gbps+ it gets into the tens of thousands (AFAICT).

Each PC is probably going to cost ~$6k and you're talking about 8000W of electricity draw. That's going to consume multiple 20A circuits even at 240V. And the electricity ain't free either. A Mac Studio seems to draw ~500W max.

Oh and the Mac Studio has an upgrade route to run 1T+ models too by chaining them together with TB5 chaining. OSX supports RDMA this way. That's comparable bandwidth to the 100Gbps Infiniband option.

So you're talking about $50-60k of hardware and more power draw and more heat for something that will I'm sure beat the MS M5U option but at huge cost. Also, at that kind of price point, I'm likely to get a workstation PC and put 2 (or possibly 3) 6000 Pros in it.

Others have chimed in on cost and size, I'll chime in on power. The 8x 5090s will require a dedicated datacentre grade power source. The Mac Studio runs on a plain jane wall socket.
>8 of the 5090

Where are you buying 8 5090s for under $10k? With CPU, RAM, and (checks comment) infiniband hardware???

You're probably looking at a lot closer to $60k when all is said and done, and that's before you hire an electrician to run a sub panel for your homelab...

openai make a npu google make npu (tpu no mater) amd buy tellas

every company make his own npu (without xai)

probaby in 2028 we will have more concurent firm on market place

Not forgetting of course that an RTX5090 is what 600W+ ? And the Mac is probably half that at most ?
Certainly not forgetting wattage. A 5090 is 575W. The M5 Ultra Studio is 480W.

nvidia-smi -pl 450 for like a 4% reduction in throughput. I tend to set it around 350W because it's a comfortable temperature blowing on my legs under the desk without warming my office in the summer.

I put together this system two years ago, so it's a little out of date, but it only cost $3000 for the same performance and capability as an Ultra. I don't think I would spend $7000 to save 100W, though.

> nvidia-smi -pl 450 for like a 4% reduction in throughput.

Yeah people don't pay enough attention to those settings IMO. The first thing I do when I set up a new machine (or upgrade my OS) is to restore all my powersaving configs.

For example I've got all but one of my virtual desktops that put the CPU in powersave mode: I don't need max Ghz when browsing the Web, not even on demand. But when I switch to the virtual desktop where my development environment is, then I want power on demand.

Now I don't do it to save the planet: I do it because I love a quieter computing experience (coupled with Be Quiet! PSU and Noctua fans, this makes for a very quiet computer). That it consumes less electricity is a nice side-benefit.

When you are doing matrix math, compute is compute. Apple cant be more efficient due to physics. The only reason Macs are more efficient in general is that they have tightly bundled hw and sw for specific tasks.
That's a dense model. Of course it will do worse.

Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).

I assume those are non-batched. I think the M series GPU can do 4X to 8X depending on model quant, which means if you can batch queries you'll get almost 4X to 8X performance.
On my m5 max 27b model does 75tps on 256k ctx and starts at 80 on the 8k ctx when you add https://huggingface.co/collections/z-lab/dflash-2 to it. So yeah base might be 30tps (I used iq4) but mtp or dflash help a lot and should be used when checking what is useful and what is not for running models as it is not fare to judge without them.
The next Ultra, supposedly on deck in 2028:

> Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators, according to a new Bloomberg report published by Mark Gurman...

We've already heard that Apple plans to release a base M6 chip this fall for entry-level Macs, then skip the Pro, Max, and Ultra versions of that generation and move straight to the M7 line. However, Gurman now reckons that we'll see a base M7 in the first half of 2027, M7 Pro and M7 Max at the end of 2027, and the M7 Ultra in 2028. Apple reportedly began taping out the M7 about six months after it started the same process for the M6, which is what has enabled the company to pull the schedule forward.

https://www.tomshardware.com/tech-industry/semiconductors/ap...

I really appreciate seeing these dense model numbers. For a large unified memory system though I expect that MoE numbers are what people are more interested in.

These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens/sec to start and slows down to ~110 tokens/sec over 128k context and can do the max 256k. These are for batch size 1 and throughput goes higher with batching. This is due to speculative decoding, efficient all-reduce inter-gpu compression, and custom GEMM kernels for the specific hardware. Note each R9700 only has 644 GB/s memory bandwidth.

A basic llama-bench on Qwen 3.8 27B UD-Q4_K_M gives pp512 3920 tok/s / tg128 81 tok/s on a 500W RTX PRO 6000 (should be similar speeds to a 5090, chip is basically the same, just less VRAM). With MTP3 this is 140 tok/s on mtp-bench.

This is with llama.cpp. You can of course use vLLM/SGLang well on these cards and they're even faster. On vLLM w/ NVIDIA/Qwen3.8-27B-NVFP4 baseline has a prefill of about 13,000 tok/s. The baseline tok/s is 72 tok/s, but at mtp7, it's 157 tok/s, and w/ dflash7 that goes up to 215 tok/s. On mtp-bench, DFlash2 gets a hair under 300 tok/s w/ the code_python prompt.

Yeah that can't be right, my M5 Max gets almost those speeds and certainly lot faster than what they're claiming the M3 ultra gets. Maybe they didn't have the model setup right or were running it with some unnecessarily high quant (>=8bit).
On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.

That’s like 12 years worth if OpenAI Pro subscriptions

Yeah, anyone who thinks local AI is going to save them money is likely to be disappointed, at least if they want to run models that are even remotely capable.

Plenty of other reasons to get excited about it local AI, but I don't think cost is one of them.

[delayed]
Agreed, plenty of other reasons to get excited about local AI.
I'll try that argument with my wife next time I want to buy a $15k mac.

Its a bold strategy cotton, lets see if it pays off for em.

Despite being on a site called Hacker News, we seem to often overlook the simple aspect of wanting local AI hardware to hack (not necessarily in the cybersecurity sense) with. I got my local AI hardware because it's an enjoyable hobby for me with bleed over into professional pursuits (but work provides plenty of compute access so this spillover is a very minor factor).
apparently if you ever point out HN starts for hackernews and thus expect related attitudes you get downvoted by shocked ( what I guess are zoomers and not bots ) that desperately opine the name is a random abberation doesn't mean anything and one should not deviate from our corporate overlods in any manner.
A lot of people commenting on how bad idea buying local hardware for inference is also miss the fact that even in 3 years that hardware gonna cost something.

Might be if RAM prices get much more reasonable its gonna be 1/3 of the price, but it's very much possible its gonna be half or more.

And if you're buing Mac Studio and not some AI-only locked down board it's possible to reuse it for other purposes.

Yes it surprises me too…
On a personal level maybe not yet, but for a medium business upwards it may make sense.
Having no debt and owning in the long run always works out better than a lifetime of renting, if you don’t have to, the massive rent letting these days, is very frustrating at some point don’t you have to draw the line?
Hard to guess, it can go either way. If you will need to be in a syndicate to use non-sterilized models, that mac makes sense. But if there is mandatory registration of personal cyberarms, you risk going to mines once they check you purchases. You could try to play normie and pretend you simply wanted to show off, by keeping your actual work on external disk, but that leaves traces on system. Counting on someone in the Gap renting you gray iron works as long as you can swap credits. Still, this gear is tiny. Put it in your e-car, with uplink, and leave it at uncle's farm. Discreet.
It was a dark rainy night in Neo-Tokyo as Blake puffed on his vapor cartridge and watched the Mac dealers prowl below. Almost 15k Union Credits to get one of them to meet you in an e-cafe with a fully loaded M5, but man, the inference rush from one of those things was something else.
superior zero latency local skooma
I'm sold on "personal cyberarms" as a concept

Do they include footguns from pointer bugs?

512 option isnt worth it imo, you get severe slowdowns when weights are that large. 256 is the sweet spot, you can run large open weight models at decent speeds for full private inference.
> 512 option isnt worth it imo, you get severe slowdowns when weights are that large.

I think most people are getting 512 for running Chrome with a bunch of tabs open. /s

a) we don't actually know what the prices will look like yet, b) what about same weights + huge context? or, same weights that you'd run on 128gb/256gb, but multiple models running for different tasks?
I think it's safe to start the conversation as about bad as the jump from 256 to 512 on the M3, which was a little more than double base to 256. If it's surprisingly different at launch then it can be a party, but there is no sense getting your hopes up for that at the moment.

Longer context also slows token prediction proportional to the context size. If it wasn't regularly referenced then there would be no need to keep it in RAM.

Usually the pitch for more memory is "I can run a massive model/context and get my answer in a while instead of next weekend from disk".

> you get severe slowdowns when weights are that large.

Not necessarily for MoE

The model being tested is 18k as configured.

I didn't expect this to make the 5090 to look like a good deal.

5090 has 32GB VRAM.

It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.

Is it that silly? You could run multiple 27B models in parallel.
You actually don't need more RAM to batch multiple inference tasks of the same model.

(Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).

You definitely need more RAM if you are not satisfied with small context windows, especially if the weights take a large % of the total memory to boot.
While it is silly to get the huge VRAM model and then not use the VRAM...

> You use models like GLM Flash and Qwen Next which won't fit on a single 5090.

Do those run with usable levels of performance, though? There's some novelty to it running at all, but can you actually do anything of value as a result?

can run multiple subagents of Qwen 27B though, right? Unless I am fundamentally misunderstanding how VRAM constraints work
You might be. Running another agent doesn't load a set of new weights. It creates a new KV cache for the agent and adds the prompts to the queue. Its just another inference turn.
(comment deleted)
thanks, I naively assumed when, for example, Claude Code starts subagents it loads a new instance with empty context
> "It also happens to be a Mac, with an operating system that looks nice and doesn’t suck"

Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.

What exactly is missing from macOS that makes you feel the need for Linux?

I get it on Windows systems, at least when someone wants to use Linux-type tooling. But macOS already supports pretty much all of that natively?

> What exactly is missing from macOS that makes you feel the need for Linux?

For starters, the source code.

And why would you need that to run LLMs?

Apart from that, for the UNIX part, the source is available for quite a few components:

https://github.com/apple-oss-distributions

notably also the kernel

https://github.com/apple-oss-distributions/xnu

> And why would you need that to run LLMs?

kokonokko1337 already said it was good enough to run LLMs, presumably RunSet isn't saying the source code is needed to run an inference server.

They say, not doing anything with Linux’s source code other than mentioning its existence.
Comparing to Linux is too general. You need to say what distro and what install level. I really only use headless Linux, but I've never logged into a fresh install and not had to do some sort of 'apt install devel-packages' equivalent for which ever distro being used. That's the same thing with macOS after choosing which package manager to use. I don't see how Linux vs macOS is very different
The Asahi Linux group goal should have been organizing themselves by taking out a license on Arm processor (probably too late with the present owners) and building a new Linux Distro to go with it instead of being a parasite on someone else’s existing hardware, if the three ex-engineers from Apple could do something like that why couldn’t they aim higher? Why waste time spending years reverse engineering. Wouldn’t it have been easier to raise money for such an endeavor? I hope someone in kindergarten, junior high or high school who doesn’t know any better will try something like this.
This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.

It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.

The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.

An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.

> This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. I think even the SOTA models barely code well, with Opus 4.5 being the first, good coding model.

That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).

With the latest codex (weekly quota burn) fiasco I tried open weight alternatives for the first time. And tyeah... open weight models cant compete with likes of astra yet. But, my hope is that by the time I get my Mac studio at end of november an open weight models would have closed the gap (which i think is realistic at the speed of progress). Now its true a better gpt version will also be available then but it also seems the gap is shrinking with time so theres that.
> And tyeah... open weight models cant compete with likes of astra yet

I think this is true, but also misses that a lot of us are just doing basic flask apps with a react front end. We don't need astra; Something sonnet 4.6 level locally is perfectly sufficient 95% of the time, and maybe 99% of the time.

This. People have convinced themselves that the absolute frontier is what is needed, anything below it is an unacceptable compromise, and we seem to be speaking different languages when it comes to discussing model capability.

It's like watching a discussion about cars available to take on a 100km road trip. A new car gets released that is on par with a Toyota Corolla but it is dismissed as completely useless for a 100km trip because it doesn't have the seat massagers and air ride suspension that the new Escalades have.

The reality is that something like Sonnet 4.6 is still amazingly capable for so many programming tasks, especially if you already have some reasonable level of experience to steer it in the right direction.

And if you think Sonnet 4.6 is still worthwhile, then it seems undeniable that something like Qwen 3.8-27B is also worthwhile.

The problem is that even if you're doing CRUD apps, Sonnet level will be good enough... 95% of the time. But the 5% will kill you.
Local models can be widely used as productive assets. Yes the infrastructure of SOTA API models is engineered specifically for you to be that utility, but the blanket statement that local isn't up to par is intensely short sighted. Billions of tokens per month on local pays for the hardware when compared to sota costs per month.
I believe they can currently be used productively for non-coding tasks (classification, light summary)... but they definitely are not even close to SOTA when it comes to software development.
Defining productivity is a use-case scenario, and a wildly generalized assumption for most people in this argument. Local infrastructure doesn't need to be sota for absolutely every single need for a dev lab, but it absolutely can be delivered with non-api frontier class models.
Just to be clear, I'm specifically talking about coding. I think local models can help with productivity today, just not coding.

I'm also a huge fan of local models and think it's absolutely imperative that they continue to advance so we can move off of the Anthropic/OpenAI hosted models. It's important to accurately asses where we are in that journey though.

Like the other commenter, I'm confused about the 'just not coding' conclusion. I'm using Qwen 27B on a 5090 at > 100tk/s with 150k context (which isn't enough admittedly), and DeepSeek v4 Flash with 1million context on a gb10/spark. Both of which are performing surface level, and deep needle precision infrastructure architecture. They code 24-7, stupendously.
It would be interesting to hear more about how you’re actually using them. Do you have sophisticated feedback loops around the models so they can verify their work and converge on good solutions? And how do you decide what to give the 5090 vs the Spark vs a frontier model?

Correctness matters much more than speed to me, but if I can get both, that’s obviously very interesting.

Local models are undeniably capable of "helping with coding" today.
I so want this to be true, but for the kind of coding I do (not Flask apps), it's definitely not the case. Like I said, SOTA models just barely, barely work for me. My projects are usually 100k-1M lines of Rust or Go.
Out of curiosity, what do you find the SOTA models are simply incapable of when it comes to your Rust and Go projects?
The SOTA models now work really well in my codebases, but that's only been since Opus 4.5/4.6-ish. Prior to that, and with current local models, they simply couldn't work holistically and would just thrash around. Now I feel as if SOTA are approaching my coding levels if not surpassing it. I still need to guide on architecture, but I can see that going away within the next year or so as well.
Thanks, that makes sense. When you said they “barely, barely worked” for you I assumed that meant something different.
Astra-Ultra? Even the largest open model to date (Kimi K3) is nowhere close to Astra level, and it will be quite slow even on the highest-spec M5 Ultra, with achievable speeds of about 0.5 tok/s at most due to having to stream weights from SSD (~13 GB/s on the highest storage capacity M5 Max machines so far). This is OK for doing simple Q&A in the background but it's far from a genuine coding experience. You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights (and this is where the "Ultra" part sort of becomes relevant; Kimi series models have good support for agent swarms) but this would decrease single-session performance even further. It would only be usable for background jobs, though the hardware would then have a chance of paying for itself if it was fully used on a 24/7 basis.
> You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights

isn’t this very straightforward to do..? I thought batching for Qwen models is already proven out.

> but this would decrease single-session performance even further

Well let’s take Qwen 3.8 27B. Throughput for M3 at 8 agents is 4x compared to single agent. [1]

It’s really not clear to me that 8 concurrent agents at half speed will be worse task completion latency than 1 agent.

And that’s M3 studio benchmarks, not even M5 ultra, and without the many software improvements we will see

If you haven’t tried Qwen 3.8 27B xhigh on a task you might not get the hype. Idk.

If you’ve tried doing this and don’t like it sure, and be specific about what isn’t effective, but let’s not speculate.

[1]: https://omlx.ai/benchmarks/performance/69kzkrv8?utm_source=c...

That's all well and good but Qwen 27B is a small, dense model; that's favorable to both batching and MTP. Batching of large, sparse/MoE models like Kimi K3 (requiring slow SSD streaming even on a single maxed out Mac Studio) on local hardware is an entirely different game that's mostly theoretical so far: many people would even call it outright pointless. (MTP clearly fares even worse, though - unlike batching, it ends up wasting scarce weights-fetching throughput on wrongly predicted tokens.)
>Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?

Ehh, the actual elephant in the room is:

"why bother with local AI at all when you can lease a GPU for $5/hr?"

To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.

> unless you have a bunch of money to throw at hobby projects.

there are lots of people with very expensive hobbies, see sailboat racing for example.

"Lots" of people are billionaires too, just like "lots" of people like to watch TV in their free time.
$5/hr = $3600/mo.

Unless you only need the AI available some of the time, $5/hr is pretty expensive. That's an RTX Pro twice a year.

If you're using it for discrete sessions of coding or something, that might make sense for you but if you're using it for an always-on assistant, that pricing kinda sucks.

I would imagine extremely few people are utilizing an H200 for every hour of a month. Especially for something like an assistant

5090's are like $0.20/hr

For a lot of people involved in local LLM, ownership and control of their data matters to them a lot. It's what pulls a lot of people toward local models. Renting a 5090 off of some random infrastructure provider that is reselling some random dude's 5090 does not fit the bill.
While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.

This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.

Speed vs task-completion is a new conversation and a great point. Whereas the cost to compute doesn't exist in a vacuum, making mistakes costs less, is easier to maintain with granularity and a whole host of other factors when you own the lab.
Imagine spending a trillion dollars on data centers and then reading this article. Nightmare fuel for OpenAI
And nightmare fuel is just what they'll be selling at the UN this week, for this very reason.

Sam's address will probably be more riveting, imaginative, and terrifying than the last couple of Terminator screenplays. Legislators will lobby him to write the laws for them, and the ghost of Harlan Ellison will threaten to sue him.

I don't see how that math works? This is a $15k rig under benchmark and per the results it competes very acceptably against... one consumer GPU.

I really don't see who buys this, except people who want the Studio for some other reason. But nothing in the story says you want to fill racks with these instead of Blackwell or TPU parts; it's not even close.

Your math is correct, but it’s math based on today’s economics.

Think of a company like Apple moving onto your turf. They’re not going to cede AI to the cloud. They want their part of the pie.

So in 7 years, how much AI will be handled locally on your iPhone. And will you have repaid all the debt on your balance sheet before Apple eats your lunch

I think most important thing is that Nvidia doesnt want to give 100% of market to frontier AI labs either.

It's way too easy for 1T+ frontier labs to ditch Nvidia. So Nvidia will also put effort to make sure there are open weights models and local hardware available.

And Apple will benefit from this too.

There is zero chance that an LLM approximating a modern frontier model is going to be running on a phone in the next decade. Even if you grant that you could stack enough DRAM dies on top of each other in the package, that would be a three order of magnitude improvement in power efficiency just for the compute.
This dream machine costs over $15k (not including the Apple Studio Display)? Nah, that dream is SO OUT OF TOUCH!
The irony here is, thats hobby hardware. You spend 20k and can run slow hobby models that are barely capable of anything unsupervised.

Wntrg level serious hardware starts at 100k, and a bit better but still almost-useful grade is 200k (8x rtx pro, plus a nice epyc pairing). Thats the sort of thing a salaried expert lets their employer buy them for sort of serious work.

Anything really serious is well north of 1M - not including the housing and commercial grade mains connection. And at best that buys fast Kimi K3 or GLM.

You can run Deepseek 4.1 flash on a 128gb Mac, comfortably with lots of cache and context on a 256gb one.

These models are not 'barely capable', they're contemporary near frontier.

I do have a 128gb mac, and the experience is, for my use cases, hobby grade. I use models on a daily basis for design work (chiefly), and sometimes implementation work. I can't justify a drop in output quality and significantly increased waiting time. Sure, if frontier did not exist, I'd use them, but way differently than fable or astra.
If I had 0 privacy concerns I could just use deepseek directly in the cloud but the whole point is to not put your data on someone else's computer.
If you are buying expensive hardware to run LLMs "on your own machine" you will soon find your ladder is on the wrong wall.
Ordered one for OpenFOAM. Excited for it. Will be live to not have my laptop running CFD 24 hours a day, but my M1 Max is currently my fastest machine… I’m expecting about 3.5x from the M5 Ultra.
I think Apple is really sleeping on making this run a Linux server. These things are very capable and draw very little wattage when idle. It would make an excellent homelab device, but MacOS currently holds it back in this regard.
i use linux a ton too, but still wonder what you want in a Linux server that macos as a BSD server does not have.
I'll probably manage with Mac OS well enough, but my linux distro comes out of the box with all the latest OSS tooling I'm familiar with, plus a package management system, and it has linux cgroups and namespaces that power the container technologies we all know and love.

If I switch to Mac OS, I have to sort out a package manager and install all the stuff that's missing, and when it comes to containers... they're just linux VMs. I'd happily cut out the weird proprietary middleman if I could.

Linux for argument’s sake, may have a few things that are better than Mac OS but Apple being the last vertical computer company from the 1980s, I don’t think they have any interest in using Linux, not after Next, Motorola, IBM, Intel and Nvidia in the past. They don’t need to they appear to navigate thru tech very well in comparison to Microsoft or Intel, for example.
My mac is 5 years old. I don't think I can comfortably buy a new one right now. It has a 16GB unified RAM. Honestly that would be enough for so many local models that I want to use but can't use. Because RAM usage (even with literally every single user installed app quit/stopped) the RAM usage is very high that I can barely safely get 6-7 GB (I am supposed to get ~10 GB, but it goes up and down real fast!). That's a shame. If only I could install an alternative OS that uses very little amount of RAM :-)
How are you measuring memory usage? top tells me that 45/48GB is "used", but Activity Monitor shows me that 24GB is cached files.

I'd be quite surprised if Mac OS alone needs more than 8GB, given that they sell the Neo with 8GB of RAM today.

When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.
At this point I think I will get the DGX gb300 workstation though I will wait a bit more for the cold season. It is double the price but at least is the real thing
The Year Of Local AI will be here no later than 2040, coinciding with the Year Of The Linux Desktop.
The year of linux on the Desktop was 26 years ago for me.
"a total cost of $0" Uhh ... how much is that hardware?
Another comment approximates at around USD$15k, so yeah, not zero.
For HN readers earning that sweet sweet VC money, that is zero!
I think the power itself will probably be more than you're paying in a subscription
Lol, try generation of images & videos on these, they ought to improve perf on Deep learning not just llms
What are good options to run local models nowadays? Something good for coding and personal assistant kind of things
I think we'll eventually get to the point where folks will have a local AI agent but I think people need to temper their expectations to a degree. You aren't going to have data center level tok/s from a box sitting under your desk and you don't need instantaneous responses for many workloads. Having a local agent that can execute tasks over a couple days with your supervision that might otherwise take you weeks is perfectly acceptable.

However I also think that Agentic AI is very much not an out-of-the-box solution, local or otherwise, and it takes a high level of technical knowledge to create an effective AI agent. And there's a problem now where most orchestration is fixed on what models are used for what tasks with no ability to weight constraints like cost, speed, and security.

The solutions will come in the past, big iron, thought they were immune and not too long afterwards, personal computers took over and big iron was history well, the same thing is going to happen all over again.

Those companies building those big data centers are going to find out that they overspent. Because we’re not going back to the mainframe era, no matter how much OpenAI, Anthropic, Microsoft, Meta or Google would like to.

Fantastic review, this is how it should be done with local AI.
i can buy a house with that amount of money. it used to be car money.
Once you hit the memory you need, generation speed is mainly set by bandwidth, and every Ultra from M1 thru M3 has ~800 GB/s. IMO best ROI for most people is 'cheapest used Ultra with enough RAM'

I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)

I think M1 through M3 were compute bottlenecked in prompt processing (hence the very large gap between M3 and M5, in this page's benchmarks, that's not explained by memory bandwidth alone).

For generation speed in isolation, yes.

The M5 generation added tensor instructions to the GPU cores.