69 comments

[ 1.9 ms ] story [ 21.6 ms ] thread
I am really surprised you’re seeing half the generation performance of 3.6 with 3.8 at the same parameter size and quantization (and same prefill performance to boot) - is there just an optimization in the stack somewhere for 3.6 that hasn’t landed yet for 3.8?

Curious if other folks are also surprised by this apparent discrepancy or have a ready explanation.

The local LLM scene needs a Draw Things equivalent for Mac. Too much fiddle for things that doesn't make sense (Qwen 3.8 27B should be exactly the same speed as Qwen 3.6 27B). It feels like that I am teasing (I am the author of Draw Things) something, because it is.
> Qwen3.8 27B (Q4_K_M, 17GB) generates at ~14 tokens/s on my Mac Studio M3 Ultra

>~14 tokens/s

For anyone reading that has never ran local llms, please understand that anything under 100 tok/sec is worthless. You are faster typing stuff into Gemini free version that you get with a google account and copy/pasting it in (and you can easily build browser automation with playwright or any other js runtime to have this available in a chat window)

What? Below 100tk/s is worthless?

I agree that 14t/s is pretty tedious for interactive use, yes. But 50-60tk/s is faster than I can read. 100tk/s is outright fast. Don't forget there is a limit entering content into meatspace.

Also, Gemini may be free but what if I don't want to give all my data to Google? This is precisely why I have a lot of stuff locally.

And will it remain free? How are they going to make back all those trillions of investment?

But yeah I would kinda balk at 14tk/s too that's why I use old datacenter/workstation-class GPUs.

I get 80 tok/s on the same model, and it's pretty usable. I'm not chatting with it; it's either given a bag of tokens to generate an answer or it's doing some agentic coding.

https://openrouter.ai/anthropic/claude-opus-5 is it worthless because its 65 tps?

re: gemini

https://openrouter.ai/google/gemini-3.7-flash worthless as well?

That being said, 14 tok/s is pretty slow.

The standard is not what you can do with it, the standard is what is the free alternative. Local LLMs need to be able to beat the rate limits on all the free models like Gemini to be useful. The only way to do that is to have high enough tok/sec, especially for agentic loops.
I mean, if you don't care about your inputs being trained on, you can just use one of the many free models on openrouter.
In agentic scenarios, an LLM has to read far more tokens than it outputs. I think focusing only on the decode speed is somewhat misleading. 14 tok/sec for decode is actually okayish. 93 tok/sec is what's abysmal, my RTX 5090 goes above 2000 tok/sec with 5 bit quants.
Quality of tokens matters almost as much.

30 is worth it on a meaningfully better model, though annoying. 50 is pretty much unnoticeable (good old 60 FPS). 80-100 is heaven (aka 120 FPS - once you get used to it, it does suck to go back).

I would still rather have 30 tokens per second of GLM 5.3 Flash than 100 tokens per second of Qwen 3.6 35B-A3B.

My lower bound is 20 t/s. 100 t/s would be great but even Opus doesn't sustain that.

I'm at ~7.2 t/s on a Shoehorn'd obliterated Qwen 3.8 on my M3 Pro, which is rough.

I'm encountering the same behavior. I've tried 4-8bit quants and get 14-17 tok/s with one run that achieved 19. I'm eagerly awaiting dflash2 support in Unsloth or LM Studio, as allegedly that should increase throughput to around 30tok/s, which is the baseline for what I consider at least somewhat interactive.

Jealous of the folks with 5090s running ninfer and getting >100tok/s. At those speeds it's a true frontier replacement IMO.

I run similar workflows as the author on my 5090. It's really good and reasonably fast at ~90 tok/sec on LM Studio. I haven't tried ninfer yet. The only problem is having to be mindful about the context size. I am jealous of the folks with RTX 6000.
Ninfer is the real deal. 160 tokens/sec. You can parallelize 4 chats at once and get ~400 tokens/sec I’ve heard. Although that just compounds the context size issue, still absolutely incredible for running locally. I made a web based battle chess game with sound effects running on a pi harness in about 15 minutes, complete with a (very unskilled) “AI” (not an llm) that plays against you if you want!
Just tried it, really cool and works as advertised.
dflash2 is atmost 10-20% above mtp, it won't get you to 30 tps
These numbers are a lot lower than I expected from such pricey hardware.

I'll stick with my old AMD datacenter cards thanks. An Instinct MI50 32GB would cost around $550 or so. It has 1TB/s vram bandwidth thanks to its HBM2.

any numbers to share? also, what inference engine do you use and with which api?
I use Llama.cpp and OpenAI-style API of course (who doesn't, unless you use Ollama perhaps?)

Ollama sucks, they dropped ROCm support for my card out of the blue, and the Vulkan replacement wasn't ready. The developer just shrugged and didn't care about the people affected so I moved to llama-server. Also, here's some other reasons to avoid ollama: https://sleepingrobots.com/dreams/stop-using-ollama/

And really, ollama are trying to sell their cloud platform and are always running behind llama.cpp in new features like the KV Cache quantisation.

So numbers, I don't really measure it from day to day, I'm not really interested in benchmarks at all. But I just did a quick test, llama 3.1 8b q8_0 does about 70tk/s (with Vulkan up from 45 when I was still running ROCm), and qwen 3.5 9b q8_0 does about 45, I think that includes its time spent thinking, not sure. Overall I'm happy with that.

Smaller sizes and quants are of course faster than this.

It's not terribly much faster than a Mac Studio, but it's one hell of a lot cheaper of course. And for the price of a Mac Studio you can get a lot faster hardware. The only thing where the studio excels is if you need large amounts of vram like 64GB and up and run the models to go with that (I run multiple small ones and spread them over multiple cards)

I am also running some old AMD datacenter cards, 2x MI25 in my case. Getting around 30 tokens/second with short context.

Tensor parallel in llama.cpp using RCCL (disabled by default in llama.cpp for some reason). Surprisingly, for these cards HIP is actually faster than Vulkan, unlike the 9070 XT where Vulkan still wins.

ROCm nightlies do actually support these old cards, just not the ROCm stable releases.

It's a software issue.
How so? I've had no issues running llama.cpp with vulkan compute on my AMD graphics card (9070 XT). llama.cpp exposes an OpenAI API endpoint, which seems to be the lingua franca for AI applications. Does llama.cpp not work with datacenter AMD graphics cards?
Yes it works great with them through Vulkan. ROCm is more hit and miss, especially because the cards which are affordable (and PCIe, the new ones aren't PCIe compatible!) are already fairly old and have already been dropped by ROCm.
Because OP is running it on an M3 Ultra with Ollama. He'll get much better perf with llama.cpp or MLX, both of which Ollama wraps, albeit very poorly.
This low-effort slop post is misleading. It suggest that 3.8 is half as fast as 3.6, but this is almost definitely because MTP isn't enabled by default. The two models should perform the same. When you get AI to think for you, you lose.
Apart from learning, I cant see the point of spending that kind of money and energy to get such awfull token generation speed. Assuming memory bandwith is the bottleneck, is it just a matter of time until we start to see hbm4 based chip able to run qwen3.8 for normal people running at more than 500 token / seconds? Is memory speed the only technological bottlenecks that prevent us from having fast local model?
A slow AI tasks that can process my self hosted personal journal, my medical history in fasten, my diet and exercise in Mealie and Sparky Fitness to offer insights once a day or even once a week is better than nothing. Because I would never upload that data to any of the AI companies.
I have 2x 3090s and I get 260 tokens/sec peak (110 avg) with Dflash2 and about 1500t/s prefill with a Q4 quant of Qwen 27b. I didn't buy an overpriced Apple product and it performs much better. It's extremely reliable for Agentic coding and I can run 2-3 simultaneous agents with a full ~260k context window.

I realize the price of NVIDIA has gone up but there are plenty of GPU options from others like AMD and Intel with reasonable performance.

Pretty impressive how the Mac ends up less expensive than the Strix Halo boxes, at least here in Ireland. A 128GB Mac Studio with an M5 Max (the Ultra can only have 96 or 256GB) still costs less than the "GMKtec EVO-X2" or the Nvidia DGX Spark with similar performance.
That seems unlikely, even for Ireland. Current prices in USD:

128 GB 1 TB M5 Mac Studio: $5399

128 GB 1 TB GMKtec EVO-X2: $3499

128 GB 1 TB Framework Desktop: $3748

That's not right. Mac Studio M5 Max 128gb, in Apple Store Ireland, starts from 5,909e (~6,847 USD; tax included). On the other hand, GMKtec Evo-X2 128gb can be currently purchased for 3,299e (~3,822 USD; tax included) from GMKtec online store (free shipping from DE warehouse). It's quite the difference in price.
It’s the price I got from amazon.ie
I figured it's dodgy amazon ie listings. DGX Spark/Asus Ascent GX10 prices are inflated there too. They'll also be happy to sell you an RTX 5090 for 10 grand.
Qwen3.8 has the same architecture and the same parameter count as Qwen3.6. Something is not right with the GGUF if it's 2 times slower.
I am also seeing slower speeds, roughly the same ballpark, sometimes even lower - 10-11 tok/s on M5 Max. If there is that ONE version (GGUF or MLX) that runs roughly as fast as 3.6 used to run, please let me know.

What would be incredible is the 3.8 35B MoE version too, I can run 3.6 with 60 tok/s which is a really, really nice speed.

Ollama? Running a q1 quant? I don't think the writer knows what the are doing here to be honest.

You will get better information cruising r/localllama for about 10 minutes.

This is a problem that Hugging Face could easily help solve.

Crawling through Reddit or forums to find the right incantation to run a model is frustrating.

AI is a killer adblocker ! imagine connecting it to instagram and curating all the images that you actually care about removing all ads !
That seems like an expensive way to block ads considering traditional ad blockers have been doing that with a tiny fraction of the compute for decades.
Have you used a social network in the past year? Facebook, YouTube, Pinterest, et al. have become increasingly hostile to blockers with techniques that break pattern-matching on hostnames and element selectors. It may be inevitable that adblockers move toward visual identification.

  > Facebook, YouTube, Pinterest, et al. have become increasingly hostile to blockers
maybe at some point we should just stop visiting these sites if they are so hostile...

we keep going to them despite all of this and it just serves to delay any kind of alternative getting pickup (not an easy task, i know)

We go to them because that's where the friends and organizations we care about have chosen to post their content.
Slop; all those LocalLlama threads "you" read had #s too. And Q1 quant? Why? You have the memory...
The cheapest card that will run this model very well is a ln unlocked CMP 170HX. But you can run it on a 3090. I run it on an old spare A6000 Ampere. I think I wouldn’t use anything lower than 60 tok/s though, which you can get with MTP etc. I just use a full vllm stack but some people see a lot of speed with ninfer (there are non 5090 ports).

The large RAM Macs are unusable for inference of dense models as of now. Token generation is too slow.

Yea Qwen3.8 wasn't fun to use on my 64GB M4 Max either (better than these numbers though), so my new daily driver is Ornith-1.5-35B-A3B-MLX-4bit. I recommend giving that a whirl if you're on similar hardware, it's definitely better than Qwen3.6 35b-a3b which was my go-to before.

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B

Can you tell more about your experience with Ornith? I've come across it, and the benchmarks on its landing page are unbeleavably good. But it isn't featured on more established benchmarks like ArtificialAnalysis and it's not on OpenRouter, so I wrote it off as scam.
I tried Ornith-1.5-35B-A3B-MLX-4bit on a 32G M2-Pro mac mini with pi and OpenCode. I didn’t get very good results coding in Python, Racket, and TypeScript. I have seen several positive comments like yours so I was probably doing something wrong. I amgetting the new 64G mac mini in 4 weeks, and I made a note to try Ornith-1.5-35B-A3B-MLX-4bit again.
I'm running the 5 bit quant on a 3060 12G. Get a handful of tokens/s. Prompts take 10-20 minutes to complete but it works.
[flagged]
is it really worth to run urself? the watts drsin and that 100% gpu usage
Short answer: no

Long answer: some dude on youtube calculated that even if electricity was free it would take you 37 years to get back your investment.

Slightly longer answer: ... but if you absolutely must have local LLM, a GPU with 24-34 GB VRAM is much cheaper and faster than a Studio

FWIW, on my Mac Studio I get ~24-27 tok/s generation between 0-16k context in - that's on the Q6_K GGUF with speculative decoding on. I have spent zero effort optimizing/improving this so far but will be trying the 4 bit MLX next (I've tended to find models drop off somewhat below 6 bit but maybe that isn't the case nowadays).
Peaks at over 300 tok/s on my 5090 with Dflash2. At maximum context (252000 or so) still get 60 tok/s.
I've been thinking about buying a system to run LLMs locally but the price for one that'll run Qwen3.8-27B well is quite offputting to say the least.

What I've been looking at instead is inference providers that use TEE and E2EE to provide cryptographic guarantees that my prompts and responses are only visible to me and the GPU itself.

Despite their docs and assurances of what their guarantees mean, I'm having trouble getting to a point where I'm actually comfortable trusting them with secrets though. Phala for example seems to be E2EE only to the gateway and will then forward prompts to (potentially third party) providers.

Has anyone been down this path and found a provider they feel safe with?

Not suggesting a provider but if you're willing to get a Chinamod GPU, go get RTX 3080 20G or RTX 2080 Ti 22G. Get a couple of them and you can probably run Qwen3.8-27B at a reasonable speed.

I've got a RTX 3080 20G for $450 like a year ago. With llama.cpp, bf16 kv cache, kv cache offloaded to RAM, Qwen3.8 27B UD-Q4_K_XL, single RTX 3080 20G, I got 10 tok/s initially and it dropped to 5 tok/s at 50k context. I'm thinking of getting another Chinamod GPU.

Building a dual Chinamod GPU machine would only cost like $1000~$1500. It's not too bad compared with the alternatives!

A year ago was prior to the memory cartel. You're probably off by 2x.
I've checked. I can still get a 3080 20G for $525, or 2080 ti 22G for $380. The price has increased but not too bad.

Probably the modding had prevented it from gaining too much price increase.

> single RTX 3080 20G, I got 10 tok/s initially

Add $200, buy a used 3060 and you won't need to offload cache to RAM. Yo'd have like 50-60 t/s with MTP enabled.

aww. I used to have a 3060 and I upgraded to this 3080 20G.

Maybe it's time for me to get a 3060 back again. :P

Or, theres the option of a 32gb v100 (about $600USD on taobao etc), where you can get 1200 of prefill at 80 of decode: https://github.com/geoffwatts/ninfer-v100 - that's _really_ cheap inference, and it's not a modified card - you just need to add a blower or water block.
I run qwen3.8 on 5060ti 16G RAM. If you can do with Q3 and 64k of context. It works great at 25t/s
the fact that this author cannot get qwen3.8-27b run at the same speed as qwen3.6-27b, says the article is not worth reading. the author does not know anything about how to run local AI. 3.8 and 3.6 are the same model with different weight.

both tg and pp speed are so terrible on author's machine.

>I realized I’d been sitting on the one thing most of those threads were missing: a machine that can actually run it properly, and time to measure it.

This was written by AI I take it. The one thing, huh? All those idiots running it on their RTX 6000s at 140tok/s must be feeling pretty stupid for not getting a Mac Studio instead.