Not long-horizon coding but for a lot of other things
like batch processes with structured outputs, quick checks/fixes, making sense of unstructured data etc..
That's a solid 2200 words to spend on operating parameters and conveying state, leaving a generous 700 word window for them to decide and respond in.
When the bonsai/prism 1bit models dropped and I saw how many prompts a minute I could get from a dusty m2 mini I started hooking it up to all sorts of shit, like a traffic simulator that translates the car state/surroundings/immediate goal into text, it responds with a seqeunce of actions defined in the system prompt, which then get translated back into NPC input.
What I was hoping for here was that it would result in fucking chaos, all sorts of stupid decisions and epic car accidents. I cannot overstate my disappointment (and terror) when they were perfectly reasonable, safe drivers. I had to cut the tire grip by 75% without telling them and make them control twice as many cars to delay their ability to respond before I saw anything resembling an enjoyable traffic accident.
> We currently have 14x nodes of CG480-S6053 ready to ship.
Oh, okay, so this is an ad.
I do still think it's well written and interesting... But if anything, it's just making me more curious about the newest generation of M5 Ultra. (and less and less interested in PCI-E Gen 5 anything)
I'm curious whether actual inference workloads actually push to 600W (and not 350W) and what the last 250W get you. Rare is the (generic gpu) workload where I get >5%, some rare light inference benchmarks up to 10%...
Yes that was the second part of my comment. I've been benchmarking many things AI or not on those boards and while one can sometimes get to 600W I haven't seen more than 10% gain for those last 250W. If someone has a workload that gets more from those watts I'd be interested. For now on anything I run (full-capacity, continuous, batched...) capping at 350W seems better in bang-for-bucks when including energy costs (from the GPU + added cooling).
I mean, I can't afford a new F80, but you don't see me going out of my way to post on /r/Ferrari about it. I seriously do not understand the point of these posts.
The topic of the story is running 8x RTX PRO 6000s, so whining about how much they cost or what you could/should buy instead is completely off-topic.
My bad, didn't mean to derail things with my little quip, for what its worth I did read the article and it was an interesting read. Overall I do hope for more smaller models in the future which I can run on my 3090 that deliver similar performance, or for hardware costs to drop in the next few years to prosumer levels.
4x RTX 6000 Blackwell cards is a good place to be if you can't swing 8 of them, or if you don't have the power or cooling to run that many. A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision [1] from a US-standard 120V 20A circuit and give you a better pelican than Fable 5.1 [2]. What's not to like?
Flash will run on 4x, but at the time I ran that test there were no 4-card quants for the full 744B-A40B 5.3 model. There are now, though, with KLD figures close to the FP8 level. I need to do some more benchmarking to see if they live up to the hype.
These people have zero idea what they're doing. Not a single mention of pipeline parallelism that would actually make the setup useful to run a big model.
I can’t stand it. Very engineering-y over specified formal language around a complete lack of core understanding. Is damaging other people read this and try to learn things from it.
Makes you realize how insane the M5 Ultra Mac Studio is. 1.2GB/s bandwidth 512GB memory. Its rated max power draw is just 480W. And it also has amazing M-series CPUs. It costs less than just one of these GPUs which each take 700W to run.
There is comparison actually. I spent all day researching this a few days ago.
Memory-wise, the RTX PRO 6000 can barely hold two 1M context Qwen 3.8 27B models at 8 bit quantization at the same time. The 512GB M5 Ultra Mac Studio could hold around 14.
At such high concurrency, batch performance is usually limited more by memory bandwidth than compute. The RTX PRO 6000's memory bandwidth is just 50% faster than the M5 Ultra.
So yeah, I think if we are talking about many short context requests, sure. But if you are chewing through a backlog of coding tasks with Qwen overnight, they might actually be comparable.
Im sure in Nov when the M5 Ultra comes out we'll see a lot of interesting benchmarks.
Pass. When articles keep mentioning models like DeepSeek R1, or Llama 3.1, or Qwen3 32B, it is a pretty robust indicator of AI slop. LLMs love to suggest DeepSeek R1, etc. - training data cut-off?
No person with real practical experience and real use cases will be using these ancient models as examples, when talking about local LLMs.
Utterly awful article. R1? Llama3.1? Not being able to serve larger llms on 8(!) RTX pro’s? You can literally run open weight SOTA models with relative ease. Even 4 GPUs get you there with a bit of elbow grease and compression. Pure slop.
34 comments
[ 0.25 ms ] story [ 38.3 ms ] threadstopped reading after that. What 4k context would be usable for?
That's a solid 2200 words to spend on operating parameters and conveying state, leaving a generous 700 word window for them to decide and respond in.
When the bonsai/prism 1bit models dropped and I saw how many prompts a minute I could get from a dusty m2 mini I started hooking it up to all sorts of shit, like a traffic simulator that translates the car state/surroundings/immediate goal into text, it responds with a seqeunce of actions defined in the system prompt, which then get translated back into NPC input.
What I was hoping for here was that it would result in fucking chaos, all sorts of stupid decisions and epic car accidents. I cannot overstate my disappointment (and terror) when they were perfectly reasonable, safe drivers. I had to cut the tire grip by 75% without telling them and make them control twice as many cars to delay their ability to respond before I saw anything resembling an enjoyable traffic accident.
Oh, okay, so this is an ad.
I do still think it's well written and interesting... But if anything, it's just making me more curious about the newest generation of M5 Ultra. (and less and less interested in PCI-E Gen 5 anything)
https://github.com/aikitoria/open-gpu-kernel-modules
The hardware supports it, but Nvidia disabled it if the driver detects cheaper cards.
Sure, let me just buy $60,000 worth of GPUs to run a *quantized non-frontier model*.
For that price you could:
- put a down payment on a home in a large % of the US
- buy a brand new car in cash (possibly two!)
- take a long sabbatical and travel the world
- pay all 4 years or your child's college tuition
The topic of the story is running 8x RTX PRO 6000s, so whining about how much they cost or what you could/should buy instead is completely off-topic.
1: https://huggingface.co/local-inference-lab/GLM-5.3-NVFP4
2: https://crimson-jeri-74.tiiny.site/
It actually runs fine at FP8 on this hardware too, with the full 1M context.
No comparison.
None.
Memory-wise, the RTX PRO 6000 can barely hold two 1M context Qwen 3.8 27B models at 8 bit quantization at the same time. The 512GB M5 Ultra Mac Studio could hold around 14.
At such high concurrency, batch performance is usually limited more by memory bandwidth than compute. The RTX PRO 6000's memory bandwidth is just 50% faster than the M5 Ultra.
So yeah, I think if we are talking about many short context requests, sure. But if you are chewing through a backlog of coding tasks with Qwen overnight, they might actually be comparable.
Im sure in Nov when the M5 Ultra comes out we'll see a lot of interesting benchmarks.
No person with real practical experience and real use cases will be using these ancient models as examples, when talking about local LLMs.