I have a framework desktop w/ 128GB that I bought last Christmas and if I’m looking at it right it costs $2000 (CAD) more now because of the RAM shortage (and in any event is apparently out of stock). Would love to have 192 GB but I’m not sure I can justify buying another.
Does anyone have a sense of how this might progress, e.g if I can get a 256 or 512 GB in a year if I wait. In any case I’m jealous this exists and I don’t have one.
One last thing, I assume this isn’t exclusive and there will be other builds with this same config same as current Strix Halo?
Likewise here, my setup costs 2k+ more eur as well. Given that memory bandwidth remains the same I don't think it's worth buying if you have the 1st gen framework desktop. Usecase caveats apply.
The market research I have seen indicated DRAM prices will not come back to “normal” until 2030 at least and probably 2031. You can wait, but it will probably be longer than a year
Unfortunately we don't have any indication of memory pricing coming down in the next year, and many indications pointing to memory pricing continuing to increase pretty substantially over the next six months, especially for the LPDDR5X we use in Framework Desktop.
Yes, others should be able to do this too, and the change is larger memory chips.
The next big step is Medusa Halo, which will have a 384-bit LPDDR6 interface. Those should be able to support 256GB at release, with 512GB coming later with larger chips (But I don't know to what extent people should trust the memory vendor roadmaps.) I'm not sure if they will be out in a year. Probably will be in a year and half.
Mmm.. AI395 at 3.5k ( and out of stock mind ). And I still want to give it a shot when it is out. I suppose I just explained to myself the ridiculous jump in prices. The demand is crazy.
Memory bandwidth certainly is the bottleneck on my 395+ 128GB. It's been a great server. Think I'll wait another year maybe hopefully the prices come down and there's one more gen of hardware and local llm gets even better. Excited for this!
Im less interested in the memory than the memory bandwidth. The current system with 128GB can load pretty big models, but its meaningless unless you want to wait 40 minutes per prompt.
I've had best results with Qwen3.6-35B-A3B, which uses 40GB of memory, but only uses 3 billion parameters per token which helps with throughput.
Until memory bandwidth significantly improves I just can't see myself wanting to use all that memory. Unless it's just to keep a wide variety of models in memory.
For those wondering, the AI Max 395 has around 256GB/s of memory bandwidth, whereas this new 495 has 273GB/s. So a very modest improvement in bandwidth.
Woof. What's that going to cost? And will we actually be able to buy one, because the 128GB model has been sold out pretty much since the LLM craze started...
We can't share pricing yet, but certainly will before we open orders. The 128GB has been selling well, and we just re-opened orders with stock landing within two weeks.
I understand their reasoning but it’s still a shame Framework opted for a proprietary motherboard instead of mini-ITX with a socket CPU and GPU. I value repairability over a little bit of extra performance.
IIRC, Framework did ask AMD to explore a more repairable approach with socketable CPU and RAM for these chips and they did but came back and said it wouldn't perform well enough that way.
There is no socketable Unified Memory platform. The entire point of this SoC is that it has Unified Memory. The Framework Desktop has a Mini-ITX board, but you MUST use a solderable SoC with HBM stacked memory to get a Unified Memory outcome. For local AI workloads, that's the only part that matters, which is the entire point of wanting 192GB of RAM available on an APU.
I like the repairability and modularity of my Framework 13 laptop, and I still bought a maxed out MBP M5 Max because for local LLM, unified memory is all that matters.
Any SoC currently available does not use HBM for its unified memory. That's just stacked DRAM, wired somewhat more optimally because of shorter distances, making it able to go a little bit faster, and maybe wider. But not as massively wide as HBM is, with its load of many lanes. This may shift in the near future, in favor of more traditional stacking, cost/production-wise, vs. real HBM with massive speed improvements, at much higher costs.
I typed my (masked) email into the input, pressed tab and enter, it redirected me to some cloudflare captcha marketing site. I thought at first it didn't want the masked email, but then it worked when I clicked with my mouse. Apparently there are 3 invisible focus targets in between!
Dear people who create websites, these things are important, they should work!
the thing about ECC memory is that it doesn't support reporting of ECC errors, so it ends up just hiding the fact that your RAM has started to go bad. this is good if you replace RAM every n years, but not if you intend to use the RAM until it starts degrading.
Tl;dr: think hard about what you’d use this much RAM for in a desktop setting and do some research about how people like the 395 for your use case. Not all use cases work well.
Anyone who thinks they are going to serve some 100+ GB LLM locally, remember that memory bandwidth becomes a key limitation for large models. While you might be able to load a model, token generation can be very slow. MoE models like Qwen3.5-122B-A10B work decently fast, but dense models of a decent size are slow and you won’t want to use them.
I’ve got a 395 system, and found that I’m quite happy with Qwen3.6-35B-A3B, generating at around 50 t/s, but the dense 27B model is 20-25 t/s and that’s the lower limit I’m willing to tolerate. So a 70b dense model is just not going to happen. That means you can’t really use that much RAM.
A reasonable use case is to have multiple smaller models loaded- you can have an image generation model loaded along with the text model. Or you can use this computer for development simultaneously with serving LLMs. Those ideas work okay. But trying to load up a single giant model is going to test your patience.
> A reasonable use case is to have multiple smaller models loaded- you can have an image generation model loaded along with the text model.
Lemonade server updated yesterday for an Arena model / ai routing.
---
The Lemonade local AI server allows easily deploying LLM workloads across Ryzen AI NPUs, Radeon GPUs, and CPUs with ease under both Windows and Linux. The most significant change with Lemonade 11.5 is the completion of the Lemonade Router that can be used for automatically routing queries to relevant models based on defined policies -- or even LLM-as-a-router for using a small LLM to in turn determine which model to route a particular request.
The Lemonade Router allows steering requests based on rule, classifier, semantic similarity, or LLM-as-router policies. The policies are defined in JSON files or can also be authored via the Lemonade GUI. More details on these router capabilities can be found via this commit.
Lemonade 11.5 also introduces a server-side job engine to let clients post multi-step recipes and managing them via /jobs endpoints. From these endpoints the muilti-step recipes can be paused / interrupted / resumed / deleted. And another one is that the lemond daemon can now act as an MCP client host to connect external stdio MCP servers.
Somewhat on-topic, if anyone here's received their Framework 13 Pro yet (which I believe began shipping this month), I'd be happy to hear a first-hand review.
This machine needs better networking options to be useful, such as built-in 100Gbit QSFP28. There is not enough space between the x4 PCIe slot and the power supply for a NIC to fit, and even if there was, you'd have to cut a hole in the back of the case to get at it.
42 comments
[ 2.7 ms ] story [ 24.5 ms ] threadDoes anyone have a sense of how this might progress, e.g if I can get a 256 or 512 GB in a year if I wait. In any case I’m jealous this exists and I don’t have one.
One last thing, I assume this isn’t exclusive and there will be other builds with this same config same as current Strix Halo?
The next big step is Medusa Halo, which will have a 384-bit LPDDR6 interface. Those should be able to support 256GB at release, with 512GB coming later with larger chips (But I don't know to what extent people should trust the memory vendor roadmaps.) I'm not sure if they will be out in a year. Probably will be in a year and half.
> 192GB coming soon
> The most powerful Framework Desktop yet is coming soon with an AMD Ryzen™ AI Max+ PRO 495 processor and 192GB of LPDDR5X memory.
Abd wake up without crashing?
I've had best results with Qwen3.6-35B-A3B, which uses 40GB of memory, but only uses 3 billion parameters per token which helps with throughput.
Until memory bandwidth significantly improves I just can't see myself wanting to use all that memory. Unless it's just to keep a wide variety of models in memory.
Feels like they blatantly ripped that off the title of a Cory Doctorow book [0].
[0]: https://en.wikipedia.org/wiki/The_Internet_Con
I like the repairability and modularity of my Framework 13 laptop, and I still bought a maxed out MBP M5 Max because for local LLM, unified memory is all that matters.
That aside, ever heard of https://en.wikipedia.org/wiki/CAMM_(memory_module) ?
How does a CEO read this and not immediately fire their entire marketing team?
Dear people who create websites, these things are important, they should work!
And into the trash it goes.
Refer to https://news.ycombinator.com/item?id=48969530
What would you say is the likelihood of an LLM hallucination vs. a correctable RAM error?
Anyone who thinks they are going to serve some 100+ GB LLM locally, remember that memory bandwidth becomes a key limitation for large models. While you might be able to load a model, token generation can be very slow. MoE models like Qwen3.5-122B-A10B work decently fast, but dense models of a decent size are slow and you won’t want to use them.
I’ve got a 395 system, and found that I’m quite happy with Qwen3.6-35B-A3B, generating at around 50 t/s, but the dense 27B model is 20-25 t/s and that’s the lower limit I’m willing to tolerate. So a 70b dense model is just not going to happen. That means you can’t really use that much RAM.
A reasonable use case is to have multiple smaller models loaded- you can have an image generation model loaded along with the text model. Or you can use this computer for development simultaneously with serving LLMs. Those ideas work okay. But trying to load up a single giant model is going to test your patience.
Lemonade server updated yesterday for an Arena model / ai routing.
---
The Lemonade local AI server allows easily deploying LLM workloads across Ryzen AI NPUs, Radeon GPUs, and CPUs with ease under both Windows and Linux. The most significant change with Lemonade 11.5 is the completion of the Lemonade Router that can be used for automatically routing queries to relevant models based on defined policies -- or even LLM-as-a-router for using a small LLM to in turn determine which model to route a particular request.
The Lemonade Router allows steering requests based on rule, classifier, semantic similarity, or LLM-as-router policies. The policies are defined in JSON files or can also be authored via the Lemonade GUI. More details on these router capabilities can be found via this commit.
Lemonade 11.5 also introduces a server-side job engine to let clients post multi-step recipes and managing them via /jobs endpoints. From these endpoints the muilti-step recipes can be paused / interrupted / resumed / deleted. And another one is that the lemond daemon can now act as an MCP client host to connect external stdio MCP servers.
https://www.phoronix.com/news/AMD-Lemonade-11.5
https://lemonade-server.ai/