I meant serious players often do serve on a single node. They can beat API pricing as well. Multi-node can add gain, but also adds a lot of deployment complexity so its just not always possible or optimal.
> Not to mention no one serious is serving this on 8xB200 instead of multiple nodes: the vast majority of Moonshot's inference work is focused on PD-disaggregation The GPU price discourse is absurd, but many are serving…
I think it's fair to expect an extensive review of an article before publishing. Not everything has a set of serious flaws.
https://chatgpt.com/share/6a6f09ff-2830-83ea-9578-de3016cfae...
Wafer is making themselves synonymous with slop in the inference space. Exaggerated unfair comparisons in all their results, twitter hype posts with alarm emojis etc. > $2.50/GPU-hr for the MI355X, $6.00 for the B300,…
There is barely any effort, a simple GPT5.6 sol pro query rips the post apart.
I meant serious players often do serve on a single node. They can beat API pricing as well. Multi-node can add gain, but also adds a lot of deployment complexity so its just not always possible or optimal.
> Not to mention no one serious is serving this on 8xB200 instead of multiple nodes: the vast majority of Moonshot's inference work is focused on PD-disaggregation The GPU price discourse is absurd, but many are serving…
I think it's fair to expect an extensive review of an article before publishing. Not everything has a set of serious flaws.
https://chatgpt.com/share/6a6f09ff-2830-83ea-9578-de3016cfae...
Wafer is making themselves synonymous with slop in the inference space. Exaggerated unfair comparisons in all their results, twitter hype posts with alarm emojis etc. > $2.50/GPU-hr for the MI355X, $6.00 for the B300,…
There is barely any effort, a simple GPT5.6 sol pro query rips the post apart.