This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some range of $/MTok for a 3T model. Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".
Also interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.
Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)
AISI is capped at 100M tokens and K3 is less token efficient than Anthropic/OpenAI models. There is an argument to be made, looking at AISI results, that with uncapped tokens it would be just slightly behind the closed weight players.
"Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing"." - The thing is though, that Anthropic and OpenAI have breakthroughs in optimization that none of the Chinese labs have. So it'll demonstrate a ceiling, but not a floor.
They cant push it too low. The license agreement it is released under wont allow it.
> If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
Strange communists, giving away such an expensive model to the public.
On the other note, can't wait to see 1bit quantisation soon and how it performs in benchmarks, if it performs really well in benchmarks, would be very good news for GPU hosting providers, to offer "Opus 4.5 level model at the cost of Haiku 4.5"
I think China publishing this stuff is more about prestige. The US has had export controls that make it illegal to sell Nvidia chips, and other AI hardware to China, and this is China saying "yeah, whatever". Also, it weakens western tech companies' position in AI, and pushes CCP bias perniciously. Building your product/company on top of a text-generation model that favours the CCP position on everything is just peak propaganda.
We already know that competition brought GLM 5.2 prices down roughly 45% since its release on June 16th (1.5 months ago), and the price downward slope is probably still going (I've been checking regularly and new providers keep fighting on price, I don't think prices have settled yet). For reference : https://openrouter.ai/z-ai/glm-5.2#providers
I saw arguments like "Providers cannot price less than their costs" in other comments. In economics, it's generally admitted that they shouldn't price less than their marginal costs, i.e. in their case roughly the cost of electricity, since a lot of these datacenters are not at capacity in terms of graphics cards usage (speculation since it's very easy to rent a GC for a couple hours on some providers). My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong
Seminanlysis is estimating sub $1 cost per MT for ~2Trillion models. The numbers change based on throughput and quant, but it is conceivable that provider costs at scale are low enough that even $2.42 per MT on GLM 5.2 (current best price) is margin positive by a wide margin.
is there a realistic way to distill 2 consumer hardware friendly models with max ~200B and ~20B? Qwen did it, but would it be possible for 3rd parties (unsloth etc)?
I feel like most hardware to run LLMs on is shaped wrong for individuals.
It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards would be kinda useful (albeit NVLink or equivalent would need to be commonplace).
Obviously nobody is running Kimi K3 locally without an insanely beefy homelab and lots of money to burn, but running GLM 5.2 would be cool at like ~100 tokens per second for a single session and maybe ~60 tokens per second with N subagents.
Post-crypto, the GPU manufacturers took the proactive move to use VRAM to segment the market for the purpose of price discrimination. Sure, data centers will pay vastly more for GPUs, but Nvidia knows that the PC market is steady and reliable. They could get the best of both worlds by kneecapping their consumer cards to tiny amounts of RAM, to dissuade the cloud providers from scooping up all the consumer cards, and then charging the two segments wildly different amounts for what amounts to the same hardware (back when the cost of RAM was negligible)
You can run GLM 5.2 at a reasonable quant on a cluster of DGX Sparks. On the link below, the guy is getting decent performance and links to a repo. An 8x Spark cluster can run it even better. A Spark cluster is about as good as it gets for (a) runnable at home, (b) "affordable" hardware, (c) not going to kill your power bill.
Given the frontier-level capabilities of Kimi K3, I'm wondering if it's possible to extract the core capabilities (fundamental reasoning and tool calling) of the model into a smaller one that consumer devices could run? Not sure exactly how, but either by heavy distillation or some other surgical method since Kimi has a Mixture of Experts architecture.
I think it's very valuable to have a smaller model that doesn't have any domain knowledge or facts built into its weights, but given the right context, could accurately reason about what to do and use the right tools.
I'm aware of colibri [1], but so far I've only seen extremely slow performance.
> smaller model that doesn't have any domain knowledge or facts built into its weights
I’m not sure it works this way. Language modelling itself is a “domain”, and if it didn’t have a grasp language it wouldn’t be able to do anything else.
I also think a “reasoning engine” that had to reason through everything from first principals would likely be extremely inefficient.
It’s good went models have domain knowledge and expertise - and all of their reasoning flows downstream of that
There's no going back on this. This is putting a very capable intelligence in the hands of the masses. Private companies in the US are aching for Trump's protectionism but it'll do nothing. The hardware needed to run this is ofc prohibitive, but actually putting it out there feels like a 'RSA source code on t-shirt' moment for humanity.
103 comments
[ 2.8 ms ] story [ 50.8 ms ] threadAlso interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.
Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)
Sounds like I'm buying a lottery ticket this week so I can drop $800k on hardware.
If you're going to open source your model, why would you set your own price high enough that other providers could easily and profitably undercut you?
> If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
We won't be able to run this ourselves, but many providers can.
On the other note, can't wait to see 1bit quantisation soon and how it performs in benchmarks, if it performs really well in benchmarks, would be very good news for GPU hosting providers, to offer "Opus 4.5 level model at the cost of Haiku 4.5"
I saw arguments like "Providers cannot price less than their costs" in other comments. In economics, it's generally admitted that they shouldn't price less than their marginal costs, i.e. in their case roughly the cost of electricity, since a lot of these datacenters are not at capacity in terms of graphics cards usage (speculation since it's very easy to rent a GC for a couple hours on some providers). My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong
It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards would be kinda useful (albeit NVLink or equivalent would need to be commonplace).
Obviously nobody is running Kimi K3 locally without an insanely beefy homelab and lots of money to burn, but running GLM 5.2 would be cool at like ~100 tokens per second for a single session and maybe ~60 tokens per second with N subagents.
How unfortunate.
https://www.youtube.com/watch?v=nbHOBvLlypY https://www.youtube.com/watch?v=PV89U-PNUNA
I think it's very valuable to have a smaller model that doesn't have any domain knowledge or facts built into its weights, but given the right context, could accurately reason about what to do and use the right tools.
I'm aware of colibri [1], but so far I've only seen extremely slow performance.
[1] https://github.com/JustVugg/colibri
I’m not sure it works this way. Language modelling itself is a “domain”, and if it didn’t have a grasp language it wouldn’t be able to do anything else.
I also think a “reasoning engine” that had to reason through everything from first principals would likely be extremely inefficient.
It’s good went models have domain knowledge and expertise - and all of their reasoning flows downstream of that
Now the US government has 5 hours left to (attempt to) stop the release. (and save Anthropic)
Let competition run its course and the market (not government) determine the winners and losers.