71 comments

[ 0.32 ms ] story [ 41.1 ms ] thread
One of the strangest things in the AI industry is 'tokenomics'. It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. This pattern has continued across various labs/providers for years--there is a continuous see-saw of pricing that doesn't seem related to anything.

So what open weight models do is at least provide a baseline of inference cost to add some sanity to the price markers. And of course predictability too--if you really want Kimi K2 instead of K3 you can still use it.

So the competitive pressure and predictability offered by open models is helpful for users

Way over complicated for what most users need? Huh?
(comment deleted)
> The government should use procurement to create demand for portable, interoperable systems rather than permanent dependence on one API vendor.

Now, here is an idea that I have not heard before... and I think there is some merit to this. This is also the sort of thing that a state (looking at you CA, CO, IL, NY) could do, instead of just the federal government.

I dunno about the other states I feel like we need a better coalition lobbying California lawmakers about AI innovation because I feel like they're doing a lackluster job.

For example, Buffy Wicks introduced AB 2023 (passed the Assembly) which will effectively cause chatbot operators to ban minors because of the huge liability risk that the law introduces. This is a great way to kneecap innovation by excluding curious children and teens who are often the most innovative. California already has laws on the books to address chatbots encouraging self harm (BPC §§ 22601-22606) so I don't know why on earth they're doing this.

Then there's AB 2169 introduced by Lowenthal, which would have mandated interoperability between chatbot platforms to help people migrate to competing ones easier. I thought this was awesome but it didn't even get a vote in the Assembly.

Maybe the newly-created Little Tech Association can help here.

why would any software want to have Kubernetes moment? can't count how devop I know that is confused by it
Is anyone using open weight models for agentic coding?

What is your stack (harness, model) and how much do you pay per month?

How would you compare your experience to a typical subsidized plan like Claude Code + Pro plan?

I’m asking because i keep hearing that open weight models are cheap and efficient - is that really the case in practice?

I'm using Qwen 3.6 27B on a macbook with Pi. It's alright, it runs fairly quick (40 tps for quality version, 80 for the fast). It doesn't tend to one shot things but I'm generally comfortable fixing the bugs myself afterwards or prodding it a little bit. I find the harness matters a lot. "Continue" (the vscode extension) worked horribly, OpenCode was ok but its vibecoded internals make me view it as a security nightmare so I'm hesitant to run it, so I've settled on Pi for now.

Claude and ChatGPT are good deals right now, with the subsidies. They produce things faster and better. I guess not cheaper, in that inferrence on my macbook is basically free, although the macbook itself definitely wasn't. My focus on running local is around three principles:

1. I don't want to support surveilance capitalism by giving these companies my data anymore, when I can avoid it. And LLM companies want to vacuum up every detail of your life.

2. I don't find these companies to be remotely trustworthy, and I find them hostile to a healthy society, so I want to avoid giving them money going forward

3. I think they're going to start charging a lot more

GLM 5.2 awq4 via Opencode, it was better than corporate's fave Sonnet 4.8 and a but worse than Opus 4.8. overall quite capable of a lot of what the dev team needed. Not as good as Sonnet 5 but also never runs out of tokens.

The flip side is that it takes four H200s to run, and that will only let you cache context for maybe three users.

Fingers crossed our Blackwells show up and Kimi 3 really releases weights, because at some point devs are spending a significant portion of their salary on tokens and it's somehow cheaper to buy these ridiculous DGX servers and rack and run them.

At least with the Kimi Code plan, its limits are pretty abysmal compared to ChatGPT/Claude plans.
(comment deleted)
Dario is a FUD-spreading douche
In fact more countries should have government funded models. There are some obvious issues in China completely dominating open weights space. Kimi had funding of just $2B and could literally create national security threat. A lot of countries could fund something in the range of few billion for something so important. At the very least US and EU could fund few companies.
Enormous amounts of money is being invested in the development of AI models. Investors expect returns on their investment or they will not continue investing. Open weights make it harder for investors to get their money back, so it harms the industry.

Once the weights are out, it makes no sense to ban them in the US while the rest of the world takes advantage of it. But that doesn't mean developers of frontier models shouldn't take steps to prevent their weights from being stolen.

Sadly until china scales production of hardware it really isn’t economical to run this stuff yourself. It is good it exists though to put pressure against the labs.

Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.

Pre bubble prices (~= “we stop building data centers with subsidized credit / circular loans / hidden debt”), a 128GB halo strix ran for $1400, and 200-ish watts. Four of those in a cluster will run a 1T parameter frontier model:

https://www.amd.com/en/developer/resources/technical-article...

At 7 months of claude code subscription per node, the cluster pays for itself in 28 months. On a 5 year (60 month) depreciation schedule, you can buy two of those clusters for basically break even, so you get two concurrent request streams (each of which can batch, etc).

The next generation hardware has already been announced, and should ship roughly two Moore’s law doublings later. It’s likely its steady state price is <= $1400 USD (2024), and it is faster.

So, once the bubble pops (because the financial machinations eventually will come to an abrupt halt), and the labs stop buying hardware for data centers, local inference will be extremely practical and cheaper than a subscription.

My main question is, when that happens, will UNIX Surplus be selling inference servers for pennies on the dollar (like after the dotcom crash), or are the power requirements too exotic for home use?

The sentiment in this article is nice. But open source software is a weak analogy for frontier models. Principally because software requires zero capital investment (actually zero) while frontier models demand billions. Open models can only survive in the long run if they can (eventually) generate significant cash flows or if they are paid for by governments. Now China essentially has a monopoly on open weight models. And so supporting open source models means either supporting long term economic capture by China or supporting Chinese government control of your intelligence. Both of these outcomes are unequivocally bad from an American perspective. If you live in the valley and benefit from the US venture ecosystem you should be highly skeptical of open weight models. Banning them may very well be the best course of action.
They don’t have a monopoly and if they do, it would only be because the EU and the United States let them or I should say they let all their decisions be made by private companies who scraped the public Internet and now want to fence it in for their up-and-coming IPO.
Open-weight and OSS are wildly different and the article makes a poor comparison.

What's the incentive for the Chinese labs to continue releasing weights 5 years from now? It's not a stable equilibrium and cannot last.

- The lab spending large sums on research and training does not get the inference revenue to fund those efforts.

- Unlike OSS where a single volunteer can keep a project going, training costs run into the $billions.

- OSS is often a two way street where features and integrations are built that the original author benefits from. Open weight models are largely a one way street because the marginal benefit is so much less than training costs.

In the short term, it means Chinese labs can attract talent and, I suspect, funding from their gov. Similar to every other industry the CCP subsidized to take over.

China has no end of money to support these companies. The reason this equilibrium is unstable is that the autonomous agentic coding aspect of the models has been so successfully improved that it will soon be a threat to China state security.
Eventually I think to truly be like Kubernetes, you would need an AI model that has public training data and that a lot of companies collaborate on.

Might make sense eventually. Same logic as companies working on Linux. "An AI model is a business necessity. But making an AI model is so expensive we should not make our own. So let's just use the open one, and contribute the stuff that we need."

FTFA: American labs need to release frontier-grade open-weight models under licenses that startups can actually build on.

oh now i see, the Chinese government is funding the training and release of their best models to pressure OpenAI, Anthropic, and others to do the same for competition's sake. I don't buy it, this seems more like a way to get SOTA models RL'd to comply with Chinese government approved information distribution. If I have to trust a black box of answers to questions i would trust one from a US for-profit publicly traded company subject to market forces over one approved, and heavily subsidized, by the Chinese government.

No, it’s consistent with what China is doing in other markets, which is dumping product to drive others out of business.

I had a shower thought on how to counteract this, specifically related to the AI dumping. If China is losing substantial money on every token, why wouldn’t an adversary try to maliciously increase consumption? This strategy is not really viable against physical goods dumping because demand is finite and there are large environmental costs. Software demand is infinite and the environmental costs are quite low compared to the financial cost to make it, even with ultra cheap Chinese tokens

A "US for-profit publicly traded company" is what brought us the Cambridge analytica scandal. I'd rather not trust them with the next wave of consciousness and opinion shaping technology.

Also, oligopolies aren't famous for being strongly bound by market forces, especially when their decision makers are non ironically being treated as if they were heads of state.

https://www.nytimes.com/2026/06/17/world/europe/g7-summit-ai...

Kubernetes is a system/infrastructure orchestration tool. I completely fail to see how it is comparable to open weight neural nets. In either application or function.

I'm sorry to do that hn comment thing where we all just race to contradict or talk in opposition of whatever was said before. I'm aware. But really guys, was this article really not just a miss?

like most of people I think you didn’t read the article entirely.
Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.

So, any solution to this “problem” must include ALL open-weight models. As far as I understand this is exactly what they intend to do. Axios article linked in the post mentions that. As in this quote:

“The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”

It doesn’t say “Chinese” open-source models. Because they already know that it’s not feasible. Any regulation must cover all the models.

Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution. If a company wants to run an open model in their own servers, they can only use approved and certified pure “American” models. This of course creates a monopoly for the big labs who are authorized to train and distribute such “open” models. A company can fine-tune the model for its own needs but of course can’t distribute the derivative model.

I’m sure there are other solutions but all of them would be equally ugly. Also these regulations can’t be enforced to other countries easily so only Americans will be restricted.

> “The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”

That's going to hit first amendment grounds pretty quick, the same way that software in general did.

The modern version of the decss flag will be a character that says "I think good weights are {...weights go here...}"

They could, however, ban any payment to a chinese entity, or any entity owned by a chinese entity for inference/ai services/etc

It would basically make America behind as every other country would use open, cheaper models for all tasks but the ones requiring frontier models.

And that list of tasks grows smaller every day

> Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.

Historically just asking it about tianment square or getting some random answers turn into chinese (as latest interation of online deepseek likes to do recently) is enough

> Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution.

I am very worried that's where consumer hardware will go to. All so AI companies can license local use of their stuff, and once that happens, less of an incentive to even have model be open.

Possibly even have DRM that counts number of computation done per model in pay per use model

You need to ask what happened in Tienanmen square
It’s not really feasible, in my opinion.

US Gov could make US companies comply, like have Huggingface take down models out of compliance.

But most likely, a foreign-to-US Huggingface replacement would be made and everyone would go there instead. Lose-lose for US.

It's simple, the US government will put any Chinese open model companies on the entity list which blocks any company which does business with the US from also doing business with the Chinese companies. This creates a chilling effect where even if it may be harder to tell, no US company will be able to provide or use any overt Chinese open model and won't even risk trying to go around as the punishments for trying to evade the ban are severe.
> Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them.

It is not possible to 100% ban open weight models getting released in the same way you cannot stop leaks.

Just ask Meta with the original Llama leak.

> So, any solution to this “problem” must include ALL open-weight models.

What about the EU? Would they follow Uncle Sam's order to ban all open-weight models? Lately the EU hasn't been that cozy with american companies: there are EU companies and institutions moving to EU clouds, the EU just fine Google a cool billion, several are switching away from Windows to Linux, etc.

Or is it just the US that'd ban open-weights models, while, say, the EU and Japan would still allow them?

Aren’t the models just software ? You could ban them in say high assurance environments like to be FedRAMP certified you wouldn’t be allowed to use them.

There is no great firewall so banning their import via the Internet is impossible

Regulatory capture? One of the two companies that stands to benefit most has a founder who, together with his wife, donated $25 million to MAGA Inc, a pro-Trump super PAC, and another $25 million to Leading the Future, whose stated mission is to advocate for policies “friendly to the artificial intelligence industry.” Open-weight models aren’t necessarily aligned with the commercial interests of the industry’s largest incumbents.
It'll be pretty easy. Some US gov entity will create a list of models from Hugging Face, decree "thou shalt not provide access to these models", wrap it around some scary legalese for the pirates who try and that will be that.

Though the legalese might not even be necessary. The list alone will make sure that no American company runs these on their servers, including the hosting providers.

Sounds to me Anthropic’s marketing strategy of “AI is so dangerous and we’re the only responsible shepherds” is a great success if this open weights ban will indeed happen.
A Chinese model doesn't think that anything of note happened on June 4th.
It will go as well as the banning of music piracy and BitTorrent sites. They can't even take down those open scientific paper sites.

0% chance they can ban these models practically

> So, any solution to this “problem” must include ALL open-weight models.

I think that is what the "leading AI labs" actually want. They don't care where the open weight models come from, they just don't want to compete with them. The fact that a lot of the open weight models come from china is just a convenient circumstance they can leverage to get the government to give them what they really want.

yep. just like how US gov gatekeeped Mythos/Fable and Ant just repackaged them and call it Opus 5.
You can't ban open weight models either because everyone will sell their models for a penny, making them legally proprietary.

Worst case the models will be sold for a fee that covers the training cost. That would actually be much worse for the big players in the long run.

>I’m sure there are other solutions but all of them would be equally ugly.

Why do you subscribe to some weird "conservation of misery" theorem without proof?

Technically the following must be true in the steady state: the cost of training must be amortizable by its utilization, else no one would train the model.

Technically a computation (like training) can be proven to result in an output (open weights) given the used corpus and a deterministic training algorithm: publish the whole corpus, the (custom modified) deterministic training algorithm, the RLHF datasets etc. And in theory one could verify that the model is derived from the accessible data efficiently: every deterministic calculation can be paused for a thousand (or a million) checkpoints, each checkpoint signed together with the elapsed number of steps since either starting state or last checkpoint whichever comes last before the current checkpoint. This does increase storage requirements. Because it is signed, anyone can recalculate just a small segment of the training computation and verify that the hash on the last checkpoint equals the hash of the proclaimed next checkpoint. Observe that if the source wishes access to a market, they can host the series of snapshots and signatures, and anyone can recalculate a small part of the training, and report a provable difference in outcome ("they said they put all their cards on the table, but when I repeat their overt reproduction instructions, it doesn't reproduce from step 534 to 535" and it only takes 1 person pointing it out and then its cheap to reproduce the discrepancy). It could involve escrow of huge funds, returned only when the model is effectively retired without incident.

This doesn't only protect against Chinese or other foreign influence (let's not ridicule genuine threats like others do on this forum), but also from domestic interference or regulatory capture.

I'm pretty sure the Pentagon wouldn't like Big Tech seizing absolute control of US, neither would a White House regardless of Republican or Democrat.

It should be easy to convince the Pentagon or White House to require all promiscuously shared open weight models to provide this forensic training traceability in standardized machine readable form, regardless of whether its a base model or LoRA fine-tune.

So hobbyists can still train or fine-tune models at home, but when they want to share it OR alternatively when they want to sell or license their work for US workloads, they just have to make sure they enable the build reproducibility in the training harness.

Every time Big Tech refloats the "let's blanket ban all open-weight models", we should reply with this because this sane proposal is actually holding a knife to their financial throat: to fully prove the origin of the final weights, not only does the machine readable archive need to contain snapshots of the process, it also needs to publish the exact training algorithms (a hypothetical mathematically equivalent training speed up trick would not be bit for bit equivalent to the slower computation), the exact corpus dataset, the exact datasets for RLHF, etc...

So basically it would involve forcing model providers to voluntarily publish all their moat, all of it, from the corpus, to custom trade-secret algorithmic optimizations in training, to sensitive RLHF datasets used.

The saner the proposals, the less moat is left untouched, so trying to push for a blanket ban on open-weight models, is a recipe for surfacing such saner models, and thus a very retarded move for big tech to make.

In fact any POTUS, present or future, Republican or Democrat, could probably gain a lot of credibility by enacting such a law.

Just ask the model about Tianamen Square and you know if it’s a Chinese model or not.
> American labs need to release frontier-grade open-weight models under licenses that startups can actually build on.

To be fair, OpenAI has released a couple of (then very good) OSS models. I run the 20B version at home and it is excellent for reviewing text and common tasks like drafting bash scripts. There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too. I wish OpenAI updated these models more frequently though.

Yes, and yes, again the only way to compete is to build the best not hide in a corner and once again the rest of the world will go on in AI without the United States if we flub it. Circling the wagons, isn’t the long range answer.
What’s interesting is no one is talking about political censorship in models and how DeepSeek, Kimi, and the rest have to abide by CCP rules. It’s a big opportunity for China to control information.