I have been only using GLM models since last December and have had the best experience without any drama about tokens and geopolitical restrictions. The quality has been great and I am doing more and more with the latest 5.3 and am really excited that consumer hardware will develop in the next few years where I can run these at home.
Do you use it to write HTML/CSS? Javascript? C++? There's a huge difference in ways people use models and if you are not specific about it then your comment means nothing, unfortunately.
I really like how it doesn't have that Claude talk. It just does the thing without Claude's "load-bearing honesty." It's probably my favorite model to interact with, even if it isn't the best or most reliable.
My second favorite model by now is GLM 5.3 flash which is very capable of day to day task. I use it as the main model and GLM 5.3 for task that is more complex
It's actually slightly more expensive ($0.50 vs $0.48), but there's a temporary 50% discount.
I've seen dozens of conversations about it in last 24 hours, and every major inference provided added in first 24 hours. I think it's gaining plenty of traction.
OpenCode Go is becoming less of a good deal by the month. I pretty much only use it for mimo 2.5 pro now, and everything else is either ollama or openrouter.
OpenCode Go is probably using quantized down DS4Flash. They outsourced to 3th party providers to keep the cost down, and being able to provide that $30 value (instead of the initial $60 > $15).
We saw the same issue with GLM 5.2 when they still published publicly who the providers are on their website. Most ran FP8 but one was doing FP4, so you had this issue where one moment you had the better FP8 and another session you had the FP4 provider.
You can check the internet archive, it was in the FAQ part before they hide/removed it. So if you looked up the providers, and the published quants, yea, ...
Given that a lot of complaints are coming from people that felt OpenCode Go Flash feel like a step down compared to old OpenCode Go/DeepSeek API directly, it smells of a quantized down provider is mixed in.
I've been using it quite a bit too. My main complaint is that it can be really slow sometimes — like, really slow — and the speed feels pretty inconsistent.
Not in my experience. Tasks that would normally cost $0.08 on DSV4-Flash have cost me $0.30+ on GLM-5.3-Flash. These costs are after Deepseek's recent increase. Also GLM-5.3-Flash is so slow compared to DSV4-Flash. I would be fine with GLM-5.3-Flash if it was cheaper and at the same speed as DSV4.
I have seen you advertise your website a few times. I like the idea of not having to trust the router, so I took some time out of my day to critique your website: https://files.catbox.moe/v68cf7.png
My visit to your website went like this:
1. Visit models page
2. Try to find GLM-5.3-Flash (which is among the ~5 models that 90% of people currently care about)
3. Give up scrolling (which would have taken OVER 50 SCROLLS!!!) and use Ctrl + F
4. Try to find input/output/cached price
5. Scroll all the way up to find out which column is what
6. Notice that output price is cut off
7. Notice that the scroll bar is over 100 scrolls further down the page
8. Use Shift + Wheel to scroll horizontally (most visitors probably won't know this trick)
9. Notice that cached price is missing
10. Conclude that this is probably not a serious offering and bounce
There are probably more issues later on, but this is how far I got.
I would suggest you to:
- Deslopify all pages that a user may visit before conversion
- List important models first (see OpenRouter rankings)
- Move the most important information (model name/input/output/cached price) to the left
- Disaggregate the prices per provider (maybe subtables per model? not sure)
- Measure cache hit rate and compute effective price per provider (see OpenRouter)
(- Optional: Fix the broken link on your HN profile page. Currently, the only way to get from this comment to your website is a search engine.)
Have you guys been having a good experience with OpenRouter? I tried it out recently with Claude, and it cached no tokens, charging me $200 for one conversation of 11 messages.
I tried using deepseek v4 flash with OpenRouter. It switches between providers too eagerly which resets the cache. Then, each provider begins to rate limit me for providing so many uncached tokens, so it just keeps on switching providers. I'm paying for every token... why rate limit me? It was unusable compared to just using the official Deepseek provider which has a much better cache rate.
GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better.
Assuming you’re willing to drop a fat stack of cash on the upcoming Mac m5 ultra with 512 gb unified memory, you can even run it locally, quantized to 4 bit. Whether it’s even slightly reasonable, well, my wife would probably skin me alive but maybe yours is more understanding.
One could also run it locally on a used dual xeon (or amd-equivalent) server with 512GB RAM, albeit slower, if you have a useful workflow for it that's like "take this day's efforts and run it through various analysis agents", combined with giving it one-shot tasks/modules to build overnight. You would want a place like a garage or basement to put the server because it'll be loud.
Well if you did get the m5 ultra could you obliterate the guardrails and then your wife can ask it pertinent but unsafe questions about how to punish you. Seems doable.
I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.
Honestly I suspect neither of them will be performing terribly well but with DDR4 3200 RAM I wonder if you'll be counting tokens per second or seconds per token. I mean, you do at least get a lot of memory channels at least, compared to consumer PCs. I am curious to hear what performance you get, I feel there is not enough information out there on what different setups manage to eek out.
What model are you interested in? DS Flash 0731@Q4KXL I'm about 25-30tps. Same as the new Qwen3.8 Flash Next. The new GLM 5.3Q3KXL at 10tps. I've got 2x3090s which I didn't mention in the original message.
I’ll be very curious what you get with DDR4. I also almost went that way. I have an Epyc DDR 5 rig and the best I see is 10 tok/s. Caveat being that’s at Q8 and a 4090 doing pre fill so it could be pushed up.
The surprising thing for me is how much work you will need to cool the banks if you’re near your memory ceiling. My memory starts soft throttling at about 74C (dies may be hotter, that’s the bank temp) and will turn down speed to try to stay below 80.
Happy to send my llama.cpp config settings if you want it.
I find 10 to be very usable. It’s not (that) interactive but it chews through tasks. I let Kimi churn away at 4 overnight and it gives good results that are ready for me in the morning.
I am getting 10t/s on unsloth's Q3kxl with 2x3090s@250w. It's enough for me for now. I will probably upgrade the GPUs down the line. DDR5 would have made the price of the machine double and I just wasn't prepared to pay that much.
Typically computers with these larger memory amounts have fans that scream like a banshee trying to move impossible amounts of air over the memory and CPU. Getting something both cool and quite can be a bit difficult.
I built a dual epyc server with 64 cores and 1 TB of DDR4. Draws around 800W or so under load. I used off the shelf liquid cooling. It is audible but not noisy.
The trick is to turn on the cooler's RGB in your 6000€ server to get a free speed boost. I am not liable for sysadmin's heart attack upon reading this.
Yes, I thought when I was starting that 1u and 2u form factors were to save space. Maybe they are, but they also have the advantage of moving air front to back very effectively through and over the components. Though there still must be some need because I see even those boxes have optional manufacturer built memory shrouds to try to force airflow between the DIMMs.
I had some 120x38mm fans from another server box that I pulled out because they were too loud and I didn't need the static pressure they were giving. They went in here. That 13mm (and the extra 1k rpm) moves so much more air.
Oh please, as though significant numbers of tech workers on this website are not buying LV, Gucci, and yes Hermes for their pampered wives.
My wife can have an opinion on my tech purchases when I get an opinion on her 20,000$ bag addiction.
And the worst part is both of you should unironically accept these purchases, as somehow luxury bags, high end hardware, AND rolex's often out-pace inflation and are objectively good stores of value in a world of rapidly increasing income/wealth inequality.
and even if you were making such a salary, the quesiton of if the investment on hardware to run llm's locally is still a big if, its OK if you buy the HW cause you'll use it and you get the extra capability as a nice extra, but doesnt make sense to spend so much when you could just get 200$ subs with almost infinite SOTA tokens a month etc (if you dont need the local/privacy aspects of it)
* Models are evolving quickly with high worldwide competition
* Hardware is evolving despite RAM shortages
Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what you’d spend otherwise on serverless or dedicated GPU spend…
It is absolutely not worth buying hardware to run models for purely (long term) cost reasons. For open weights models the economies of scale means the cloud beats local significantly and your payback time is like 10 years.
However there are other reasons (e.g. privacy) that might make it worth running locally for some people.
I think the privacy argument that keeps coming up is overrepresented. Certainly ZDR is enough for an absolute majority of use cases? I see so much talk about local inference but I doubt most of it has privacy as a valid argument (not arguing it doesn't exist). It's fun to do things locally though. I've tried it as well but cloud is just faster and cheaper.
These companies have displayed zero respect for everyone's intellectual property getting these models trained.
I think not giving them your complete trust is reasonable! I'm not saying zero trust, and ZDR is fine for most things but I understand the people who don't want to stream their whole codebase out token by token.
Then use other providers hosting open models. Companies and individuals already put their whole code base on the cloud. I'm genuinely interested in privacy-oriented use cases where ZDR is not enough.
I'm not that worried about the codebase itself. I'm worried about the fact coding agents poke around the terminal and system so much that there is almost a certainty that some of your other personal data ends up in the context somewhere which is getting logged in to a training dataset by random hosting providers.
ZDR is built on trust. Given that end-to-end encryption fundamentally doesn't work with LLMs, as they need the content to be unencrypted to operate on it[1], you have no way to prove that once your plaintext data is on somebody else's server they aren't doing whatever the hell they please with it. All you have to rely on is their pinky promise that they won't do anything with it. Trust is a valid option, much of our society runs on trust, but you can eliminate the need for trust whatsoever by running on your own hardware.
[1] Yes, I'm aware of experiments to operate on encrypted prompts, but these are only research attempts, not something that could actually be used with frontier models in production.
Privacy isn’t only, I don’t want anyone to have access to my data. It could also be, I don’t want anyone to know my use case because it’s niche and highly profitable.
You do, there's like 20 providers for any model on openrouter. You can also just spin bedrock or gcp and download the weights for later if you're worried. It's never going to make cost sense when the token rate is so low with how expensive ram is
> I think the biggest reason is to own the stack so your model can't be changed out from under you,
The concern would be future regulations that prohibit you from buying a hosted version of the model. Even that could be bypassed with a VPN to another country but it's more work to go through the payments.
As long as there is demand for a model, it will be hosted by multiple providers.
What if the model is hopelessly obsolete, and thus no demand, but I want that specific model? Owning the weights and hardware is not just solving for one problem. It eliminates all the classes of problems that occur outside of your building, if you have a solar and battery setup.
Also, on a more practical basis, what if the way it's served is bad. Maybe I want my specific KV setup, or ultra low quant for entertaining garbage at 200 tk/s
> What if the model is hopelessly obsolete, and thus no demand, but I want that specific model?
You can still find a lot of old and completely outdated models on OpenRouter. The providers can scale serving of models up and down as demand arrives, so models don't generally disappear. They're just kept in the mix and the clouds will allocate hardware to it if someone is willing to pay.
In the odd case that it disappears completely, buying the hardware 2 years from now is probably going to be a better deal. That wasn't true if you selectively check the time period before hardware got expensive, but as new hardware comes out we're going to start seeing Strix Halo and old Apple hardware hit the market as people upgrade. It's already happening.
There is a certain personality type that cannot tolerate any uncertainty and must lock everything in right now against all future possibilities. If you fit that description then there's nothing anyone can say to discourage you from buying your own hardware, but for everyone else I do not recommend buying hardware to self-host LLMs just to save money. I self-host and run a lot of tokens through my setup (non-coding work) but I'm not really saving money.
> There is a certain personality type that cannot tolerate any uncertainty and must lock everything in right now against all future possibilities.
I thought HN banned personal attacks. I'm in this sentence and I don't like it. /s
I just buy the good apple hardware because it's good, and it also happens to run local models. It's not as good for the dollar, don't get me wrong, but I'm not going to develop iOS without a mac, that's even more questionable than buying a strix or whatever.
I live in a place where using VPN is illegal and akin to "terrorism" because why would you want to hide what you are doing. Only bad guys hide. So if you use VPN, you are a bad guy.
“Out of the 15 individuals identified, five were minors who were counselled and advised in the presence of their guardians, with emphasis on awareness, lawful digital conduct, and the consequences of violating lawful orders,” he added.
I'm actively uninspired to write high quality code when using Anthropic/OpenAI models given the high chance I'm a customer as well as used as dataset generation tool for them.
But currently cloud does beat costs of hardware ownership, particularly with ridiculously high RAM/GPU/SSD costs....again due to these same companies.
I think that's overly pessimistic. Here's [1] a video of somebody running it on a ~$6000 rig and getting around 14T/s for complex prompts (about double that for simpler prompts). Payback time is going to depend on your electric cost/consumption. In most domains cloud providers end up charging a significant premium rather than a offering a scale enabled discount, relative to local at retail costs. That will almost certainly end up being the case with LLMs as well, if it isn't already.
Furthermore we continue to follow the path that image gen neural networks took. In that domain hardware requirements reached a peak and then started sharply declining to where we are today where a plain old video card can rapidly generate images that took a supercomputer not that long ago. So it's reasonable to assume that performance of such a system could potentially even increase over time.
With roughly 2.7 million seconds per month, times 14 tokens per second, you are getting 38.5 million tokens a month at most.
That’s less than 164USD worth of GLM5.3 tokens on the inference market. So that 6000 USD rig will take 3 years to break even - and only if it runs continuously.
And this is being generous, as it’s not even taking quantisation into account.
I think if you steel-man what I'm saying, what you're saying falls apart. 14 tokens per second was rare. It only dropped that low in one scenario where he had it single shot an entire game (flappy bird clone) from scratch, with different assets, all self created, and so on. It ended up resulting in the LLM doing stuff like plotting out a some odd 100 item long to-do list, requerying it repeatedly, and so on. And it succeeded.
Also as the video mentions, the guy wasn't very familiar with what he was doing, and so there are almost certainly various optimizations on the config side he could work out, especially as he was using a 5 GPU system, which default configs are probably not well optimized for.
But I think we've rapidly moving along the same path as image gen stuff. Local generation has gone from purely theoretic, to requiring supercomputers to run relatively incapable models, to where we are today - where with a fairly basic high end setup, he's comfortably running a frontier level model. There's definitely an argument for going local that's only growing stronger by the day.
I agree it’s probably not representative token speed. But I do believe the overall observation holds: The value of local inference is bound by the wall clock.
I agree that there are many other reasons than cost alone.
I think the “killer app” is doing inference without sending the data to China or the US. At home it’s overkill but imagine you are an EU consultancy with a lot of client data to work on, or a company/institution with a lot of sensitive data, buying the hardware to make sure the data stays private is a big selling point.
Some of that is that EU providers need to up their game here.
Needing an EU native option is really the one and only reasonably objection I've heard against using LLMs from the cloud, the rest is tin-foil hat level unless you're actually intending to meddle with the inference or fine tuning or something beyond just querying.
I mean, I think it depends. At home 3 of us we use AI for multiple reasons, from coding apps to asking general questions, and if we would have to pay equivalent subscriptions that would be ~1k a year on AI + submitting all your data to external services. I payed around ~8k on 2 DGX Sparks that, at the moment, serves perfectly fine as a ChatGPT/Claude replacement at home (DS4 Flash peaking at ~170 tokens per sec with 6 concurrent sequences), and even once the technology is obsolete for inference in a few years, I will still have 2 pretty powerful machines for whatever I need + some pretty fast NVME Storage. I don't think its a terribly bad idea.
> It is absolutely not worth buying hardware to run models for purely (long term) cost reasons
This is especially true when it's trivial to have the LLM itself write you a script/tool that can rent a GPU node for you (via API calls to providers) and then download and set up an open weight model for you.
So far I don’t regret buying an M1 Max device with 32Gb of RAM. The models available for it keep getting better (running just about okay for interactive use) and 400 GB/s of bandwidth is still considered a lot.
The models are currently improving much faster than the hardware and this doesn’t seem to have plateaued yet.
Try 3.8 27B in MTPLX; I get about 30 tok/s with the same hardware as you. (Although it does use around 90-95W of power, compared to the ~60W that 3.6 35B-A3B uses to generate 55 tok/s. That’s about 3 J/tok instead of 1.)
NGL: I don’t really have a good way to find out right now. It also doesn’t matter that much because the way the models use the tokes varies a lot. Qwen 3.8 is known for overthinking while Muse Glimmer may be a little slower per token, but it uses them very efficiently, caveman style.
Generation speed isn’t the bottleneck anyway, at least on pre M4/M5 devices (the newer chips got significant processing acceleration). It’s prompt processing time. OpenCode’s system prompt can take up to 3 minutes to process, which is why good prompt caching is essential.
For that I use omlx, which can persist the KV cache to disk, chunked so you can reuse parts. This helps with the usability a lot, when an agentic session is warm it runs pretty smoothly. New requests can take a couple seconds (sometimes many, which must be fixable somehow).
So: It’s not fast, but I also don’t find it awfully slow. My use is typically semi-interactive, for fully interactive use you have to wait a bit, but it’s possible. I personally am still regularly amazed that something even close to this is possible on completely local hardware.
That's basically the question I'm trying to answer.
If you're paying Anthropic or OpenAI to use their models, harness, governance, etc., I could see the local inference potentially coming out ahead. They're already starting to ratchet down what your money gets you on their platforms, and that can be expected to continue as the leaders of those companies continue to seek the road to the El Dorado that is being a trillionaire.*
If you're looking to get into the guts of AI development instead of having it handed to you by a provider, that's where it gets murky. I'm wanting to write some sort of agent that does things and get into making outputs consistent in the like, and I'm not sure whether to host something on GCP or buy an M5 Mac.
*Note: El Dorado is a mythical city and many people died trying to find it.
Part of it is knowing that whatever sort of enshittification the cloud providers do, my local programming environment won’t ever be less effective than it is today locally. It’s the same reason my entire development stack from editor to compiler is open source. I don’t need to modify it today, but I always must retain the option to do so later.
There are several things I do in my life that only pay off in the event of a big disaster, like an extended internet outage, civil unrest, supply chain disruption, war, etc.
I like to be able to do the things I do even if offline for weeks.
I spent a lot of money for flash in the big iPad so I can keep all of offline wikipedia and OSM in it, for example. It’s sort of like being a digital prepper. (Being a prepper is a spectrum, from anyone who keeps food in their pantry to people building bunkers under their house - how much you invest is a personal prudence and threat modeling decision.)
> Part of it is knowing that whatever sort of enshittification the cloud providers do, my local programming environment won’t ever be less effective than it is today locally.
Is that true though? Many of the core LLMs need to be retrained as languages evolve to incorporate changes (language specifics, compilers, tooling, etc.). To some degree this can be handled via context injection in a variety do forms (agents looking up documentation and so on) but inevitably it’s not stationary in time, just as your OSS stack (probably) isn’t (depending on the languages, technologies, and use cases).
So your hardware is to some degree dependent on the good merit of groups like Z or Alibaba or whomever pushing out updated open weight models that dumped loads of capital into to train. You can keep using the existing models but at some point I suspect they’ll start to have more friction due to dated specs in language and so on. Again there are tuning and ways of layering this information on, and in theory you can even do some training on your own but I don’t think it’s as stationary as being portrayed here.
Those updated open weight models may not always be there (updated on new data). The usability of them is probably fairly long to be fair, but I suspect you’re going to see explosion in everything from libraries to languages etc due to LLMs so even the rate of change across your OSS stack may cause these models to be dated quite quickly, at least in the core model which will require layering fixes.
To be clear I’m on the fence thinking about much of the same issues and as close as I am to pulling the trigger, I keep thinking of very valid counter arguments as to why it’s me just wanting this thing I own. Which may be enough.
With every newly released open weight model, the clock on the issues you describe is reset. I can see a marketplace arising for paid updates to common lines of open weight models, which will incentivize those with the hardware to train to fix the problem for those who only have the hardware for inference.
I would say when this comes to pass, we are already 5 years along?
> Part of it is knowing that whatever sort of enshittification the cloud providers do, my local programming environment won’t ever be less effective than it is today locally.
I think this is quite understated. It basically is freedom from a growingly antagonistic relationship between you and some remotely hosted API managed by faceless corporates at the whims of their board, shareholders and governments.. It really is such a mental burden to need to constantly manage this relationship (watermarks, silent downgrades, random false refusals, downtimes, model sunsets, changing ToS's, fucking ads). These companies will need to squeeze you for every cent that they can before open-weight models are simply good enough for the valuable tasks we can throw at them.
To have your own hardware is to no longer have this mental burden.
With competition we kind of have guarantee up to what providers can do, they don't have that much control, the most radical thing they can do is to go bankrupt.
I have a Strix Halo and dual 32GB GPUs in my desktop, that sit idle right now, because the electricity to run them and to cool them in 110F weather Texas is currently experiencing pretty much nulls any savings I might see over getting better models from cloud providers. While I mostly use Claude or Codex with subscriptions for agentic work, for API use DeepSeek has usually been my go to, but now I guess it's GLM 5.3 or the Flash version. And, for security work that Anthropic or OpenAI models are likely to refuse, I've been using Kimi K3 (also via subscription, though their subscription is extremely stingy), but I guess GLM is now the one for that, too.
Anyway, yeah, even at the prices I spent on my local AI stuff (I bought before RAMpocalypse really kicked into gear, so I bought old server GPUs for about $350 each and the Strix Halo for a little over $2k) it was never going to pay for itself; I just like to tinker. But, I can't imagine spending today's prices for hardware for local AI.
When the memory shortage ends, I'll be down to the Apple Store (or, more likely, clicking refresh on the Apple outlet every few days). But, until then, there continues to be a glut of cheap and free models in the cloud that are better than anything I can run locally and they're faster, too.
If they "should" cost 4k in the sense of marginal cost, then you will be spending more running the same at home, because your home hardware will always be less efficient.
There is a big difference in the cost of a 5-nines up time system in a heavily space constrained environment compared to a home hobby white box used for some coding. The GPUs alone cost 10x for the data center versions compared to the gaming versions even with similar specs.
The cost of online services is also largely a result of the cost of training (though hard to say exactly what that number is). Assuming you are using open weight models at home, you aren't paying for the training - someone else is.
> The cost of online services is also largely a result of the cost of training
OpenRouter prices are somewhat simmilar to Antrhopic/OpenAI API prices. So I conclude that the hardware plus operating margin alone can genuinely produce prices way above what you'd pay if you had a subscription.
Of course the primary unkown factor is average token use per subscription. Without that it's all wild speculation.
Yeah, I guess, but it feels like there isn't really an opportunity for anyone to do that, given how competitive the market is. If Anthropic decides to demand API rates for everything (which would make my $100/month turn into a few thousand, I guess), I won't be seriously inconvenienced by switching to GPT. And, if both of the major American providers do a pricing collusion and GPT also becomes thousands of dollars a month to use, I can choose between Kimi K3 and GLM and so on. I'd rather use Opus 5 and Fable, but I'm not going to be seriously put out if I can't. We've got three or four open models to choose from that are as good as or better than Opus 4.8, which is Good Enough, and the competition isn't slowing down. We're seeing more new competitive models more frequently than even three months ago.
So, even though there are more models to run locally that can be useful for the stuff I do, it makes less sense now to do so than it did when I got it. There are more extremely cheap options, now, and it seems likely to continue to get cheaper and better and faster, while my local hardware will always be slow and hot and only gets better via software (which has roughly doubled the speed it can run stuff since I got it, but it seems like there's less room for improvement in software now, and even at twice the speed, it still kinda sucks to use local models interactively especially on the Strix Halo).
I wish Texas would write up a regulation allowing 'balcony solar' as I could easily generate 1000-2000w of solar in my small back yard to take a bite out the sizeable cooling bill I have.
Seems like it's easier to ask forgiveness than permission. And, I wouldn't bet on this legislature ever doing anything that would disempower fossil energy or reduce their profits, even a little bit.
There are good reasons you aren't allowed to plug random power generators into the grid. You might be allowed to have your own ones not connected to the grid. Remember that graphics cards run off poorly regulated 12V, although you'd want to regulate it anyway because they're expensive to replace if I'm wrong.
As I understand it. They banned solar inverters as well as severely tariffing the panels. In some places it's also illegal not to get your electricity from the grid.
Nope, because of laws written by energy company lobbyists for the "safety" of consumers, so if you're mixing grid and solar on the same site (even if completely isolated circuits) you have to get all kinds of certifications and permits for inverters and breakers installed by certified technicians that make it cost prohibitive.
I posted more on the other branch of this thread. But cost prohibitive is about all I can find about this and not illegal. Which for some people might be the same result of course.
Only reason to spend a bunch of money on hardware to run LLMs locally is if it's a hobby to you to an extent that even renting the GPUs temporarily won't satisfy you.
Or if you need stuff that APIs don't / can't provide. Or for future proofing your workflows. Running things locally gets you "the same thing" in perpetuity, while APIs might change, models can be deprecated and features can be removed.
Cybersec is also hit and miss, depending on what provider you choose, verification systems and all that jazz. Also, running locally allows you 100% data privacy, in any situation and for whatever usecase you might have. ~100k for hardware for a small team of devs to code locally is not that expensive in the grand scheme of things.
Lastly, local models allow for training / finetuning on your own data and processes. $/tok is not everything for everyone. Sometimes you can take a hit on value / speed if you get something else that matters for you.
Or Qwen 3.8 27B, AA score 52 (which is utterly insane given the size of this model); I have been testing Qwen 3.8 27B since a week now, as an intensive GLM-5.2 and Opus 5 user - I can say that I just can't believe my eyes i.r.t. to how good this model is.
There's certainly a mental difference between a computer you can use as much as you want for a fixed upfront cost vs a rented server you are being billed by the hour for.
But the cost difference between api and self hosted is so incredibly large now it's almost impossible to ignore the fact self hosting is a terrible deal. I'm waiting things out until the dust settles on what the required specs are and consumer hardware gets cheaper/more capable.
”less touchy” is an understatement, it happily complies with running exploits, reverse engineering and decompiling. Asking Claude to do that will give you an error and make you feel like you’re about to get banned.
I was going to try out pi and I set it up with some of the recent big open weights models, but then I realized that if I just use pi the way I use claude code, it doesn't have an auto mode so it's basically just --dangerously-skip-permissions and you're supposed to sandbox everything yourself. What's your sandboxing setup like?
how feasible is it build a SOTA specialized model for some use case e.g. deal sourcing by using this as pre-trained model or a LORA or similar pattern on top? Gonna shoot my shot at a billion dollar business
Not feasible. We’ve seen again and again that generalized models—somewhat surprisingly—dunk on specialized ones in almost all cases.
The first case of this I remember was Bloomberg thinking that their extensive private data about markets would give a home-trained model better performance in finance tasks. The next version of ChatGPT beat them.
With how often new versions of frontier models are released, you likely won’t finish your work before it’s obsolete. The labs have whole teams dedicated to specific getting training data for specific industries (finance is one), and more powerful generalized models make better decisions even without specialized knowledge.
Your best bet is to get really good at training for something and then sell your company to one of the frontier labs for their post-training efforts.
yes would also be interested in that - using knowledge distillation and other special knowledge sources to post-train on top of an open model like GLM-5.3. I was wondering that when Elon Musk tweeted "Specialist AI’s (single language, single area of knowledge) are another 100X" (src: https://x.com/elonmusk/status/2089968914596045178) - maybe he knows something regarding specialist model training the general public does not know?
How much usage do you find you get on these kinda models (I know the pricing changes a bit) compared to a $20 sub say for Google AI Pro in anti gravity?
I hate how difficult it is to compare prices when looking at subscriptions.
Would $20 in open router, using models like GLM get me more or less?
I think it'd get you less than a $20 sub to any of the big three. I've used it on OpenRouter and found it kind of expensive for the results, but that might change now that it's open weight and other providers can host it/compete with Z.ai. For the work I did with it, I would've rather used DeepSeek V4 Flash just because it's more economical and still gives good results IMO.
Z.ai does have their own subscription, but I haven't used it because their privacy policy was pretty buns last time I checked.
Yeah I used Ox alpha earlier this week when it was free and deepseek flash when it was free on opencode. Both were great. Z.ai’s subscription doesn’t look very good versus the others.
I burn through my current Google AI pro sub for the week in about 2 and a half days so wanted something extra to add to it, but don’t want to buy any expensive ultra plan. Flash models have got me about 98% of what I need, but quotas are still a bit low.
"The Company do not store any of the content the Customer or its End Users provide or generate while using our Services. This includes any texts, or other data you input. This information is processed in real-time to provide the Customer and End Users with the API Service and is not saved on our servers."
I made a little project to calculate that. It gets the prices for API access and subscriptions, and calculates an average cost per token, for a given rate limit. So the cost is cheaper if you get a higher rate limit (cuz you get more tokens per dollar): https://codeberg.org/mutablecc/calculate-ai-cost
tl;dr API cost (openrouter) is always more expensive than a subscription (for the same given model). you should always use a subscription first and only go to API pricing if you run out of your subscription.
In terms of which subscription is best, different ones provide different models, different amounts of tokens, different rate limits. So it depends on what model you want and how much you need to use it. The frontier ones are always more expensive than open weight ones, but a few subscriptions are starting to include frontier models like GPT-5.6 Luna (which is a great deal but not necessarily the best price-per-performance).
I'd like to ask Sam Altman if he still thinks that it's too dangerous to publish GPT-3. I mean, no one would use it, but what is his reasoning for not publishing it now, in 2026?
Maybe I'm reading too much between the lines, but I suspect the reason is to rub his nose in the duplicity or naivety depending on how generous you're feeling. Publishing the model would be a confession that he was wrong.
AI policy is being shaped somewhat by the things Sam and Dario say. So even if you're not feeling vindictive, it's probably good to keep a track record of the previous things they have said as a Bayesian prior. People who don't know better listen to these people, and maybe they shouldn't.
If you want to dunk on sam for "it's too dangerous to publish GPT-3", you hardly need the release of gpt-3 to prove your point. All the other open model releases already provide pretty good evidence. Moreover the fact that the model wasn't release hardly points to the fact that he wanted to save face or whatever. Most AI labs don't release their old proprietary models, so the fact that gpt-3 wasn't released tells us very little.
I can't read Sam's mind, so I don't know the point. I was just trying to answer the other dude's (apparently insincere) question.
> I think the release of kimi k3 is definitely arguably dangerous
I understand the argument, but I don't see how it matters. Get Sam and Dario to lobby to make it illegal to run in the US? Ok, then the bad guys who follow the law will run Kimi K3 from outside the US.
Maybe you'll petition the Chinese to lock down access to their models?!? Even if they did, it means the Chinese government can still use them.
The only solution I see for the cyber security side of things is to fix the bugs or migrate to more secure systems and approaches. We all had to do something similar in the mid 90s when the world came online and TELNET and FTP weren't good enough any more.
I think it would be an important historical document as well. We are potentially looking at the dawn of AGI and one of the most important models ever created. Each model is also a kind of ultimate time capsule, containing a snapshot of the entire human collective mind. If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.
> If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.
The weights arent the truth tho, maybe a timecapsule-vhs but i wouldnt trust llm weights more than more hardcore deterministic media that might get preserved to infer facts from an era.
The companies doing the training are becoming the "winners" that are "rewriting history" as they train their models.
I agree that they are historically important, but if you want to query old thinking in 2070, you'd probably be better off having a modern model analyze archive.org. If that ever goes down, we're sunk.
> containing a snapshot of the entire human collective mind.
I say this with kindness: Anyone who believes this absolutely needs to turn off their computer for the week, go outside, travel a bit, and experience reality with other humans outside their regular bubble.
The “entire human collective mind” is not digital. It’s not on the internet. These models could’ve syphoned literally every piece of digital media in existence and still wouldn’t have it. People don’t exist inside computers, and it is naive to believe the sum of what’s online makes the sum of the human experience. It doesn’t.
There are risks associated with releasing historical proprietary models that were not designed for open release:
- It is trivial to extract samples of the training data that was used, which can bolster existing lawsuits/foster new ones.
- Older models are not as safety-hardened, so it is easier to coax unsafe behaviour out of them, which is a PR risk.
- It may be possible to divulge proprietary secrets from the model (e.g. architectural details that may still be relevant still).
For these reasons, and more, it's unlikely that GPT-3/similar models will be released until these concerns are no longer relevant (e.g. when they become a purely historic concern, similar to the open-sourcing of other proprietary software from decades ago).
> It is trivial to extract samples of the training data that was used, which can bolster existing lawsuits/foster new ones.
At least to this point, the legal teams could get the model via discovery. IDK that the lawfirms realize that they could get experts (or that they'd have contacts that could) to search the model weights.
There’s not such a straightforward relationship between safety and model sis.
According to the book The Thinking Game, lower quality models at that time were considered less safe, because they could be easily tricked into doing harmful stuff. In the book, Dario (of Anthropic) was the head of safety at openAI and was responsible for pushing for 10x scaling in training to make the models safer.
It does make sense, a smart model is going to be way harder to jailbreak into telling me how to synthesize anthrax (or something).
Models are quite safe when they're useless, actually.
In the times of GPT-3 I'd scoff at the idea of an LLM doing any hacking; today, I'm running several AIs on my code before publishing, and they are finding (and demonstrating!) RCEs on my localhost server.
For example, one found a missing check in a third party JWT library which allowed full account takeover, which I'd have never even looked at.
Hence I don't believe a single word coming out of these people's mouths. Their "beliefs" are just marketing.
From today's perspective, it sure seems like it, probably because increased capabilities have generated a new kind of danger. Back then, they were worried about stuff like the model telling me dangerous knowledge.
I certainly think the labs have muddied the waters using safety for marketing, but that doesn't mean less capable models weren't more dangerous at one point.
Extremely weak justification on their part, bordering on trolling. It's just marketing.
Easy access to malicious information hasn't proven to be the disaster these anti-prophets proclaim. For the last ~3 decades of widespread internet and search engines, you could trivially find all sorts of information (drug synthesis, explosives, etc.), and we're just fine.
(Corollary: easy access to good information did not make non-intellectual non-curious people smarter. Easy access to information does not have the consequences people immediately assume.)
There’s this strain of thinking that’s complete alien to me, I can’t interpret what’s being said and it sounds incredibly aggressive. ex. our lead post asking if Sam Altman still thinks GPT-3 is unsafe to release, and I don’t understand what beliefs you don’t believe and who they is and why they’re just choosing to lie for marketing.
My guess is I’m unintentionally refusing implicit signal that you think it’s safe to release all models openly, because you have observed models finding a vulnerability in a JWT library. But that sounds like a straw man instead of a steel man. Idk. :/
In some interviews, OAI mentioned that they didn't think that GPT-3.5 would be a success. They thought it would be a cool toy and they decided to launch it to see how users react. That means that they didn't think GPT-3.5 was intelligent enough. But somehow once GPT-3.5 became a huge hit, people conveniently ignored the anecdote, and started to believe that AGI had been eminent.
I previously posted that DS4Flash was _good_ but not _great_ on two DGX Sparks, but I have to say that GLM-5.3 is pretty amazing. It's been able to tackle all the random hard problems I've thrown at it and it has the intuition that DS4Flash seems to lack.
We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.
It's early days, but GLM 5.3 Flash is the first local model that feels good enough to me to be a "main" model without debating whether each problem needs to be sent to a stronger model. DS4 Flash is good enough at implementing given a plan, but I wasn't always a fan of what it came up with when asked to plan something.
The good news is it can only get better from here.
> We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.
I'm wondering of you could clarify your thoughts on this. I've had a hard time evaluating what Fable-class actually is capable of that sets them (or really it) apart from other models in a very significant way.
Have you tried it out? I doubt there is a general accepted definition. For me it is just more capable of deep reasoning/handling complexity. Still can mess up, still does not count as strong AI - but a level above Opus and co.
Is it possible to fine tune this model and unlock / extend its cybersecurity capabilities? I'm scared that maybe we are not ready for an open-weight model with high cybersecurity skills.
You can fine tune a model from a year ago to get extended cyber capabilities. Fine-tunes dramatically increase capability in specific use cases and don't require a lot of investment. Attackers have been doing this for a while now, they aren't waiting for someone else to make them a security model.
I get how you feel, but it's too late to be concerned. The cat's out of the bag. It's like being scared of moving from the bronze age to the iron age... when everybody already knows how to make iron, and the raw materials are everywhere. People are already making iron spears. We need to make iron shields.
We need open-weight models that are good at finding security holes so we can apply them to all of our software by default, and close every possible security bug, before the attackers find them. Every piece of software in the world should be held for release until it's scanned by a high-powered security model.
This is the same debate we had in the 1990's when strong encryption was considered a munition and not allowed to be exported. This just made the world less secure. And it was pointless anyway, because you can't really stop it being developed and shared. Eventually good sense prevailed and now we all have strong encryption. The same thing applies to security bugs.
Looking at HF, it looks like the unquantized version is half the size of glm-5.2 756 GB vs 1.51 TB. I wonder how they were able to optimize it this well
I doubt it - AWS hasn't added any non-western models since GLM 5 and MiniMax M2.5 in February, afaik. Might be a deal with OpenAI (GPT 5.4 was the first to be available via Bedrock, in April) or might just be that there isn't a lot of demand due to corporate skittishness around models trained in China.
Yes I really don’t a reason to not switch all my team to use GLM 5.3 for planning, and Flash for implementation. This really does seem apocalyptic for Anthropic and OAI if more of the industry switches.
I’ve been trying out GLM 5.3 Flash the past < 30 hours, and although I’m only running at Q3, it feels different than most 100B to 200B models I’ve run before... More measured, thorough in thinking, and has so far passed all my private tests (and at Q3).
I find GLM 5.3 Flash more interesting than 5.3. The fact 5.3 does not have vision is kind of a deal breaker. Also 5.3 Flash seems to be better at making pretty UIs.
Its more about "man, imagine if we didnt have competition", which is what the us is trying to eliminate by soon labeling all openweight llms as dangerous and illegal.
What's very promising here is the number of tokens-vs-accuracy ratio. I am assuming their "output tokens" means tokens generated as part of thinking and any tool calls (what are referred to as "input tokens" from billing PoV by service providers). The Chinese models like Qwen3.8 and GLM 5.2 are insanely overthinking in our workloads (which are highly complex data analysis tasks). It's a factor of 3-4x over Opus and GPT models. Even with cheaper prices per 1M tokens, the cost ends up being higher, in some cases 2x. So this is very promising from GLM 5.3. Looking forward to trying it.
Interesting model. It's nice to see open-weight models catching up with the top proprietary ones. The coding benchmark results are pretty impressive. Overall, it's great to see progress moving forward. Curious to see how it performs in real-world usage.
251 comments
[ 0.19 ms ] story [ 11.6 ms ] threadI've seen dozens of conversations about it in last 24 hours, and every major inference provided added in first 24 hours. I think it's gaining plenty of traction.
We saw the same issue with GLM 5.2 when they still published publicly who the providers are on their website. Most ran FP8 but one was doing FP4, so you had this issue where one moment you had the better FP8 and another session you had the FP4 provider.
You can check the internet archive, it was in the FAQ part before they hide/removed it. So if you looked up the providers, and the published quants, yea, ...
Given that a lot of complaints are coming from people that felt OpenCode Go Flash feel like a step down compared to old OpenCode Go/DeepSeek API directly, it smells of a quantized down provider is mixed in.
z-ai/glm-5.3: also Z.ai, Novita, Atlas Cloud, IO.NET
My visit to your website went like this:
1. Visit models page
2. Try to find GLM-5.3-Flash (which is among the ~5 models that 90% of people currently care about)
3. Give up scrolling (which would have taken OVER 50 SCROLLS!!!) and use Ctrl + F
4. Try to find input/output/cached price
5. Scroll all the way up to find out which column is what
6. Notice that output price is cut off
7. Notice that the scroll bar is over 100 scrolls further down the page
8. Use Shift + Wheel to scroll horizontally (most visitors probably won't know this trick)
9. Notice that cached price is missing
10. Conclude that this is probably not a serious offering and bounce
There are probably more issues later on, but this is how far I got.
I would suggest you to:
- Deslopify all pages that a user may visit before conversion
- List important models first (see OpenRouter rankings)
- Move the most important information (model name/input/output/cached price) to the left
- Disaggregate the prices per provider (maybe subtables per model? not sure)
- Measure cache hit rate and compute effective price per provider (see OpenRouter)
(- Optional: Fix the broken link on your HN profile page. Currently, the only way to get from this comment to your website is a search engine.)
https://trustedrouter.com/models
we are doing billions of tokens a day and thousands of users
https://openrouter.ai/docs/guides/routing/provider-selection
Assuming you’re willing to drop a fat stack of cash on the upcoming Mac m5 ultra with 512 gb unified memory, you can even run it locally, quantized to 4 bit. Whether it’s even slightly reasonable, well, my wife would probably skin me alive but maybe yours is more understanding.
Does inference make full use of the memory bandwidth in a NUMA system?
The surprising thing for me is how much work you will need to cool the banks if you’re near your memory ceiling. My memory starts soft throttling at about 74C (dies may be hotter, that’s the bank temp) and will turn down speed to try to stay below 80.
Happy to send my llama.cpp config settings if you want it.
Temp wise, no throttling, surprisingly cool.
The trick is to turn on the cooler's RGB in your 6000€ server to get a free speed boost. I am not liable for sysadmin's heart attack upon reading this.
I had some 120x38mm fans from another server box that I pulled out because they were too loud and I didn't need the static pressure they were giving. They went in here. That 13mm (and the extra 1k rpm) moves so much more air.
Oh please, as though significant numbers of tech workers on this website are not buying LV, Gucci, and yes Hermes for their pampered wives.
My wife can have an opinion on my tech purchases when I get an opinion on her 20,000$ bag addiction.
And the worst part is both of you should unironically accept these purchases, as somehow luxury bags, high end hardware, AND rolex's often out-pace inflation and are objectively good stores of value in a world of rapidly increasing income/wealth inequality.
* LLM usage is new for the world
* Models are evolving quickly with high worldwide competition
* Hardware is evolving despite RAM shortages
Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what you’d spend otherwise on serverless or dedicated GPU spend…
However there are other reasons (e.g. privacy) that might make it worth running locally for some people.
I think not giving them your complete trust is reasonable! I'm not saying zero trust, and ZDR is fine for most things but I understand the people who don't want to stream their whole codebase out token by token.
[1] Yes, I'm aware of experiments to operate on encrypted prompts, but these are only research attempts, not something that could actually be used with frontier models in production.
The concern would be future regulations that prohibit you from buying a hosted version of the model. Even that could be bypassed with a VPN to another country but it's more work to go through the payments.
As long as there is demand for a model, it will be hosted by multiple providers.
Also, on a more practical basis, what if the way it's served is bad. Maybe I want my specific KV setup, or ultra low quant for entertaining garbage at 200 tk/s
You can still find a lot of old and completely outdated models on OpenRouter. The providers can scale serving of models up and down as demand arrives, so models don't generally disappear. They're just kept in the mix and the clouds will allocate hardware to it if someone is willing to pay.
In the odd case that it disappears completely, buying the hardware 2 years from now is probably going to be a better deal. That wasn't true if you selectively check the time period before hardware got expensive, but as new hardware comes out we're going to start seeing Strix Halo and old Apple hardware hit the market as people upgrade. It's already happening.
There is a certain personality type that cannot tolerate any uncertainty and must lock everything in right now against all future possibilities. If you fit that description then there's nothing anyone can say to discourage you from buying your own hardware, but for everyone else I do not recommend buying hardware to self-host LLMs just to save money. I self-host and run a lot of tokens through my setup (non-coding work) but I'm not really saving money.
I thought HN banned personal attacks. I'm in this sentence and I don't like it. /s
I just buy the good apple hardware because it's good, and it also happens to run local models. It's not as good for the dollar, don't get me wrong, but I'm not going to develop iOS without a mac, that's even more questionable than buying a strix or whatever.
https://srinagar.nic.in/notice/immediate-suspension-of-virtu...
Phones are randomly searched on the streets and if VPN is found, arrested
https://www.medianama.com/2026/01/223-jammu-kashmir-vpn-ban-...
https://timesofindia.indiatimes.com/india/after-vpn-ban-in-k...
“Out of the 15 individuals identified, five were minors who were counselled and advised in the presence of their guardians, with emphasis on awareness, lawful digital conduct, and the consequences of violating lawful orders,” he added.
But currently cloud does beat costs of hardware ownership, particularly with ridiculously high RAM/GPU/SSD costs....again due to these same companies.
Furthermore we continue to follow the path that image gen neural networks took. In that domain hardware requirements reached a peak and then started sharply declining to where we are today where a plain old video card can rapidly generate images that took a supercomputer not that long ago. So it's reasonable to assume that performance of such a system could potentially even increase over time.
[1] - https://www.youtube.com/watch?v=ZWS2JVN2iBI
That’s less than 164USD worth of GLM5.3 tokens on the inference market. So that 6000 USD rig will take 3 years to break even - and only if it runs continuously. And this is being generous, as it’s not even taking quantisation into account.
Also as the video mentions, the guy wasn't very familiar with what he was doing, and so there are almost certainly various optimizations on the config side he could work out, especially as he was using a 5 GPU system, which default configs are probably not well optimized for.
But I think we've rapidly moving along the same path as image gen stuff. Local generation has gone from purely theoretic, to requiring supercomputers to run relatively incapable models, to where we are today - where with a fairly basic high end setup, he's comfortably running a frontier level model. There's definitely an argument for going local that's only growing stronger by the day.
I agree that there are many other reasons than cost alone.
Needing an EU native option is really the one and only reasonably objection I've heard against using LLMs from the cloud, the rest is tin-foil hat level unless you're actually intending to meddle with the inference or fine tuning or something beyond just querying.
I can cherry pick stats too.
The other day I heard mention of someone paying $200/mo for Claude Code.
At those rates my local LM setup pays for itself in a single year.
This is especially true when it's trivial to have the LLM itself write you a script/tool that can rent a GPU node for you (via API calls to providers) and then download and set up an open weight model for you.
And not "Hey, these companies have a history of giving you something nice now, and rugpulling you either in quality or price later."
Your "absolutely" seems silly.
The models are currently improving much faster than the hardware and this doesn’t seem to have plateaued yet.
It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B.
I’ve also used Qwen 3.8 27B but I get 10t/s on it.
It’s useable in some use cases, but I rely mostly on my $20 Claude subscription.
3.6 35b a3b, I’m getting upwards of 100
Generation speed isn’t the bottleneck anyway, at least on pre M4/M5 devices (the newer chips got significant processing acceleration). It’s prompt processing time. OpenCode’s system prompt can take up to 3 minutes to process, which is why good prompt caching is essential.
For that I use omlx, which can persist the KV cache to disk, chunked so you can reuse parts. This helps with the usability a lot, when an agentic session is warm it runs pretty smoothly. New requests can take a couple seconds (sometimes many, which must be fixable somehow).
So: It’s not fast, but I also don’t find it awfully slow. My use is typically semi-interactive, for fully interactive use you have to wait a bit, but it’s possible. I personally am still regularly amazed that something even close to this is possible on completely local hardware.
If you're paying Anthropic or OpenAI to use their models, harness, governance, etc., I could see the local inference potentially coming out ahead. They're already starting to ratchet down what your money gets you on their platforms, and that can be expected to continue as the leaders of those companies continue to seek the road to the El Dorado that is being a trillionaire.*
If you're looking to get into the guts of AI development instead of having it handed to you by a provider, that's where it gets murky. I'm wanting to write some sort of agent that does things and get into making outputs consistent in the like, and I'm not sure whether to host something on GCP or buy an M5 Mac.
*Note: El Dorado is a mythical city and many people died trying to find it.
My rented H200 server comes at just over £120 a day, £3.2k~ a month. Minimum two year contract.
I don't want to, but I am having to say bye to my 2x1u.
There are several things I do in my life that only pay off in the event of a big disaster, like an extended internet outage, civil unrest, supply chain disruption, war, etc.
I like to be able to do the things I do even if offline for weeks.
I spent a lot of money for flash in the big iPad so I can keep all of offline wikipedia and OSM in it, for example. It’s sort of like being a digital prepper. (Being a prepper is a spectrum, from anyone who keeps food in their pantry to people building bunkers under their house - how much you invest is a personal prudence and threat modeling decision.)
Also, privacy.
Is that true though? Many of the core LLMs need to be retrained as languages evolve to incorporate changes (language specifics, compilers, tooling, etc.). To some degree this can be handled via context injection in a variety do forms (agents looking up documentation and so on) but inevitably it’s not stationary in time, just as your OSS stack (probably) isn’t (depending on the languages, technologies, and use cases).
So your hardware is to some degree dependent on the good merit of groups like Z or Alibaba or whomever pushing out updated open weight models that dumped loads of capital into to train. You can keep using the existing models but at some point I suspect they’ll start to have more friction due to dated specs in language and so on. Again there are tuning and ways of layering this information on, and in theory you can even do some training on your own but I don’t think it’s as stationary as being portrayed here.
Those updated open weight models may not always be there (updated on new data). The usability of them is probably fairly long to be fair, but I suspect you’re going to see explosion in everything from libraries to languages etc due to LLMs so even the rate of change across your OSS stack may cause these models to be dated quite quickly, at least in the core model which will require layering fixes.
To be clear I’m on the fence thinking about much of the same issues and as close as I am to pulling the trigger, I keep thinking of very valid counter arguments as to why it’s me just wanting this thing I own. Which may be enough.
I would say when this comes to pass, we are already 5 years along?
> Part of it is knowing that whatever sort of enshittification the cloud providers do, my local programming environment won’t ever be less effective than it is today locally.
I think this is quite understated. It basically is freedom from a growingly antagonistic relationship between you and some remotely hosted API managed by faceless corporates at the whims of their board, shareholders and governments.. It really is such a mental burden to need to constantly manage this relationship (watermarks, silent downgrades, random false refusals, downtimes, model sunsets, changing ToS's, fucking ads). These companies will need to squeeze you for every cent that they can before open-weight models are simply good enough for the valuable tasks we can throw at them.
To have your own hardware is to no longer have this mental burden.
The object permanence of not having to reinvent the world every time a model gets sunsetted has value.
with open models, there is ecosystem/market of providers, where you can easily switch to provider you like
Could be a long time till gets released
Anyway, yeah, even at the prices I spent on my local AI stuff (I bought before RAMpocalypse really kicked into gear, so I bought old server GPUs for about $350 each and the Strix Halo for a little over $2k) it was never going to pay for itself; I just like to tinker. But, I can't imagine spending today's prices for hardware for local AI.
When the memory shortage ends, I'll be down to the Apple Store (or, more likely, clicking refresh on the Apple outlet every few days). But, until then, there continues to be a glut of cheap and free models in the cloud that are better than anything I can run locally and they're faster, too.
*$1000? $14,000? Who knows but everything in the middle there has been claimed.
The cost of online services is also largely a result of the cost of training (though hard to say exactly what that number is). Assuming you are using open weight models at home, you aren't paying for the training - someone else is.
OpenRouter prices are somewhat simmilar to Antrhopic/OpenAI API prices. So I conclude that the hardware plus operating margin alone can genuinely produce prices way above what you'd pay if you had a subscription. Of course the primary unkown factor is average token use per subscription. Without that it's all wild speculation.
So, even though there are more models to run locally that can be useful for the stuff I do, it makes less sense now to do so than it did when I got it. There are more extremely cheap options, now, and it seems likely to continue to get cheaper and better and faster, while my local hardware will always be slow and hot and only gets better via software (which has roughly doubled the speed it can run stuff since I got it, but it seems like there's less room for improvement in software now, and even at twice the speed, it still kinda sucks to use local models interactively especially on the Strix Halo).
Homeowners can’t put solar panels on their roof to use the produced electricity?
Then again I don’t live in Texas let alone the US so i might not know where to look and I don’t care enough to truly find out.
I was just surprised that they were supposedly illegal. Which seems untrue.
Cybersec is also hit and miss, depending on what provider you choose, verification systems and all that jazz. Also, running locally allows you 100% data privacy, in any situation and for whatever usecase you might have. ~100k for hardware for a small team of devs to code locally is not that expensive in the grand scheme of things.
Lastly, local models allow for training / finetuning on your own data and processes. $/tok is not everything for everyone. Sometimes you can take a hit on value / speed if you get something else that matters for you.
That said, when I bought my pair of Sparks, the best model I could run on it was GPT OSS 120B. That has an AA score of 24.
Today, the best model I can run on them is GLM 5.3 Flash at Q4, AA score 57. Just still out on GLM 5.3 mixed quant.
So from that perspective, they are many times better value than when I bought them, and will likely continue to increase in value.
That AA score is for the original model only
But the cost difference between api and self hosted is so incredibly large now it's almost impossible to ignore the fact self hosting is a terrible deal. I'm waiting things out until the dust settles on what the required specs are and consumer hardware gets cheaper/more capable.
I'm happy with all of the competition in the APIs on openrouter... I watch that like I used to watch the stock markets, lol. It's great fun.
The first case of this I remember was Bloomberg thinking that their extensive private data about markets would give a home-trained model better performance in finance tasks. The next version of ChatGPT beat them.
With how often new versions of frontier models are released, you likely won’t finish your work before it’s obsolete. The labs have whole teams dedicated to specific getting training data for specific industries (finance is one), and more powerful generalized models make better decisions even without specialized knowledge.
Your best bet is to get really good at training for something and then sell your company to one of the frontier labs for their post-training efforts.
I hate how difficult it is to compare prices when looking at subscriptions.
Would $20 in open router, using models like GLM get me more or less?
Z.ai does have their own subscription, but I haven't used it because their privacy policy was pretty buns last time I checked.
I burn through my current Google AI pro sub for the week in about 2 and a half days so wanted something extra to add to it, but don’t want to buy any expensive ultra plan. Flash models have got me about 98% of what I need, but quotas are still a bit low.
What did you find objectionable? I looked at it when I subscribed almost a year ago and I was fine with it (e.g. they don't train on your API inputs).
"The Company do not store any of the content the Customer or its End Users provide or generate while using our Services. This includes any texts, or other data you input. This information is processed in real-time to provide the Customer and End Users with the API Service and is not saved on our servers."
tl;dr API cost (openrouter) is always more expensive than a subscription (for the same given model). you should always use a subscription first and only go to API pricing if you run out of your subscription.
In terms of which subscription is best, different ones provide different models, different amounts of tokens, different rate limits. So it depends on what model you want and how much you need to use it. The frontier ones are always more expensive than open weight ones, but a few subscriptions are starting to include frontier models like GPT-5.6 Luna (which is a great deal but not necessarily the best price-per-performance).
What's the point of publishing it when it'll likely be outclassed by gpt-oss?
AI policy is being shaped somewhat by the things Sam and Dario say. So even if you're not feeling vindictive, it's probably good to keep a track record of the previous things they have said as a Bayesian prior. People who don't know better listen to these people, and maybe they shouldn't.
I think the release of kimi k3 is definitely arguably dangerous, we're already seeing consequences of elite-tier cyberoffense capabilities.
> I think the release of kimi k3 is definitely arguably dangerous
I understand the argument, but I don't see how it matters. Get Sam and Dario to lobby to make it illegal to run in the US? Ok, then the bad guys who follow the law will run Kimi K3 from outside the US.
Maybe you'll petition the Chinese to lock down access to their models?!? Even if they did, it means the Chinese government can still use them.
The only solution I see for the cyber security side of things is to fix the bugs or migrate to more secure systems and approaches. We all had to do something similar in the mid 90s when the world came online and TELNET and FTP weren't good enough any more.
I cannot stand using gpt-oss, but I miss some of the creative spark of GPT-3 davinci dearly.
GPT-3 is probably one of the last models trained without synthetic data. Meanwhile gpt-oss is rumored to be only trained on synthetic data.
I also miss the 2020 GPT-3 based AI Dungeon fine tune.
The weights arent the truth tho, maybe a timecapsule-vhs but i wouldnt trust llm weights more than more hardcore deterministic media that might get preserved to infer facts from an era.
The companies doing the training are becoming the "winners" that are "rewriting history" as they train their models.
I say this with kindness: Anyone who believes this absolutely needs to turn off their computer for the week, go outside, travel a bit, and experience reality with other humans outside their regular bubble.
The “entire human collective mind” is not digital. It’s not on the internet. These models could’ve syphoned literally every piece of digital media in existence and still wouldn’t have it. People don’t exist inside computers, and it is naive to believe the sum of what’s online makes the sum of the human experience. It doesn’t.
- It is trivial to extract samples of the training data that was used, which can bolster existing lawsuits/foster new ones.
- Older models are not as safety-hardened, so it is easier to coax unsafe behaviour out of them, which is a PR risk.
- It may be possible to divulge proprietary secrets from the model (e.g. architectural details that may still be relevant still).
For these reasons, and more, it's unlikely that GPT-3/similar models will be released until these concerns are no longer relevant (e.g. when they become a purely historic concern, similar to the open-sourcing of other proprietary software from decades ago).
At least to this point, the legal teams could get the model via discovery. IDK that the lawfirms realize that they could get experts (or that they'd have contacts that could) to search the model weights.
According to the book The Thinking Game, lower quality models at that time were considered less safe, because they could be easily tricked into doing harmful stuff. In the book, Dario (of Anthropic) was the head of safety at openAI and was responsible for pushing for 10x scaling in training to make the models safer.
It does make sense, a smart model is going to be way harder to jailbreak into telling me how to synthesize anthrax (or something).
According to me, this is nonsense.
In the times of GPT-3 I'd scoff at the idea of an LLM doing any hacking; today, I'm running several AIs on my code before publishing, and they are finding (and demonstrating!) RCEs on my localhost server.
For example, one found a missing check in a third party JWT library which allowed full account takeover, which I'd have never even looked at.
Hence I don't believe a single word coming out of these people's mouths. Their "beliefs" are just marketing.
I certainly think the labs have muddied the waters using safety for marketing, but that doesn't mean less capable models weren't more dangerous at one point.
Easy access to malicious information hasn't proven to be the disaster these anti-prophets proclaim. For the last ~3 decades of widespread internet and search engines, you could trivially find all sorts of information (drug synthesis, explosives, etc.), and we're just fine.
(Corollary: easy access to good information did not make non-intellectual non-curious people smarter. Easy access to information does not have the consequences people immediately assume.)
My guess is I’m unintentionally refusing implicit signal that you think it’s safe to release all models openly, because you have observed models finding a vulnerability in a JWT library. But that sounds like a straw man instead of a steel man. Idk. :/
We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.
The good news is it can only get better from here.
I'm wondering of you could clarify your thoughts on this. I've had a hard time evaluating what Fable-class actually is capable of that sets them (or really it) apart from other models in a very significant way.
I have a hard time getting models like GLM 5.3 to not perform on my tasks but I might be biased.
I get how you feel, but it's too late to be concerned. The cat's out of the bag. It's like being scared of moving from the bronze age to the iron age... when everybody already knows how to make iron, and the raw materials are everywhere. People are already making iron spears. We need to make iron shields.
We need open-weight models that are good at finding security holes so we can apply them to all of our software by default, and close every possible security bug, before the attackers find them. Every piece of software in the world should be held for release until it's scanned by a high-powered security model.
This is the same debate we had in the 1990's when strong encryption was considered a munition and not allowed to be exported. This just made the world less secure. And it was pointless anyway, because you can't really stop it being developed and shared. Eventually good sense prevailed and now we all have strong encryption. The same thing applies to security bugs.
The insufferable gatekeeping of the US companies is actively contributing to computer insecurity at this point.
Apparently, Cami Clark was tight with Eric Schmidt. She seems to have pursued Epstein to invest in her "luxury porn" businesses, after a divorce & going bankrupt? Wild: https://www.wsj.com/tech/ai/claude-dario-amodei-wife-anthrop... /https://archive.vn/MJI7q