466 comments

[ 0.20 ms ] story [ 227 ms ] thread
Standard API Pricing for GLM-5.3-Flash (per 1M tokens)

- Input: $0.15 - Output: $0.50 - Cached input: $0.03

I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field.

It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

Starting? This was obvious way back in 2019, when the US decided to give China a little push developing their own silicon industry.
Well the big problem with china is that they do not respect international law when it comes to technology theft. But that argument is very weak when it appears that a lot of what they do is out in the open for anyone to replicate.
yeah, America is totally out there respecting international law.

"problem" indeed.

No major power respects nor cares about international law.

Intellectual property is part of WTO agreements but enforcement is domestic.

US companies do it too, regularly, they simply hire and poach staff from competitors.

Proving it to be IP theft is difficult unless you can prove documents being passed. But often all you need is the know-how of the hired talent.

There isn't one global "international law" for copyright. There are treaties that countries negotiate with each other.

If the USA wanted a copyright treaty with China bad enough, we would negotiate one. China is not breaking any laws here, international or otherwise.

if the americans didn't want IP theft, they shouldn't have taught tens of millions of people how to make their IP
> I'm starting to think

That's good. Keep going.

> It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

It's not like we didn't try it. China first have to learn to make deals where both party benefits.

I would say the Trump Administration needs to learn this as well.
America knew how to do it. But they are learning quick from China.
> with all of this traffic served on Chinese AI chips

RIP Nivida shareholders

yay, I called it! :) (in the other thread)
Another self-inflicted own courtesy of US government policy.

While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.

The export controls were revoked before it triggered Chinese protectionism: https://www.silicon.co.uk/e-innovation/artificial-intelligen... / https://archive.vn/B2pah
Revoked or not, just ever having those controls signals to the Chinese ecosystem that you're not necessarily a reliable supplier (Would you trust US export policy to remain stable for the next ~decade given the state of US politic?) and to the Chinese government just how strategically important you see these components.

This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.

> The export controls were revoked before

Zai is on another "export control" list outside the broader 1. Doesn't help.

The export controls were not revoked, only reduced, and not before, but after China refused to buy low performing chips. Top gear was and is still sanctioned, as is any EUVL equipment.
And to add to the above: by building their own supply chain for chips, China is helping the unprivileged, those who can't front-run the market with long-term contracts. If China wasn't producing their own chips, the prices for us would be even higher.

Similar to the war-pricing of oil, China's reduction of imports is actually helping to keep our inflation from going even higher.

And it doesn't matter, it still pushed China to speed-up their AI related hardware development.
Long term it's irrelevant. The only relevant thing is that there's lots of money in chips that can do high performance inference. You see all kinds of competitor products in development or already on the market even here in the US where there are no such restrictions. Cerebras comes to mind. It's natural and expected that eventually Nvidia will either have to keep way ahead or competition will catch up with specialized products.

That doesn't mean by any stretch of the imagination Nvidia will disappear. But the entire stock market valuation, not just tech, has had me scratching my head for a while.

Cerebras "competes" with Nvidia in the same way a Vespa scooter competes with a Ford F-150. Groq and Tenstorrent are in a similar boat, ASICs don't really threaten CUDA.

Curiously, there is not a single real CUDA competitor anywhere in the world. We almost had one with OpenCL, but all of the American stakeholders abandoned it right before the crypto/AI takeoff. All of which means that Nvidia sets their own margins, exploiting American investors and taxpayers while letting China avoid their dominance. So the American economy subsumes the bulk of Nvidia's arbitrarily-priced debt, and the Chinese economy can direct SOEs to pour billions in liquid cash into real GPGPU research.

I'm an American and I'm pretty fond of Nvidia, but Jensen was right about this policy; it gives China everything they need to actually replace CUDA. It's reminiscent of America's attempts to deprive China of ARM and Texas Instruments IP, only to end up swimming in unlicensed clones after refusing to sign an IP deal.

You forgot about AMD ROCm, Intel OpenVINO and Vulkan.
Not really a brag: it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time.

I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)

> it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time.

Boy, do you have a rude awakening in store.

Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.
Has there been any confirmation about what that model even is?

Edit: Ah:

> This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.

It's also in this very announcement, in the first paragraph:

> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.

It was being served for free. They were almost certainly being overloaded.
RAM was probably the bottleneck for the amount of context they were offering.

I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful

> and it was running very slowly

... I'm at a loss for words here. It was being served for free. To the entire world.

GPT-5.6 Luna is also served for free to the entire world with a tokens per second rate nearly 10X higher.

> ... I'm at a loss for words here

No need to be so dramatic. I think it's great that they're developing chips, but the whole "RIP nVidia" claim was overly dramatic.

Do you know how much traffic luna was getting vs Ox Alpha?
Are you really comparing chatbot to agentic/code work?

Why is Luna not free on OpenRouter? :)

Lead doesn't really matter anymore. I just ported a very old cuda library to rocm, so it can be run on MI300s. 2 years ago this would have been a nightmare. Today it was an afternoon.
Exactly. Coding for inference is solved. CUDA is no longer a moat.
Ox Alpha was also serving 10T+ tokens a day for free.

When it first launched on OpenRouter I was getting nearly 70 Tokens/second.

I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini)

And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it.

So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.

I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.

If anything it's custom chips from the labs that threatens Nvidia.

I can easily see a situation where most non American AI usage is on Chinese models on Chinese chips though.
These open models serve as price / performance pressure. Not all tasks require frontier models and cheap open models can be quite good for in-app assistants, if you're building that sort of thing. We also aren't sure the subscriptions will continue to be sustainable. They're currently subsidized to the tune of 50-70x. As someone who is hitting limits weekly that would easily cost me over $10k month per sub.
Casual consumers are using American models because their usage is low. As usage scales, the economics heavily favor open weight models. The API pricing from American companies is absurd. This is particularly true in an enterprise setting.
Open weight model hosts don't have the compute to meet enterprise demand. A large part of why these models are so cheap is because overall demand for them is incredibly low. Back in May, Gemini alone was doing about a month's worth of Openrouter tokens every day.
I disagree totally. DeepSeek raised prices because they couldn’t serve the demand. But there are tons of American vendors ready to fulfill it. Many enterprises, including the one I work for, are swapping to open weights.

Why wouldn’t you?

I am not ok with handing all my data to American companies that are best friends with the American surveillance state. I still remember the Snowden revelations. Chinese companies are a much better option in that regard.
you don't have to hand them your data, the models are available so you can run them on bedrock yourself (or use another US housed inference service). and for what it's worth in my job i have access to data that gives a picture of the way companies are doing inference, and they're using a lot of chinese models (deepseek-v4 is a huge percentage of inference requests for example)
Most US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market.

Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)

> Most US companies that have anything to do with government, finance, medical, etc... That's a huge market.

Compared to the rest of the world?

I don't have any pie charts in front of me, but yes, I would estimate it's a decently big slice of the world market.
Not really. Chinese AI companies were never using NVidia AI chips.

This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference.

Also, NVidia chips are still sold out and supply constrained.

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash

I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.

I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.

I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

[delayed]
[delayed]
For me it's entirely because I have a bunch of projects with my own personal data that would be tough to do with openrouter/claude or any other cloud.

For example, I have a small posix-shell-based LLM harness that can SSH into my NAS and run organization tasks using the local DS4Flash that I have right now. It's already been a massive help for me to keep me organized, and that's just 2x DGX Spark's worth of compute.

It cuts both ways. A GPU in your basement is a depreciating asset with fixed computing power and consumes electricity. Switching model providers is trivial.
> A GPU in your basement is a depreciating asset

All decades prior and up to about a year ago, I would have agreed with you. My Framework Desktop, however has appreciated in value by 75% since I bought it. Will it stay there for a long time? Probably not. But it shows that there are no hard and fast rules about things anymore.

I just bought a Framework Desktop. Would have been nice to get it at the introductory price, or perhaps the new 192gb model refresh they’re now teasing, but I settled and got a 64 gb model. At the time, the 128’s price had already risen again, but the 64’s price was still at a lower price.

64 can still easily do a Qwen 4.8 model, so I’m relatively happy with my purchase… plus, it’s price change has caused it to quickly appreciate in value… so I could sell it if my situation ever turned dire lol

At the current point in time I'd argue it's more about opportunity cost/value.

If I'm a professional photographer chasing the best possible end product, I'm not buying cameras because they're economical. I'm buying the best camera I can get my hands on to get the best product I can produce within reason under the understanding that it doesn't have to equate to the best economic decision to be the _right_ decision.

If you're in a position to be able to take advantage of the local inference - it's a no brainer. If you're not sure how that would be done, then it's not a good move.

I'm not trying to say there is no use case. I just want to know the cost. Is it less than the API cost? Is it the same? Is it more? I'm looking for hard numbers. If the cost is the same or more, then the decision for local isn't to save money
If your usage wouldn't change with local inference and you don't have security/privacy concerns then at the currently heavily subsidized pricing, sure.. not economical.

But things change real fast when you're no longer bound by costs/apis/rate limits. All of a sudden it's not about "how can I do this right and efficiently" and more about "I can poke at and test _all the things_ that might make this better".

I think most people who can't see this value in the local inference approach are likely still copy/pasting from their web LLM ui's or don't even come close to subscription quotas. Meanwhile, 1b tokens a day is a light day for me with 3 $200/m subscriptions + some level of sub at basically every frontier level provider. Had I been less frugal and ponied up for the hardware before things got crazy I wouldn't need 80% of that - just the frontier models for the most complex tasks, the open weight models would handle the rest easily _and_ I'd get to do a lot more exploratory work without concern about quotas.

What do you do with all those tokens?
And here I am, feeling a bit guilty for using between 2 and 5M tokens... since 1 August!

Employer just sent an email that.. things are changing when it comes to token spend...

> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly.

There are a lot of social media posts about people cancelling their Anthropic or ChatGPT subscriptions after installing a local LLM. I’ve used local LLMs a lot and I spend a lot of time with frontier models and the difference is still huge. As far as I can tell, the social media posts about local LLMs replacing frontier models are either wishful thinking, engagement bait, or people who must be working on much simpler projects with a much higher tolerance for slop than I have.

I have exactly the same opinion

Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it

The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled to how differentiated their marginal contribution is to peers. They’re much more incentivized to spend their personal/work time optimizing for being more skilled or acquiring some kind of competitive advantage relative to baseline.

Most people don’t consciously run the numbers of “I get paid $X/hr to add $Y of value” or model pay at work as something with variable inputs (eg something that can be increased with high performance), so it makes sense to them to spend 20 hours of time to save $100 or to make themselves 5% less efficient to take home 0.5% more or avoid doing something they don’t want to start doing.

NOT saying this always happens or that they’re stupid for doing so. I didn’t even realize how much I had been doing it myself until I started recognizing it, and shifted to having my own comp/performance fully aligned with the company’s P/L.

It actually makes a lot of sense IF you can accurately estimate incremental upside (which is much harder and more diffuse than modeling downside if you’re salaried a employee) or if the upfront skill/knowledge investment that looks like bikeshedding pays off in the long run.

These are great points. It's a little off topic but what you bring up is why i advise new grads to spend the first couple years of their career in small eat-what-you-kill companies. I think software devs who start out in large companies get this distorted view that their twice a month direct deposit is just magic and comes from the ether no matter what they do. The whole industry would be better off if everyone started out in a "you don't deliver, you don't eat" company and grew from there.
Strong agree, but I also think some roles in big companies (for me, infrastructure) or in certain industries (eg trading/finance) can help build the same understanding without as much of the variance/raw exposure to bottom line.

Now that the role of the ticket-cruncher is on the path towards full commoditization, and individuals can move much more quickly (and even more carelessly!), I think product roles will probably shift towards one where developers are more deeply embedded in the product/business process so that they own/understand what to build without as much separation between the decision-making and prioritization of what to build. Or at least, they should.

It was eye opening to me to run the math of "should X people work for Y months on this project to save Z per year?" and realize that in so many cases, the time and effort it would cost to stop "wasting" money on things is WAY more than you could actually save on it. Even "small" projects can very quickly become $1M+ investments in time and resources, and the diminishing returns add up quickly (but also a good way to justify the value of your contributions, when done). But the job only exists if it saves money or makes money...

While starting my career in the late 2000s in web design agencies wasn't good for my stress levels, it absolutely gave me an appreciation for some things that working at a ~15k employee corporate just doesn't. I can tell exactly which of my coworkers come from the "outside world" vs those who joined here as a grad and have only ever worked here haha
This is really insightful, thank you.

Can you share a bit more about how you shifted to be more aligned with P/L? And how to accurately estimate incremental upside?

I'm an early PhD student with interest in ibdustrial research/R&D, and currently struggling to understand how to think about how to navigate through my career.

I'll try to give you a full answer even though maybe only you or a small number of people will see it, given this is a couple days old now.

Regarding P/L: I left big companies to found a company, and have a majority stake in my startup. I only really can make money off of it by having the company make money or be valuable, and the credibility/accountability/etc are intertwined.

My stake is predicated on making it my fulltime job + having to go "down with the ship" until it gets big enough that it could credibly go own without me + investing into it. That essentially locks me in to the company. Because it's a C corporation I am incentivized to grow the value of the company without taking anything out for myself beyond what I need to live (if that, because it makes no sense to do that with the money I invested into it, that's just generating income taxes for no reason). I only really can make money off of it by having the company make money or be valuable, the credibility/accountability/etc are intertwined, and there's no incentive to invest my own money into it unless I actually think that's the highest leverage use of it + want to capture equity upside for myself (rather than spend other people's money to accelerate a career or pay myself a salary while also capturing a large share of the upside).

This is a common reason that people talk about "founder-led" companies: the incentives are very different vs career "operators"/executives and investors, especially when there are only 1-2 founders with majority or near-majority (30%+ well into growth/at IPO) stakes. Examples: Google, Cloudflare, Oracle, Facebook, Microsoft, Berkshire, Ford, LVMH. The incentives for long term growth of the company don't get as messed up by the principal agent problem: https://en.wikipedia.org/wiki/Principal%E2%80%93agent_proble... That's because the accountability/reputational/financial fates are directly tied to each other: all upside for the company is essentially incremental upside, with very little opportunity for incremental upside at the cost of the company's (because that would be stealing from investors/owners, and basically the only way to lose what you've built, just to steal from yourself?).

Only a small portion of venture capital-backed companies operate this way these days because their LPs (the people whose money the venture capital is investing) need liquidity on predictable timescales, generally ~10 years from commitment, 7-9 from actualy capital calls, but much better if the specific fund vintage can be derisked/validated by other investors or liquidity sooner, so they know it will actualy be valuable + have buyers when the time comes. Also, more founders + more venture capital representation on the cap table = more ownership and decision making stake optimizing for profitting off the company with similar time horizons/diffuse responsibility, rather than making the company profitable or concentrating decision making/risk/ability to manage the company in a small number of people. To be clear that's still much more of a "shared fate"/long-term horizon than most public companies' executives or investors, or employees most places, but it introduces a lot of misaligned incentives (each pumping short term valuations to exit, preferring big risks because funds manage diversified portfolios and will win across the category regardless of any particular bet wins its category, the team can just fire the CEO or start over with more pension fund/family office money if it busts).

At big tech companies (Microsoft and Google) I felt well rewarded and that there was a direct link between my effort/effectiveness and compensation. But this is how much tech companies make per employee (and the amount they're worth per e...

> the difference is still huge

Same experience here. But I think the key argument for local is about being able to leverage "non castrated" models. But maybe this is not relevant at all for standard coding tasks.

I will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be so much worse that it's not worth it.
Hopefully you also bought a switch
they have 2 interfaces each so you typically daisy chain them
That will hurt latency and latency is very important for good tensor-parallelism performance
I think I need one now that I have four - reduce is ring-oriented and still works I believe, but IIRC you only get 200gbps if you use _one_ of two connectx ports.
They’re 200gbps interfaces. The switch needed for that throughput is going to cost as much as the machine itself or more.
Qwen 3.8 27B is around Opus 4.8 level of capability on the Agentic Intelligence Index (52 vs 57). In my testing the locally hosted Qwen is good enough that looking at a given piece of work output I couldn't tell you which model was behind it.

https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...

As a counter to that - I've tried various flavors/quants/full weights and Qwen 3.8 27B has been entirely useless at anything non-trivial. Sure - it can do some boilerplate work (though, even armed with a well written spec and working within a very well known framework it went off the rails and did things in a way that were... um... questionable at best) but I don't see it as anything more than a personal assistant style model. Zero chance I'd "work" with it, I spent days trying to get it to do something for me that was usable that I didn't have to have reviewed and refined by a frontier level model or myself. Couldn't do it. The idea that qwen 3.8 27b is _anywhere near_ Opus 4.8 is laughable. Pure benchmaxxing.

DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend.

GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.

I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side.

The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity.

It's the first small local model I've felt like I can do real work with.

[dead]
Your assumption is incorrect.
They also need to tell people what quantisation they're using. Because some 4 bit version is not the same as BF16, no matter what KLD suggests.
I'm using the unsloth dynamic Q4 and getting good results. I was running Q5, but Q4 gives more context headroom so I can run two agents in parallel with ~100k context each with 32GB vram.
What harness? I've had similar results as _Implicated_ said above - it's not done well in any of the tests I've tried with it. I currently have it hung off DS4Flash as a pseudo-vision tool and subagent only because of this.
I'm using pi inside a self-made harness. I've found going super lightweight with context (AGENTS.md is maybe 20 lines) and letting the model discover what it needs to gives the best results.
I can't get 3.8 to exit thinking loops. It will just think and think and think on the most trivial topics. I wanted it to port a speed test powershell script to c#. Claude opus 5 completes it under 60 seconds. I let 3.8 churn about 6 different times for 30+ minutes and it never wrote a single line of code to disk. It wrote lots of lines in thinking.

Any tips?

That sounds like something is off - I'm using UD-Q4_K_XL on pi with xhigh thinking, and unless I'm vastly underestimating the complexity of the script that's the kind of task I would expect to take a couple of minutes (getting ~30t/s decode). What server are you running, and are you using the recommended parameters from qwen/unsloth?
I've found it tends toward long thinking loops even for simple tasks (and any quantization seems to increase their length), but those do exit eventually, unlike with Qwen 3.6.

I use the Unsloth UD_Q2_K_XL GGUF with default parameters, along with that custom template linked elsewhere in the thread, and no K/V cache quantization.

For smaller models, you'll probably find anything below Q4 will need handholding. Check Unsloth's graphs at the different quantisations VS error rates and you'll see why.
Thank you, I know about those. And I'll stick with this quant. Normally I'd be with yout there, but Qwen 3.8 is turning out to be good at self-correcting, and the free memory I can use for extra context is worthwhile.
My experience has been that anything less than a 4-bit quant has a tendency to go off the rails. There’s a threshold of coherency that is being crossed somewhere internal to the model I guess.

Try the same prompt with a larger quant (even if it runs very slowly because the model no longer fits in VRAM) & see if Qwen does better - if so, there’s your answer.

Both of you should mention what quant you're using. And as another comment said, what tasks you're doing, i.e. coding, classification, summarizing etc.
I'm using unsloth dynamic Q4_K_XL.

My use-case is coding, currently working on a project with a Rust backend and TS/React/Vite frontend, with probably tens of thousands of lines of code total (including tests).

Interestingly, I've found the output of Qwen 3.8 27B to be competitive with Opus (4.6-ish anyway, not quite 4.8), but the experience is very different.

Where Opus has seen it before and knows how to do it, Qwen knows how to work it out. It turns out it's surprisingly capable at working things out. The obvious drawback is that it takes tokens and time.

All the same, getting to run something this capable locally is momentous, and suggests to me that streaming tokens from colossal data centers might not be the long term path forward.

Lately I've been throwing tasks at Qwen and a frontier or recently-frontier model (as well as Kimi, GLM, etc) and the smaller parameter models are not really comparable to Opus when it comes to making intelligent decisions about greyer areas of good software architecture.

Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).

I’ve been doing the same thing, giving the same tasks to Qwen 3.8 27B and Opus, and the main difference is that Qwen does not consider edge cases which Opus catches. It’s good at the happy path, but even when hinting that there are uncovered edge cases and gotchas it’s oblivious to it. So I feel like I need a bigger model to do planning/review.
To be honest I’ll ask a model to specifically think of edge cases but I won’t expect any model to do the edge cases of its own volition
I got so much better experience LLM-Chunking(think RAG) with qwen-38 27B ONCE i move the thinking effort to HIGH vs XHIGH (i think is the default on Open Router).
I gave Qwen 3.8 27B and Opus 4.8 the same task in the same codebase. They both came up with the same diff. It wasn't a particularly challenging task (removing a feature flag and updating applicable specs), but it was character for character.
Qwen3.8 27B (which I adore) is nowhere near Opus 4.8 at puzzle games testing fluid intelligence, https://quesma.com/blog/baba-is-aug-2026/
yeah it's more like opus 4.6 iirc
Not really compatible on all fronts, it's very capable especially with tool calling, workflows, logic and its base coding ability, but it's only a 27b model so it does not have anywhere near the level of knowledge baked in as larger models. This does not mean that it's not a good or useful model - it is on both accounts and very efficient, but it's not similar to a large model generally speaking.
I am surprised. I've been using DS4 Flash (0731) for weeks now and it works perfectly fine as a replacement for Claude in a large variety of cases. It requires a few more iterations, sure, but it's useful enough to not need a Claude subscription anymore. Among the things I do I've been reverse engineering, writing complex C++ code...
How much did you drop on these 4 sparks?
Sparks + cables + 10g SFP+ came out to ~$21,500 CAD
Sparks don't have enough memory bandwidth, for the same 20k you're better off buying RTX or Apple M5 Ultra machines.
I ran 30M tokens through for 50c... insanity that this is possible.

and it really is opus 4.8 level.

For those who didn't read, this is the identity of the mysterious "Ox Alpha" model
Yeah, made me suspicious of how well the Ox Alpha was performing that it wasn't some 'new group' making the model.
> 320B total parameters and just 18B active parameters

This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.

Looks like the M5 Ultra Studio wait times are going to increase again. Already at 10-12 weeks, I wonder how long it'll go?
I guess like the M3 Ultra, at some point normal customers won’t be able to buy it.
That M3 had an older type of RAM. Apple hopefully secured sufficient supply of the newer variant for the M5 Ultra.
Both use LPDDR5x, they're not shipping LPDDR6 (yet).
All M3 variants use LPDDR5.
Speaking as someone who isn't really well versed in this, does 18B active parameters mean that you could potentially hold only the 18B parameters in RAM and stream the rest from a fast NVMe SSD for acceptable performance similar to how Colibri works?

https://github.com/JustVugg/colibri

Normal MoE is switch-weights-per-token so you would se substantial slowdowns that way. Apple did a More that switches weights per prompt (instruction-following pruning, https://arxiv.org/abs/2501.02086) but you have to design the model that way which I don't think the have.
from the article, pareto frontier for open source models is completely dominated by GLM now.
Well, it will be interesting to see where Qwen3.8-Flash-Next ends up landing, also released today. These are exciting times!
I find GLM's idea of fast/flash is not really competitive with the speed DS4 Flash has, and it's hard to see them as being in the same segment for that reason.
Even vision? Thought k3 might have an edge there
> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

From a biased source, but would be big if true. I've had great results with GLM 5.2.

From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.

> From a biased source, but would be big if true. I've had great results with GLM 5.2.

It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.

The recent and slightly smaller DSv4 Flash is also GLM 5.2 equivalent (or close enough)
DSv4 hallucinates much more than GLM-5.2 though.
It's only 320B, local frontier AI is getting closer, sooner than expected.
It's not possible to keep shrinking down parameters and keep "frontier" performance, it's like saying it's possible to take a 3 hour movie and compress it down to 3 megabytes, there are information theoretic limits on the amount of bits of information that can be compressed.

What I'm saying is, if you're expecting a model that can be run on a 16GB or 32GB machine with the intelligence/knowledge of Mythos or Sol, it will never happen. It cannot happen, just like you cannot watch the Odyssey saved as a 16MB file.

Smaller models can get faster and smarter, but by definition they can never compress all of the knowledge of a frontier model and they will approach a limit by which they cannot get better.

(comment deleted)
The current models are not close to approaching the limit of compression for intelligence. They aren’t even focused on it like Chinese labs are. The training of Qwen’s 27B parameter model showed that by structuring model training from fundamentals to more difficult topics they were able to drastically reduce the number of parameters needed.

The ‘frontier’ models rely on scale to achieve their results but that’s not the only approach. Eventually we will hit up against the fundamental limits but we are not close with Sol and Mythos.

Yes they are approaching the limits, try asking smaller models niche questions about almost anything, they hallucinate massively because you cannot simply pack in all the raw knowledge from a massive frontier model into something that’s quantified down to 20GB etc.

It breaks fundamental laws of information theory. It’s like saying you can extract 100 joules of energy from 10 joules of energy source. Not possible.

It doesn't really matter though. Hardware performance is still growing. The new Mac Studio could just about run this model locally (rather slowly) - something that sits on your desk, that you as a consumer can buy.

Imagine prosumer desktop hardware 10 years from now. The 2036 DGX Spark. For a few thousand dollars you will be able to buy something with hundreds of GB (maybe TB if manufacturers step up) of unified RAM, memory bandwidth in the 10-20TB/s range. Overall AI "compute" will increase 10-20x, while at the same time AI model capability per byte will increase 5-10x.

The hardware would fit today's models, something like Kimi K3, quite comfortably and give performance of maybe 100 tokens/second. So what needs data center hardware today will run on your desk.

But if we also assume the models become more efficient, a 2036 Fable-class model (in terms of intelligence/capabilities, not size) will easily run on this thing at hundreds of tokens per second.

Unfortunately it'll still slow to a crawl with 5 Chrome tabs open, and every Electron app will need at least 200GB of RAM.

Sounds like you are describing a quantized model which is a naive form of compression, not a model that is trained more efficiently.

Additionally the information theory angle is for information storage, but a model can access resources and tools to gain information and what we are really seeking to train is reasoning not information retrieval. We reduce the needs to the right capabilities and we don’t get upset if it does not know the lyrics to every song ever written.

I think the assumption here that might not hold is simply that increases in efficiency and smaller size will be achieved by linearly just training smaller models better.

You are absolutely right that there is a physical limit about these things, but very often I find that the solution is a clever way to work around the problem. Maybe the problem with knowledge of the models will be improved by them looking the information up in a better way - so smaller models will not have to have the knowledge trained in but will default to checking. Maybe Models will, I dunno, focus on training in assembler and start to only ever check the compiled output so they only ever need to learn assembler and will then compile the solution to reason about the assembler code.

Obviously that last part is a ridiculous example because I'm not gonna be able to come up with a solution myself - I'm not nearly smart enough for that. But I h ope you get what I mean. Not going the direct route but instead finding solutions people didn't think of before.

You heard of JEPA? LLM's have all sorts of garbage they have memorized. Reasoning in latent space instead of in text significantly reduces the number of needed parameters.
JEPA is a joke, let me know when those models do anything useful.
You should look at some literature around it. I don't have time to pull it up now but it's been shown that much smaller small million parameters JEPA model outperforms much bigger LLMS in some applications. Keep in mind JEPA is area of active research.
> Combined with our latest 30T-token multimodal pre-training corpus [...]

Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?

How much is the “discounted” pricing they mention?
> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.

Now translate this to physical world, robots building and optimizing other robots... getting iRobot (2004) vibes
By clicking this link you download some PDF in the background
(comment deleted)
This looks like it goes hard, can't wait to try it
When reading this type of announcements, always have keen eyes on graphs.

e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.

- This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)

I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)

If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute).

I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.

We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildouts still happening.
Has nobody from any of the companies hosting open weights models released detailed information on how much it really costs?
I’m sure someone does, but I’ve never seen anything other than vague statements like Anthropic’s “inference is profitable” comment. I suspect everyone is playing everything close to the vest because they aren’t yet public and they want to control the information flow to the street.
> I don't see how NVIDIA can keep their spot as belle of the ball.

FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications.

With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.

Why is their own coding plan always the last place z.ai release their models? Its even online, you just have to guess the model settings.
Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it.

https://deepswe.datacurve.ai/

That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost. Roughly equivalent to sol medium, at a fraction the cost.

They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts.

Congrats to them!

Opus 5 is better than Fable in this benchmark?
Even Artificial Analysis has Opus 5 better than Fable in their aggregated "Intelligence Index" which combines 9 benchmarks. Opus 5 is heavily benchmaxxed.
Depends on the benchmark but yes. I think Opus is more heavily optimized for coding. On the usability side, its output is almost intolerable to read. It seems to code fairly well. Fable is more enjoyable to use for planning/interacting with
> They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts

It's what people know. Opus is just the common target.

> Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash

The problem with this and DeepSWE is it goes for a very specific profile. I'm not convinced DeepSWE is any accurate in actual work. It's surely a different signal (compared to some that allow cheating) but it has its own issues, e.g. weak harness.

Luna is great at following instructions but bad instructions or anything not covered = death.

Deepseek is more analytical. Good for bug tracking.

GLM is a better all rounder in some ways. Better at creativity.

> I'm not convinced DeepSWE is any accurate in actual work.

They listed Muse Spark 1.2 around DeepSeek V4 Flash even though it's a much shittier model in basically every aspect.

> GLM is a better all rounder in some ways. Better at creativity.

I agree with the creativity part.

Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology.

Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium tends to be rated pretty poor on agentic tasks compared to the Chinese lab open models, but I find the latter much more likely to lose track of their own behaviour during a long-horizon task or get stuck in a doom loop.

(This isn't a comment on GLM-5.3 Flash as I've not used it!)

Opus is a yap god. I've found it much, much better with Claude Code's output style set to `Concise` and this:

https://news.ycombinator.com/item?id=49413456

We shouldn't have to resort to this, but it can be mitigated enough that it stays as my daily worker agent. Although, I mostly use Fable to farm out to Opus agents so I don't have as much exposure to what kind of blathering is going on in there.

It's crazy that sometimes I ask Opus 5 to explain what it just wrote to me, and it declares "that was word salad" (its words, not mine, without any hint from me other than "explain it").
For me it's been this: Opus (at least in my experience) is unbeatable at "planning the work". That includes a lot of things, including getting arch. sorted, a chassis/skeleton done. Filling that up and doing the actual "coding," though, I've noticed no real difference between the Claude model and GLM. So it'll be interesting to see how this flash model compares cost-wise to what I'm currently using, which is GLM 5.3, for coding. Looks like it will be reduced even further and might be great if it's faster (and better?) than 5.3 in my real world/personal experience.

(I'm someone who doesn't really care about delays of a few seconds, or even more than few seconds. But if I am trying to notice then sure Claude is definitely faster as well).

Yeah, I just ignore the benchmarks at this point. For open-weight models the provider's setup impacts performance so you can have different experience's with the same model at the same quantization from different provider's. Just have to use them on real tasks with your actual harness to really know how they will perform and hope the provider doesn't do something to degrade performance (e.g. update the middleware to a new version with a defect that impairs performance).
I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again.

I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a single line of code.

The one time I tried asking Sol to use subagents for a small project, it took a surprisingly long time, used up the entire usage limit in one go, and basically failed the project.

I’m pretty sure that plain Sol, serially, could have finished the task faster, cheaper, and far more accurately. I’m also pretty sure that any competent subagent orchestration could have gotten it done with even very simple subagents quickly and cheaply.

(Is it really that hard to set up a handful of subagents that all use the same initial context and to load that context with what actually matters? The APIs certainly support it.)

OpenAI has been doing wonky stuff with subagents, including encrypting the prompts sent to subagents in Codex. Who knows what’s really going on.
I'm in the same boat. I haven't found sub agents flows useful.
Luna max is all I use. In my experience, it works really well for overnight tasks.
I feel the same. I just use it for planning and chatting. Not real coding work.
If your code is complex enough for Luna Max to fail maybe you need to write a bit yourself so they can copy your idea
This is why I stop at xhigh.
I only use Luna (max), I find it very rarely just reasons. In fact, I find it reasons too little.
I've had the same observation that Luna will quickly fill up its context window with reasoning, but it surprisingly hasn't been a problem really.

It will cycle through like 3 /compacts, complete the complex goal successfully, and cost me like 1% of my weekly usage on the $20 plan.

Edit: this is me using Luna directly, not Sol as the taskmaster

Wonder how much it hallucinate. I like v4-flash but its very keen on making up nonsense. If 5.3 flash takes after the GLM-5.x family if might be a very interesting flash model.
On OpenRouter the pricing is: Input $0,075/M - Output $0,25/M - Cache Read $0,015 /M

How is the business model of Anthropic/OpenAI will sustain?

They're obviously in a pickle, nobody is going to continue to pay $15-50 a mm tokens here soon. There's a reason OpenAI stopped training large models last week, and it's not because of "saftey" or "alignment" they know these gigantic models are not worth the squeeze.
This is a bad model. Worse than Luna in every way; slower, dumber.
OAI/Anthropic shareholder? Speed and intelligence are not "every way". Cost is essential. Hence the Pareto boundary illustrated in TFA.
It literally cannot complete tasks that Luna can do easily. It doesn't matter how cheap it is.
I think anthropic is behind but Luna on a Jalapeno seems profitable