99 comments

[ 0.22 ms ] story [ 13.6 ms ] thread
Just after DeepSeek-V4-Pro-0813 published, is this on purpose?
Fable level performance, faster and significantly cheaper. Wow!
Fable-like intelligence, beats GPT-5.6-Sol on most benchmarks, cheaper than Kimi K3 on API and quite generous usage on Cursor subscription.
(comment deleted)
gpt 5.6 sol and fable 5 level if the benches hold
[flagged]
I generally find people prejudging without really willing to understand other people's perspective shortsighted. Funny enough, they are often very much like the people they're judging. They're just in the opposite camp.
Thats actually a lot more impressive than I thought. At least on paper
Wow, OpenAI is now 4th after Opus 5, K3, and Grok
Pricing pages haven't been updated yet, still advertises 4.5
Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations:

1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?

2) Distillation - also implausible for the reason above.

3) Benchmark hacking. AI companies have ways they can dial up performance artificially, and will reach for that to maintain the appearance of parity.

Other reasons?

Edit: Most replies are ignoring timing. It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.

It's about chips with a large enough scale up domain. Larger domain allows for bigger model, which is what's driving this jump. You've got to get the chips, test them, tune kernels, then start a big pre train, mid & post-train, and only then do you actually get the model. So it takes time. Anthropic got there first partly because they use different hardware (TPU I think, maybe Trainium) which had larger scale ups earlier.
Do they have Fable level models? Or are they all saying “hey we’re dangerous too!” and hoping to get some token spend out of it?
To the best of our recorded knowledge, nobody ran a 4-minute mile in the five millennia prior to Roger Bannister in May 1954[0], but more than 2,000 people have met or exceeded this achievement since. In fact, his record stood only briefly, being bested the following month by John Landy.

The moral of the story? People work in parallel on the same goals, they build on best practice, or sometimes just need to see something is possible (reusable rockets). Having achievements cluster like this is normal and expected.

[0] https://en.wikipedia.org/wiki/Four-minute_mile

There is a herd of companies all running a race. The technology is known. They all have roughly the same resources. It’s not unexpected that they have similar cycle times for model development and that those models will be of roughly the same quality. Then layer in corporate PR demands and you see all these models landing within weeks, sometimes days, of each other to keep the model developer’s name associated with “frontier” development.
It's a combination of (1) and something you don't list: I think the frontier labs all have multiple generations of undisclosed models in continuous training. There is no "end point" when it's magically "ready". It's just getting better and better all the time. What they release with a name and a version number is just a marketing / branding exercise.

So what you experience as a "near simultaneous" release is just their decision of when to peel off a release from their current set of in-training models, likely based on how they perceive market and regulatory conditions. They likely see a competitor release and then baseline what they should release based on that and it takes a month or two for them to package it up and push it out the door.

What I can imagine is that for some of the labs, they are being forced to publish models closer and closer to the frontier of what they have in training. Effectively, "falling behind" is your forward pipeline shrinking. Google ran out of forward pipeline. So far Anthropic and OpenAI didn't - but probably, one is shrinking.

I've been hearing this myth since ChatGPT came out - that the labs have superAI that they're just slowly trickling out as competition forces them to.

Everyone in silicon valley has a cousin who's supposedly seen Anthropic's new unreleased model that changes everything forever. I distinctly recall sitting in a work meeting where a coworker was insisting GPT4 was AGI that anthropic was too scared to release. It's amazing to me that this playbook is still working at least somewhat on folks.

I think a lot of it is just time. The quality of a model is E * C

Where: E = Efficiency, and efficiency gains come from quality of data, quality of algorithms. C = Compute (Size of model, flops of train run)

So a better company can train a bigger and better model with less required compute which let's anthropic get there first. If another company does the same thing with a worse: model architecture, kernel, optimizer, etc... They will get there as well if they just run there train run with more flops for longer

Mythos was actually ready about 6 months ago. So if you have 6 months later or hardware setup and time to train you can get a lot done.

My theory is that it all boils down to better data and longer post-training period. Cursor got curated data from the trillions reactions of real world developers in real jobs. xAI bought is and used it for its post-training and got Grok 4.5 . Longer post-training on the powerful Colossus cluster helped it get Grok 4.6 , although both versions use the same model with the same number of parameters. Thus, both must use the same pre-trained model as a baseline. See also an article infers the training and release timeline of popular models featured a few days ago here on HN.

Chinese labs must follow similar trajectories plus their specific efficiency improvements. That also explains the jump from DeepSeek 4 performance in April and July releases. They both use the same pre-trained model as well.

Now GLM 5.3! And they explicitly confirms my theory: "Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — all running on the long-horizon task environments we have been accumulating. Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them."
4) Algorithmic improvements are either relatively easy to find if you already know the system can do better, or they don’t provide an edge that can’t be overcome by increasing training compute.
I predicted this exact event several months before Fable. ,not in a provable way, but the reasoning was related to a paper I read from here that I basically self-internalized as variability knowledge. Two very similar papers, one unfortunately named.

I also stated recently (in informal conversation), based on the performance posted, that said variability was only applied to specific fields of information.

So allow me to make a more provable prediction:

There will be another significant jump related to full field converage, followed by another and from there (we'll call this v3), it will then be capable of automating ASI.

4) Elon has access to some Anthropic models because Dario is desperate for compute and bought some from a competitor
One company making a big release both reduces the risks of training a big model (you know it can work) and increases the risks of not doing so (you are bleeding market share).
>Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass.

As a designer, I'm always hesitant to believe these statements until there's independent comparisons between the old & new model, as well as comparisons to human made flows. Design can be so subjective that blanket statements like this seem almost useless.

As polarizing as grok is, it was basically inevitable for it to start being a real competitor given how much investment SpaceX made into its own inference capabilities.

Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price.

I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation making it less appealing to many.

Is it inevitable? Still waiting for (also massively invested) Google or Meta competitors at Fable/Opus/Sol levels.
Still not dead somehow even though they've been renting out datacenter capacity and other (seeming) problems with people leaving and so on. Quite impressive unless it's just been benchmaxxed.
Tangental, but has anyone else noticed grok's voice mode got stupid and terse ~2 weeks ago? I've absolutely loved grok's voice mode since it came out (incredibly useful for brainstorming on walks and helping conceptualise and get the verbiage for expressing ideas) but it seems so have lost about 40 IQ points recently, and if the question is multi-part, it often answers just one part with no elaboration or explanation of the other parts or interactions between parts. No clue why.
Yes, that was Grok's strongest point for me previously, and it's been basically unusable these past few weeks. I think there have been a few posts in the Grok subreddit (maybe on r/LoveGrok). It has a lot less personality which is a shame, but I'd take that if the answers themselves were good - but they lack information, have zero nuance, and repeat themselves pretty often too. Such a disappointing change.
It's crazy that I'd literally trust a Chinese AI company with my data over anything Musk is involved with.

Like, even if you don't care about (or even like) his politics and can look past how unlikable he comes off as, the damage he's done to his own reputation in this domain just makes using his products like this a no-go. He's literally so rich that he can get caught personally looking through chat sessions and it wouldn't slow him down a bit. He's too rich to be held accountable, and that makes it impossible to trust his businesses. It's a funny dynamic that I don't think is appreciated enough, but I know that if Google or Amazon or OpenAI or Anthropic (etc.) got caught doing something like that, the backlash would be astounding and the reputation hit they'd take would be brutal. Here, Musk would just awkwardly come out attacking people for not letting him behave unethically even more than he already is, and that'd be it.

Beyond that, the obvious astroturfing that occurs on this site (along with reddit, etc.) when it comes to Grok isn't helping. All I hear about Claude, GPT, Gemini, etc., are how terrible they are, yet any discussion of Grok seems to always revolve around sensible, but confident, assertions that it's actually a great product and every new release is the point where Grok finally catches up.

> the obvious astroturfing that occurs on this site (along with reddit, etc.) when it comes to Grok isn't helping. All I hear about Claude, GPT, Gemini, etc., are how terrible they are, yet any discussion of Grok seems to always revolve around sensible, but confident, assertions that it's actually a great product and every new release is the point where Grok finally catches up.

this is the exact opposite of my experiences on HN and Reddit. In my experience, Grok is typically reduced to hitlerbot and CSAM generator and rarely taken as a serious competitor. People let their hatred of Musk blind them to the tech of his companies

Musk is in bed with my authoritarian government, what's China gonna do to me?
So did they distill Mythos in the "Macrohard" data centers? Can Grok hack now and get a free AISI commercial?
I'd let the dust settle rather than trusting benchmarks. But in general a third competitive frontier model would be great.

I still think that it's very possible Gemini gets its act together and becomes the true competitor to the existing frontier models (on more than just cost). But they sure are taking their time with this one, and recent org changes don't exactly signal confidence