22 comments

[ 0.25 ms ] story [ 23.1 ms ] thread
Is the fact that everybody almost catches up with the frontier a sign that we are entering a new region of sigmoid curve?
meta fails at everything yet is frontier on this one
No because the frontier keeps advancing very fast.
Meta has an enormous amount of compute. They are either going use it making and inferencing models or they are going to sell their excess capacity to model providers. Zuck had to completely rebuild his AI team after the Llama 4 launch mess.
Progress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up.

Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether.

In 2024, there was a ton of talk about the plateau. Reasoning was an iteration on chain of thought, but it didn’t really work. Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. That small iteration catches the eye of OpenAI and Anthropic, turns out to be way more important than even DeepSeek could have ever expected when it comes to improving LLMs for coding, and last 18 months have been an exercise on riding that insight to the nth degree.

That one small iteration brought us a lot of progress. Now we’re seemingly exhausting the impact of that one insight, but there may be another soon enough.

> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data.

What was the difference between what deepseek did for R1 and what OpenAI did for o1?

openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.
I don’t know why people think DeepSeek did reasoning models / RLVR before OpenAI, there was a gap of months.
o1 was first, and Anthropic were doing a bit of it; DeepSeek brought it to the masses, but did not invent it.
Totally, RLVR as a concept predates DeepSeek; but they proposed a version that was simple and scalable. Popularizing a specific version of a technique is exactly what I mean by iterations on a theme. It’s only 5% different from what others tried before, but that 5% difference showed a lot more potential than other versions of the same idea.

Since DeepSeeks GRPO, they’ve been improvements as well like AliBabas GSPO that have gotten wide adoption. Again iterations

Yes. It's really up to OpenAI/Anthropic to release a new paradigm to shift the curve now, before everyone catches up entirely.
Even if all the big ideas are gone and we are entering a new part of the curve, there is still an enormous amount of improvement possible. Just iterating on data mix/quality etc, training pipelines, reward functions, specific ways of reasoning (which i guess is mostly just data still) for the next 20 years will yield a looooot. And that's just the models. The harnesses/application layers/whateveritgetscallednext space has 20 years of progress to make.
I think it means that we should be aiming further ahead
They could have just called the article "struggling to remain relevant"
Meta is the last big tech come to AI race, so I would give prop to them for catching up
So was AMD for a while and then consumers kept getting the same repackaged CPU from Intel for years. Competition is great and you should always root for the underdog.
(comment deleted)
Any idea what size this is?
Mark Zuckerberg said that they will release soon Muse Spark as open weights, in which case we will see the size.

However, the statement did not include any details, so it is not clear if the open weights variant will be the same that they are hosting now, or some scaled down version.