174 comments

[ 0.25 ms ] story [ 16.7 ms ] thread
> My two favourite hypothetical questions regarding this used to be:

> If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)

> If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"? My new preferred hypothetical for this is:

> If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?

Maybe just a rumor of a high value target having their API keys accidentally consumed in the context...
LLM can't be trained that easily. More like actual human are checking your logs and stealing valuable things from you.
Or searching anonymised logs for mentions of this problem and using that as part of the context or training.

This would work just as well and have plausible deniability.

Exactly! This is the real Occam's Razor explanation.
They wouldn't appear in weights but could be added to the context. My conversations regularly go "regarding your Java problem"... which was a separate item in the history from earlier. As long as I only see these (and nobody else sees mine), it can be helpful.
What's missing from the story I think is the part about "...then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear...". If only the two of them were working on the problem in secret, how did their breakthrough become a rumor?
Basically no new info here, not really sure why this post needed to be written tbh.
Yes, you have ability to get information very quickly, but not everyone does.
On the contrary, a level-headed summary that gathers information from all the different sources is necessary.
Sounds like a great use case for an LLM
They are busy generating pelicans on bicycles. The summary therefore needs to be written by a human.

(This was a joke. I value Simon's role in the community.)

It's got 'max' and 'ultra' but I don't see a setting for level-headed.
Yeah, this was pretty shitty by OpenAI. Not surprising, sadly.

Them being assholes, trying to exclude an author just because he worked at Anthropic, shows the kind of culture within (that part of) their organization. The focus wasn't on supporting academics or expanding research. It was on getting great marketing.

If they had to burn millions of dollars solving a problem _that they thought was already being solved_ to do so, they'd do it.

The concept that "knowing something has been done" allows others to find the solution to an previously unsolvable problem is an old proven one. For example when Germany launched a rocket (V2 prototype) the British knew its rough trajectory and from spying it's rough size. Although they had previously believed that ballistic missiles weren't possible because no engine could provide the necessary thrust to weight ratio, given the obvious German launching of one, they went through all the known chemical compounds to arrive at the combination (Ethanol and LOX) used. [Story from RV Jones "Most Secret War"].
I’ll go back to the point about authorship. I’m Not a mathematician but I am in academia. if you are fucking around with authorship you are immediately suspect.

That aspect alone would/should be unthinkable to any serious academic. Authorship reflects who did the work and changing it for business competition reasons should be a red flag for multiple different reasons. They include, the sheer tactlessness of treating a major theoretical advancement as a competitive posturing first, the norms of academia second, and all the misunderstandings of the culture of the disciplines culture that people will now suspect are hiding beneath the visible surface (insert topography joke).

Math as a field is fairly unique even in how they list authorship. It was long the norm that authorship to be alphabetical because the idea of first, second, senior etc authorship is harder to define than many other fields.

“The stated rationale for alphabetical order is that it treats co-authorship as intellectually joint work: every listed author’s name carries equal weight, and no one has to negotiate, or be seen to negotiate, over billing. That is a genuine advantage over position-coded conventions, where disputes over who is “first author” are one of the most common sources of authorship conflict in fields that use them” [0]

That norm is changing, slowly, but one option people are pursuing is notable: randomized author order. Their is a perception that alphabetical is too biased…that’s the world OpenAI is stepping into when they make that offer of authorship to one scholar with a demand that he exclude his partner.

I can’t speak to the facts of anything else in this, but if a grad student came to me and said someone made them that offer, I would tell them to run and if they were brave report it.

[0] a to the point lay description of the history of math authorship can be found here: https://casrai.org/guides/mathematics-alphabetical-authorshi...

Yes, there is no doubt about scientific misconduct. I think the entire math community agree on that.
OAI's side of the story is that they discovered the approaches (and indeed solved problems - Euler equations vs NS equations) differed. They then offered Buckmaster lead authorship of OAI's proof, without Alpöge. But they never demanded that Alpöge be stripped of coauthorship on resolving the regularity of the Euler equations. At least that's the claim.

https://xcancel.com/SebastienBubeck/status/20973794116915163...

[delayed]
Neither side has offered hard proof for their interpretation (i.e. malice vs innocence), so shrug. If one is dead set on forming an opinion that's fine but it's all vibes until we know more. Science & math especially is no stranger to bitter disputes over precedence / authorship.
Authorship is weird.

Kind of a tangent, but in some fields authorship is actually more of a formal rank: once you're in the group you appear on every paper. A good fraction of the "authors" won't even know that a paper is being published in their name, and the vast majority haven't read a word of it.

In these fields it can be quite a challenge to not appear on a paper.

My take home from this entire drama is that one should not use LLM services for confidential or proprietary information as they all seem to be run by assholes. And you’re sending them everything you are doing. Would you send your lab notebook to an asshole? Hell no.

I say that as a mathematician (on paper) who perhaps surprisingly doesn’t give a crap about the problem itself.

> as they all seem to be run by assholes.

Wait until you learn how many other SaaS and web 2.0 and cloud based things are also run by assholes.

You know this metaphor? https://en.wikipedia.org/wiki/Turtles_all_the_way_down

But instead of turtles, it's assholes.

But more seriously, no, none of what I wrote above is an attempt to excuse or play down the specific role of assholes in large AI companies.

Of course. I work for assholes. The thing is they’re getting out assholed by several orders of magnitude here.

  > Wait until you learn how many other SaaS and web 2.0 and cloud based things are also run by assholes.
Though it's not the real acronym, Larry Ellison himself has stated that Oracle stands for One Real Asshole Called Larry Ellison. He famously prides himself on it.
Or use Lumo from Proton. Are there any other privacy first companies offering LLMs?
Lumo seems the most reputable. There's confer.to, duck.ai, and nanogpt.
Chutes.ai models are served from a Trusted Execution Environment, so the GPU owners can't see your prompts.
That they do it is just concerning to me in that it says that home-ran models just aren't good enough. Surely researchers like this have the processing power to run them at home, they just don't have the processing power to train models of comparable level.

This is something I feared would happen and where open source would be left behind. Maybe they can do something with crowd-sourcing computational power from volunteirs. They were after all able to get Leela Chess Zero to be comparable to AlphaZero by training from volunteer processing power but it seems to me we live in a world now where the best models keep their stuff closed.

In imagine generation too. I'm not sure how well Stable Diffusion can compete in following instructions with all those advanced models that are kept secret.

> Surely researchers like this have the processing power to run them at home,

Nobody has that power. Certainly not mathematicians.

Yes people have to service the use case of a single person. All sorts of models exist that are designed to be ran at home. Yes, you need a relatively beefy graphics card for it but if your machine can handle the latest video games at good graphics it can handle these local models. How well they compare against these things remains to be seen but apparently OpenAI released gpt-oss-20b which is comparable to o3-mini apparently in terms of reasoning power. There's also oss-120b which does require at least a company server to run but this should well be within the budget of a university.
> one should not use LLM services for confidential or proprietary information

That’s obvious, isn’t it? Just like you wouldn’t upload your confidential documents to an online spellchecker, or your proprietary code to an online compiler?

Or you wouldn't write confidential emails in Outlook and have confidential conversations in Teams. Obviously. Right? Right?
Or if you're going to trust one of them, maybe it shouldn't be OpenAI
LLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence», I wonder if we will see similar examples by LLM’s soon. It seems to me currently impossible that LLM’s can replace human mathematicians, because of their (assumption) likely dependence on human input in the sense of enormous amounts of pre-existing attempts/data.

If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?

Extreme temperature levels (>2.0) can push a GPT of its manifold, essentially producing predictions barely distinguishable from random noise (it flattens the probability distribution of the next token). In theory this could predict anything including the next field of mathematics (infinite monkey theorem) but realistically that would never happen.

However, how to we know the next field of mathematics isn't a novel combinations of several other sub-fields? That level of mathematics would be indistinguishable from magic to most people and so in their eyes the GPT did something truly inventive.

A lot of what humans do is combining old ideas.

And an LLM could in theory also stumble upon entirely new ideas: there's randomness in how they generate their reasoning and answers after all.

I suspect that we are seeing a lot of advances coming from the combination of existing but somewhat obscure knowledge coming from LLMs at the moment, because LLMs are really good at this. At least compared to humans.

Even before our AI friends became good, they were already known for having read approximately every paper and every textbook published in any language. You only need to increase intelligence a fairly small amount from there to get to something like the 'convex hull' of human knowledge.

Compare https://slatestarcodex.com/2016/11/17/the-alzheimer-photo/

The gist is that basically whenever anyone comes up with a new method you get a big burst of activity of picking up all the now lower hanging fruit, that was previously out of reach.

This is also what I've been thinking. The result itself is amazing but it's not like this was completely unexpected. There has been a huge amount of progress on the problem in the last 10 years without which it seems unlikely today's full resolution would have been possible. It is not clear what strategy was taken but it sounds like it borrowed heavily from the two spanish mathematicians. Experts will scrutinize the proof and it will be interesting to see if anything truly original or unexpected was done, outside of known techniques, a move 37.
That doesn't really matter. Currently we're in the "this is a dog" (marks a muffin) stage of "AI as a scientist". It's totally reasonable to expect that considering how rapidly AI is advancing, the frontier of knowledge and research won't be universities anymore, but rather AI companies running their models in a loop.

Imagine that 20 years from now nobody really does science by hand because AI is just better at it, everyone just runs models, but these models require so much memory that only datacenters can realistically handle them, and it just so happens that the public gets access to nerfed models, while privately, companies actually break all asymmetrical encryption ciphers.

If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt a human mathematician would be useful at all on their own.
How have we got to the place we are today then? Someone must have made the first steps onto uncharted territory, otherwise we would be in a homogeneous state frozen in time.

I am not saying that LLM intelligence can not be the same, that they are uncapable of dicovering new fields/problems that they have no training on. I am just pointing out that historically it kind of "must" be true that humans are capalbe of this, but we have yet to see an LLM do something like this, something radically "new" in a sense. All of these breakthroughs appear to me (not a mathematician) to be more a case of "digging" through millions of existing attempts/work, patching it together into a result.

This would already make LLM's one of the greatest tool mankind has ever made, but it has yet to display what I would consider a necessity for human level intelligence, which is this ability to discover entirely "new" things.

Would an LLM, given enough time and only the currently available trainingdata with no further input from humans, be able to solve something that was discovered tomorrow?

For humans my answer would be: maybe, probably, because this has been done historically.

For LLM's I would not be comfortable in claiming that they could. I think they would not be any better at this than traditional computational bruteforce.

What a weird conclusion to come to looking at the world around you
The kind of humans who invent entire fields of science on their own come by a few times a generation. It’s fine to say AI isn’t anywhere as close to them in intelligence, but instead is comparable to the “average” mathematician who is building on the work done by others before them and taking it a bit further.
Yeah, thats likely true to some extent. But my point is perhaps a "counter-thought" to the idea that this is ASI level intelligence and that math is now a computer task and not a human one. If it is incapable of moving maths forwards into the "unknown" it simply can not replace humans, and this might be a hard structural block for LLM's.
All your datum are belong to us!
LLMs can't contribute good code to some of the good OSS math libraries, How is it even solving these problems?
That's a good question.

It is able to contribute code, but maybe not good code.

It's the same in math: it's able to solve problems, but not necessarily in a good way with a human readable code.

Math papers are a lot like software:

- theorems are like API

- lemmata like internal/private function API

- definitions are like types

- the proofs are the implementation

The proofs of ChatGPT are not necessarily readable or maintainable.

You don't need a GOOD proof, just A proof.
No one knows if it's actually LLM doing the heavy weight. It could be just human written brute force algorithm running on their massive computer cluster.
Math problems are often stated in a way that makes it possible to automatically verify if a solution is correct. Which means a loop that speculates an approach (LLM and/or prompts), implements it (LLM), then checks (automated) can work. You still need to have a very good LLM, and probably very good prompts with interesting research directions otherwise you can probably loop forever.
Notably all the major announcements so far are counterexamples or formalizations of existing results to my knowledge. Not necessarily something you can just brute force, but areas with high return on elbow grease.
It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing).

What I cannot reconcile is the timeline and the concern in this specific case.

I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)

I would assume a lot of codex data goes back into training. A well steered session is extremely valuable data.
The breakthrough was on August 15th, but Tristan and Levent had been working towards it (with the help of various models) for the best part of a year.

I personally doubt that their work influenced the OpenAI result - OpenAI themselves say "While unlikely, we cannot rule out that..." - but that "we cannot rule out" is exactly the problem.

If even OpenAI "cannot rule out" the influence of their usage of ChatGPT on this layer result then my discomfort at not understanding how my own usage of ChatGPT affects its training is magnified.

Is there a good tldr on this topic?
This article is the tldr.
I think this drama was blown up a bit out of proportion. The entire discourse I am seeing online seems to revolve around this:

> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models

I mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messages, but on the actual OAI discovery here.

If they have zero-retention, then it is not possible.

So what they are saying is that they don't have zero retention.

A very simple "these two pipelines don't connect up in our architecture, here's our internal high level network diagram combined with our data ingestion opt-out feature flag that we will stand by in court" as opposed to "yeah, we don't even entirely know how our own customer facing systems are connected to our training pipeline, but it probably didn't happen".
Have the other researchers opted out? On all their accounts? Through the entire time? And did they discuss this with anyone else? And did those people ask ChatGPT stuff? And did they disable it? If I was OpenAI, I would be very careful about my wording here when making claims of "we have never trained on any of their ideas directly or indirectly".
I can give you some context. 1. Terence Tao's mastodon explains the way this problem was solved does not in itself contribute much. LLMs (and in this case) produce massive, often unintelligible proofs that do not further understanding. It is often that in pursuit of solving these problems, many other discoveries are made. 2. There is a more serious question about scooping. If OAI is using chat data from researchers to make discoveries, essentially every researcher who chats with an LLM can get scooped. You could be 80% of your way to solving a problem, and LLM could solve the remaining 20%, and get all the credit. Years of your work could be scooped in an instant. If you're a PhD student, this is even worse. Here it's a world famous problem. But imagine you're a PhD student, working on your small but extremely career/progression critical problem, and you get scooped by an AI you talk to. No one is even going to care.
They could explain how they use data from their users "to improve our models".

In the absence of that, how can we make informed decisions about how to use their tools?

(And "you can opt out" is a weak answer IMO, because I can't guarantee that everyone I email has also opted out.)

The claim that they 'cannot rule out' using Buckmaster's data isn't true in principle. Trivially, imagine that the model had been trained on data only since Jan 1, but a discovery was published on Jan 2.

Whether it's true in practice that's a matter of how their systems' functioning & its current state. If they really care about proving this, it's worth an audit. In fact, if anything, they may be able to prove that it hadn't used Buckmaster's inputs, but wouldn't be able to prove that it had.

> ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ...

I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years.

I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days. Competition is a hell of a drug, and frontier LLMs aggressively compound that energy.

It's something that happened before LLMs - multiple discovery. Calculus is a classic example.
Yes but in this case, the allegation is Leibniz literally looked into newtons notebooks
Yep - it's a different case. But the idea that without LLMs, it's unlikely the same idea would be discovered independently is not true - it does happen.
> I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days.

If you read the history of major scientific discoveries, this has been the case for a long time. There are many things that were independently discovered by different people at nearly the same time. Once people know something is solved or solvable, it gets a relentless amount of focus.

Maybe that shows how scientific discoveries come to be. It's not a genius sitting alone in their chamber for a decade and then suddenly they emerge with this huge thing. That's Hollywood fiction. Scientific progress is the colaborative effort of countless researchers over long periods of time, communicating, exchanging ideas, many of them wrong, tweaking, trying, thinking, arguing. When a breakthrough happens then it's the tip of a mountain of work that came before it. If two individuals stand on that mountain and feel there is something somewhere then it's not too strange that they take the last step at roughly the same time because conditions were right. The preconditions were in place at that time, the results required for this were available and the focus was on this specific thing.

I'd say it illustrates well that this last piece, the person celebrated for the achievement, is disproportionally overvalued and the rest of the work they are standing on is disproportionally ignored.

If you believe the totality of the document, there was more shadiness in how OpenAI acted than just timing. Save other things they are accused of, the progression from rumors to replication would attract much less scrutiny. With those in mind timing begins to look suspicious at best.

Would anyone be surprised if major model companies had tagged the accounts of competitor employees for extra tracking? Given the concerns about distillation and bench marking it hardly seems irrational, but how it is used matters quite a lot.

I have a similar story, but perhaps even stranger.

I work for a startup. We often bring a wooden arcade with us to conferences as a marketing gimmick.

The arcade runs a single side-scrolling video game. You're running from a monster and dodging obstacles. The goal is to survive as long as possible, and your result is measured in meters.

There are always a few competitive guys who spend the entire conference taking turns to play it. And every single time, the same thing happens.

Say the current high score is around 200m. Everybody fails somewhere around that number: 190m, 186m... Maybe someone manages 210m. And the high score moves up at a snail's pace.

Then, a new guy shows up and gets something like 500m on his third try. From their next turn on, everybody easily does 450 or more, even though they were struggling to get past 200 just one turn ago.

What makes it stranger is that the game is dead simple. It's not like the new guy discovered a move that unlocked this capability. And it wasn't a lack of motivation either - they'd all been playing for an hour already. They just started performing better after seeing it was possible. There has to be a name for this phenomenon.

I wonder if it's the same with the sub 2 hour marathon which was broken this year (it certainly was with the 4 minute mile).
Reminds me of amateur table tennis. There is a strange dynamic where you down regulate your performance unconsciously when the opponent is playing worse and vice versa.

Could be described as some physical form of this effect: https://en.wikipedia.org/wiki/Asch_conformity_experiments

The term would be conformity / normative social influence.

Surely this has nothing to do with specifically table tennis, nor that it's amateur.

A better example would be mixed boys/girls sports classes in school, where the boys deliberately hold back as to not injure/scare the girls.

It's a pretty obvious and human thing not to go out and completely destroy a much weaker opponent. We're social animals after all.

There also may be an element of energy conservation, there's objectively no need to put in any more effort than necessary. Inefficient.

I think the difference is that, in the table tennis example, it happens unconsciously and you can't help it. But you're right, it likely happens in most sports.

The mixed boys/girls classes example is, as you said, deliberate. And I too remember holding back on purpose in such situations when I was a kid.

With the arcade, no one was deliberately holding back. I'm sure of it.

Is it that you’re holding back against weaker opponents, or that you’re more excited and engaged when you’re playing against someone that forces you closer to the edge of your ability?
I am pretty sure if you play at pro level you train this out of yourself. But then, the marathon / running examples where certain milestones get broken by one person and then suddenly by a host of others, is a counterargument.
I like to leverage that when brainstorming solutions to hard problems. Instead of contemplating small percentage improvements, try to think about what's in the way of improvements that are orders of magnitude better (e.g., don't take time to run a big task from 100s to 90s, take it to milliseconds). Sometimes it unlocks big ideas.
I work in software deployment for large orgs. I use this technique to optimize. "How can I provision or deploy this with one package install, and launch instance." Typically take multi-page or multi-step deployments down to fully automated.
Achieving a high score requires focus and effort.

There's little reason to aim for a score a lot higher than the current best; and when getting close to the score you're aiming for, it's easy to get agitated and make a mistake.

And if you do beat the high score, you're likely to loosen your attention right after that, and it's even annoying to keep going for much more.

I think it's largely due to psychology. If 210m is considered the best, then as they approach it they may start to tense up and choke. When the goal and possibility is known as 500m, then there's no point being concerned near 210m.
This is like the mile record. Everytime someone breaks the existing one a bunch of people surpass the old record shortly after.
Climbers know this phenomenon as the "send train". A group of climbers has been working on a problem for a while. When the first person sends it, it often happens that multiple people are successful in the immediate attempts that follow. Part of it is watching the successful beta, but it happens even when everyone knows the moves. Once you've seen it's possible, you stop climbing tentatively and commit to the hard move instead of hedging for a fall.
That's the other thing. Even if it somehow magically turns out that they didn't plagiarize Buckmaster and Alpöge, that an independent audit goes through all their processes and finds that they could not possibly have stolen anything, and ignoring their outrageous attempts to force a coauthor off the author list, and ignoring that their employees act like small children on social media, this is a company whose culture is “we heard of a breakthrough being possible, let's scoop it”.

Normally, when you tell a coworker that you're wrapping up a result, unless they're some kind of sociopath, their natural inclination would not be to try to steal it from you.

Not the hope, but telling them that an answer does exist. This ofcourse means now we are going to go into an even more darker cave next time we are looking for some gold and pathbreaking discoveries will become even more rare. Add to that, the fear of people not winning against ai and having fewer rewards, then fewer people even enter those fields or attempt problems over the next generation. AI erodes skills not at individual but at civilisational level.
You see this in many things. Look at climbing for instance, or running. Once the first v17 had been done, suddenly many people did them. Or running with a sub 2 hour marathon.
I run a small SaaS[1], like so many others, that uses AI to generate and optimize SQL. Getting this to perform optimally has been a lot of work and now I wonder if OpenAI is outright stealing this knowledge, which without a doubt is highly valuable to them.

[1]: https://www.sqlai.ai

SQL is so ubiquitous and the use case so obvious, there's no way they have not already been tracking performance and benchmaxxing on SQL queries for years.

But I don't think openai will bother to release a competitor, the real threat is that anyone with a decent LLM and a harness to try a few queries will land at the same or a better query within minutes.

It's extremely unlikely that there is anything interesting or novel in the optimisation of a small SaaS SQL
But what if they added "be interesting and novel" to their prompt? Did you consider that?
I did not.

Having spent considerable amounts of time undoing the effects of overconfident junior programmers who decided to be interesting and novel on small business code bases, I guess if we're going to replace juniors with AI we may as well ask for the full experience.

No one wants to “steal” your vibe coded crap. Your prompts aren’t as unique or valuable as you think.
It could be much worse.

OpenAI can easily identify these outstanding human behind their accounts. Human in OpenAI constantly check their logs for breakthrough. When they find something interested, they brute force the result using their massive computing power.

No LLM is even needed.

Or, using anonymised data, search for anyone seriously trying to tackle this problem - probably about 10 people in the entire world and use their ideas as a starting point. They wouldn’t even need to be watching specific accounts or using de-anonymised data if they know what they’re looking for.
Maybe someone on the NYU team forgot to opt out of “improve the model for everyone”.
>While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?

But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."

I’m not a mathematician so take this with a grain of salt. Apparently terry tao commented that the approach used for the Euler paper can “probably” be used for solving NS but it’s still technically challenging and can probably be done with an LLM with a lot of compute. To me the crux of the issue is whether the insight were stolen so that the problem becomes something that is in the domain of LLMs. This is much different than LLMs coming up with the insight. OpenAI wants everyone to think the LLM came up with the insight and solved the thing by itself even though they have perhaps an army of researchers.
Having the chat logs enter the training data and having them have a meaningful influence on the ultimate result the model produces are very different things.

The text for all the Goosebumps books are certainly in the training data and to some small amount influenced the solve. But their contribution was so vanishingly small it would seem absurd to say R L Stein should have recourse for contibuting to the solve.

But this is different, right?

The equivalent would be taking a (fully offline) LLM and asking it about the ending of one specific Goosebumps book, and it revealing the twist. And although that specific book was (probably) only once in the training data, a high parameter LLM can usually "remember" the twist.

The only way it would be able to tell you the ending is if it was somehow given more importance in pretraining, loaded into context, or represented in multiple sets of training samples. I have a blog that I make very LLM friendly and usually load posts up into context when I’m working on something relevant. I’ve also opted to improve models for everyone. Despite this, the model can’t recognize my site or any of my posts when I ask it to recall without internet usage (I also turn memory off btw).
That's not correct for a SOTA model with trillions of parameters. Those have immense amounts of knowledge trained into their parameters. Try it. It "remembers" the ending of random books, Goosebumps and otherwise.

That's the entire point of why knowledge cutoff is so important (if you use them offline).

I find it entirely believable that the Navier-Stokes conversation was auto-flagged as high value training data, and burned-in the models knowlegdge base.

Time to repeat the same angry mob style discussion again. Great job to the mods.
If they wanted to solve a Millenium prize problem so much, why did they not try to solve P-NP instead? It boggles the mind.
What makes you think they didn't try?
Are you kidding/trolling? The P=NP problem is FAR more fundamental, and if proven true, would basically be a proof that e.g. public key crypto can be broken (NOT a description of how to though).

Basically, it would be a proof that all the REALLY hard (combinatorial) problems out there, have a much simpler solution, if we were able to find it.

> The P=NP problem is FAR more fundamental,

Yes.

> and if proven true, would basically be a proof that e.g. public key crypto can be broken (NOT a description of how to though).

Well, only if the answer is that P=NP.

> I believe P vs NP is a different beast entirely. We don’t even know which way the answer should go.

Most people expect P < NP, and then crypto wouldn't be broken. P=NP would also break pseudo-random number generators, for pretty much the same reason as the rest of crypto.

However Don Knuth is one example of an expert who thinks P = NP is plausible.

Btw, we do know quite a lot about how a proof of P vs NP will _not_ look like. That is we are in the curious situation where we can prove that certain proof techniques won't work on this problem.

Weirdly enough, we already have the optimal algorithm, we just can't prove its runtime. Ie we have an algorithm that runs in polynomial time on all NP hard problems, if P=NP. (But the constant factors are crazy.)

I always thought it was enough to switch off the "Improve the model for everyone" setting on chatgpt.com:

"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy."

But apparently there is also an entire completely different route "Do not train on my data"?

Does this mean that before I submitted the "Do not train on my data" request, my data was used for training in spite of "Improve the model for everyone" being turned off?

We are getting to facebook/meta-levels of privacy settings obfuscation.

That's because what people enter into LLMs is the last gold there is out there. Everything else is already scraped or ensloppified.

Maybe next step is to filter your input client side through an unknown number of obfuscators where you ask LLMs to rephrase your question (onion router idea) such that no single provider can be certain that this is human input and not some slop feedback loop.

The last time I checked there was a loophole — if you provide feedback in-session (responding to “how are we doing” or “which prompt is better”) then they can use that feedback + relevant context. Relevant context for chatgpt might include memories / other sessions. That may not be the only loophole.

That in itself is a dark, dark pattern. There should at the very least be explicit warnings for users who have checked “do not train”; or they should not be presented with such dialogs.

I mean this is just the logical end game. The company/ies that control the uber mind will devour ALL useful or valuable work. Unlimited intelligence at unlimited scale means the value of humans for knowledge work goes to zero. We will eventually not have meaningful access to the uberminds because we will be pointless. At which point our Silicon Valley luminaries will really have no choice but to extinguish as many of us as possible for the greater good. After all if we are all pointless then that has to be weighed against are cost to Mother Earth. Clearly the only moral solution is to cull or allow to be culled some 98% or so down to a more sustainable, manageable population of curiosities. The end game of ai is the remnants of humanity in a zoo, and that’s if humans control the outcome… the machine minds might be more charitable as they would be less afraid…
Fortunately, there's plenty of competition between AI labs. Not all of them are even in Silicon Valley.