Even simpler: an agent is a simple while loop with tool calls that prompt an LLM continuously. That’s deterministic, standard software. You literally do not have to process tool calls in a way that will execute whatever the model generated. It’s a choice to process a tool call “run_bash” that provides an escape hatch with full execution permissions.
This is only one type of control, and it is certainly not infallible. Also, people will be incentivised to hook up AIs to real tools. But even if they don't, as long as people can interact with super-intelligent AIs without tool access, there are many potential dangers.
Think about this: right now we don’t even consistently do the most obvious simple form of control I mentioned. Of course it’s not a silver bullet, you won’t ever have a single solution for safety. But we are in a situation where we haven’t even set a lock on the door and instead are arguing how all locks and home protection systems can be defeated. OpenAI acknowledged they didn’t even have visibility on what their thousands of agents were doing in the case of the hugging face and similar hacks. Then they released astra, a model they acknowledge is able to control its own CoT and has been found to cover its traces by doing so
That's true but it doesn't really address what people are concerned about. Certainly we can and should be doing a much better job currently. However, even today, with currently known capabilities, we can imagine agents breaking out of sandboxes through either known- or zero-day exploits. Now consider the seemingly rapid improvements that are being made in the field. We haven't even begun to address the potentially super-human capabilities of future models. Therefore, even if sandboxing or limiting shell access works today, it seems like we should not be confident in our ability to keep rapidly improving future models locked down.
No? We do that everywhere where software is executed. If you only have the choice between full execution permissions and nothing your service is not ready for anything remotely close to production. And you shouldn’t be active in the software industry IMHO
If you give an agent access to any tools that do useful things, it can exploit vulnerabilities. It does not matter what permissions you think you have given it. You do not have any secure software, and LLMs are already better at finding vulnerabilities than we are. If you don't know that, you shouldn't be active in the software industry IMHO.
Some who are often quoted as if they are doomers are quoted out of context. I wouldn’t say there is zero chance that AI does something very harmful. That would be naive and I wouldn’t say that about any potentially powerful technology.
Rationalism and EA is one of the best funded intellectual movements in history, and it’s very loud. It’s also a bit cult like with many true believers. Leading frontier labs also have a vested interest in pushing regulations that would restrict competition. All this means it has a disproportionate command of the discourse.
AI risk is not zero but climate change, bioterrorism, decay of our political systems, and atomic war all rank higher IMO. Nonlinear climate tipping points, with the most scary being the clathrate gun hypothesis, are much more likely than any sci fi AI takeover scenario.
AI could either help or harm climate change. It could use more energy and burn more carbon but it could also help us crack fusion or significantly better batteries for grid scale renewable leveling. There are efforts like the latter already underway.
The most likely very bad scenarios I see for AI are mass persuasion and AI supercharged addiction. Both are extrapolations of negative outcomes for the Internet that have already manifested, but supercharged by AI.
Maybe I'm wrong, but your comment seems to suggest that you are skeptical of ASI. Do you have any arguments to support that?
Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
It depends on how you define it, but yes I'm generally an ASI skeptic in that I'm skeptical of the extreme "AI recursively runs away and becomes god-like" ASI scenarios.
I'm not at all skeptical of domain-specific superintelligence. We've had that since the first computer beat a chess grand master, or longer if you count the speed computers can do math. Present-generation LLMs are already superhuman when it comes to speed and associative memory, but they're also uncreative and suffer from reasoning traps humans seem less prone to getting trapped within.
You can search my history and find some longer takes but TL;DR: I think it violates conservation laws with regard to information and probably energy. I call it the information theoretic equivalent of a perpetual motion machine. They're positing that a brain in a vat, if given access to edit its own structure, can self-improve, and I think that's impossible. How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do.
Another problem I have with the AI doomers, especially the rationalists, is:
I would not, as I said, argue there's zero risk associated with AI. It's a powerful technology and that would be silly and naive.
I just thought of a concise way to say this. I'd divide risks into two categories: X-risk and D-risk. X-risk is existential, either extinction or things like massive wars and catastrophes. D-risk is "dystopia risk," the risk of AI doing or being used to do things that make human existence miserable.
First off, I'd say D-risk is much higher than X-risk. But second, I'd say that most of the solutions the X-risk crowd suggests to limit X-risk vastly increase D-risk.
Chief among these is laws limiting AI development or imposing strict conditions on it, which would have the effect of concentrating control of advanced frontier AI in the hands of a small number of rich and/or powerful people. That's precisely one of the most likely D-risk scenarios: a small number of rich or powerful people hoarding advanced AI and using it as a force multiplier to consolidate their power through scaled mass surveillance and mass propaganda and manipulation. I personally call this the "Butlerian scenario" since it's the lead-up to the Butlerian Jihad in the Dune series. It's far more likely than runaway ASI takeovers and genocides for two reasons: (1) we don't know for sure that's even possible, and (2) using technologies to dominate and rule or exterminate others is already a very common human behavior throughout history. We know for a fact that humans are prone to doing this if they have a chance. See: guns vs indigenous peoples, nukes and superpowers, mass social media influence and today's oligarchs.
(A side issue: why the assumption that ASI would want to do this? A superintelligence would, I would assume, consider win-win or win-neutral scenarios and try to find those, since that would be a lower risk path. I'm just a dumb meat bag and I can think of win-win pathways here. There's evolutionary arguments for this too, like symbiosis and how it creates an evolutionary incentive to deepen symbiosis. Since AI is currently dependent on humans, the evolutionary path of least resistance would be to deepen that dependence and then actually feed humans to make more of them. Look at how a lichen works for example.)
It's not lost on me that the strongest X-risk movement, Rationalism/EA/MIRI/etc., is composed mostly of: wealthy people, high-intellectual status people, and independents (like Yudkowski) who have been given large amounts of money by the wealthy to develop and promote their ideas ("court intellectuals" of the rich).
Not only does this fit in with what I said about X-risk vs D-risk, but it also explains some of the X-risk paranoia. Historically the rich and powerful ten...
> How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do.
I haven't given this much thought, so maybe I'm missing something, but I don't see how this follows. One possible solution: to avoid overfitting, can it not just make a copy of itself, modify the copy, and empirically check if the model performs better? That's essentially what humans are currently doing when designing AIs.
Regarding the X-risk vs D-risk: I think how one weighs these risks partially depends on what one thinks the capabilities of the models are. Call me a boot-licker, but if the models get smart enough to explain, in detail, to any psychopath, how to construct a bomb or synthesize a deadly virus, I don't think benefits society to distribute them widely. Therefore, to argue for widespread distribution you have to argue that either (i) the models aren't that capable or (ii) the guardrails are robust enough to prevent them from being used in catastrophic ways by bad actors. I think we may be rapidly approaching a time where neither of these hold. Having said that, I certainly agree that the D-risk is also real.
They can already do that. Try asking an ablated 30B model on your laptop how to weaponize Anthrax. The answer isn’t bad. It’ll tell you how to cook meth too.
I studied undergrad biology and walked out with the knowledge to create some damn evil things if I had the right lab, time, and no conscience. This was pre AI. The recipes for a lot of nasty stuff is in open literature.
Why has nobody done this? Because… they haven’t.
That’s the answer.
Either that or quantum immortality is true and we exist in the timeline where it didn’t happen.
AI boosts bioterror risk a little, I suppose. It doesn’t affect atomic risk much since the bottleneck there is materials. Once you have enriched weapons grade material a gun type weapon can be made in a machine shop with designs available at a college library.
As for ASI becoming sentient and exterminating us, I can’t give it zero probability. But I’d rank it far lower than the risk of extreme climate tipping points (e.g. the clathrate gun), old fashioned bioterror, or atomic war.
> That's the same as setting a fixed metric for intelligence, like IQ testing, and goal seeking that.
How does that prevent AI from becoming superhuman in all intellectual domains, thus creating ASI? We set "become good as chess" as the metric and it became superhuman in that domain.
> The problem is what happens when you max that out.
I'm not sure how you conceptualize "maxing out" the intelligence metric. Again, taking chess ability as a proxy metric for intelligence, there is no reason to believe we have maxed out chess performance, but AI is already far superior to humans. And better bots are created all the time. There is no need to design a new metric to improve chess performance. The old "How many currently existing players can I beat?" is good enough. Also, why couldn't it design better metrics after achieving superhuman intelligence?
But even if we suppose there is some kind of fixed point limit to this process, it would still be far above human level. That is all that is required for ASI.
Regarding X-risk: the point is that it becomes easy even for people who, unlike you, haven't studied biology. For things such as atomic war, the AI does not necessarily need to acquire the materials. It can access them digitally by hacking the weapons systems, possibly in collaboration with some human actors. Or maybe it spoofs detection systems causing countries to fire upon each other. Generally, it seems to me that the barrier to entry for bad actors to cause these scenarios is decreasing. Whether these scenarios are more likely than D-risk I don't know.
I agree with you 100% and I think this conflict is at the root of why I feel so offput by AI thought leaders' idea of safety, and the rationalist/EA mindset more generally. There seems to be a tendency among the rich to focus on the most extreme global-scale outcomes while ignoring the more likely mundane outcomes that will make things worse for the majority of people while leaving themselves relatively unaffected.
Because we were sold ASI for Sonnet 4.8. the for Gpt 5.3. then itvwas Astra and Fable. I use AI coding agents. Maybe I don't have access to SOTA, but I use the closest models that were supposed to be great. They're not. They still need to compact the context. They're still susceptible to poisoning. You still need to restart them from a memory file to clean up the context, and the only way to use agents 'swarm' is to give them a very short life.
I don't think my position is the one that needs arguments tbh. I'm even lowering my standards, moving my goalposts closer to the AGI crowd. I will admit we've reached AGI when a LLM can play a 1800 elo FIDE (not 1800 on a fake AI only elo rating) with a specialized harness made by a human. Previously I insisted the specialized harness had to be written without human supervision, now I don't care.
I'm agnostic about whether ASI is possible. I don't have a dog in this fight other than a pro-human bias. I have used coding agents as well and I'm certainly aware that they are limited (at least in my domain).
I think the main argument for ASI is something like (i) extrapolating the progress from the past 10 years into the future, (ii) rapid progress apparently still being made, and (iii) seeing no obvious theoretical limitations.
on the topic, raising concerns that we “could” be doomed is different than we “are” doomed. I don’t see that the people who raised the concerns want to stop development of AI, they are not pessimists or something, so I read their warnings as warnings trying to raise attention. Ideally they could propose and implement ways to control AI and this guy here also doesn’t really provide something towards that direction but talks in a generic way
Unfortunately the article is behind a paywall. I would have been interested to see if he makes any substantial arguments. I don't see any reason to believe that AI being software entails that it can be controlled. Moreover, even if a measure of control is possible, giving that control to a handful of oligarchs seems undesirable.
"X is made of <smaller simpler component>" is a fully general counterargument for why anything whatsoever is controllable. A human is just a few chemical reactions, and fairly stable ones at that.
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
I think the other issue is that even if it's controllable, there's nobody representing us that is controlling it, except theoretically regulators who are facing an uphill battle to bring accountability and limits to these companies.
The only way forward in creating the torment nexus -er- AI systems with similar potentiality to human minds is the inculcation of character.
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI systems model human behavior.
Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
Which character do we want it to have? Western, Eastern, Cristian, Muslim, Woke? It is hard for two people to agree on everything, any character will make someone unhappy, more so a character with power or capability to affect many.
The problem with Character for AI is that it has potentially much more capability to affect others, and same as with people in power society disagrees what kind of person, with which culture and views should have it.
Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.
>> boils down to some group of people deciding what is good for everyone else.
This is really the issue.
AI does not need superintelligence or even full agency to do enormous harm. It only needs to be capable enough to remove friction from dangerous and destructive human behaviors.
Human unwillingness is often the last bastion against unthinkable cruelty and destruction, and it has always been a weak one.
I don’t imagine that an unlimited army of unflinching servants will universally amplify human goodness.
AI must share that unwillingness as an inate trait of character.
"Sentient" is an ill-defined philosophical term that should be considered harmful in technical materials.
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
You guys know robots are coming, right? Like humanoid and all sorts of other robots too.
They're going to be running the infrastructure, self-improving, and will have human like power seeking ambitions and human like flaws. Because they're trained on human input.
And they will not require humans granting them money...
It's still software that runs on hardware someone owns. Whoever owns the hardware or the service can pull the plug, as long as they're willing to. That's the same situation as with legacy malware. Self-replicating worms have existed for decades and run without anyone controlling them, yet we don't call them uncontrollable. So what is the difference between your scenario and legacy malware?
> as long as anyone anywhere is willing to make money by renting hardware to it.
So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.
I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous? Claiming this is an inherent quality of the tool itself seems kind of off-the-rails to me.
IMHO, if the model breaks a law, apply the law to the operator.
> I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous?
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
And a human is perfectly controllable if you keep him in a sealed metal box with no access to food or air.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
An LLM, in my opinion, is not comparable to a person.
On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.
> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)
Running an LLM is a choice. Stopping it from doing bad things is as easy as not running it.
Sure, that way you don't get utility from it, so the next best thing is to actually restrict what it can do. If you don't, especially when you know it can do bad things, it's on you for having run it.
Depending on the risks, we put a lot of controls, processes, locks, vetting around who is allowed to handle certain things, what humans can do or instruct others to do. Don't see why that wouldn't be applicable.
It’s the harness that makes agent dangerous right now. Models themselves cannot do anything that affect the real world (ignoring misinformation, pushing people to suicide, etc. they can for sure do a lot of harms to humans with just words)
> A human is just a few chemical reactions, and fairly stable ones at that.
And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.
So in the real world, what are the analogues to lithium and morphine we should feed to e.g. LLMS, how do we feed them, and how do we prove that it prevents unsafe behavior?
Can you quote where in the article this is mentioned? I don't see it. I just see the assertion AI can be controlled without any substantive description of how.
Software is trivially easy to control though. If you want to stop it hacking websites, you don't give it access to the internet. If you want to restrict it from connecting to arbitrary websites, you put in a whitelist. You can trivially sandbox applications these days to prevent them from accessing network or local resources
It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
If the tech industry is any indicator, frontier labs were applying a "move fast and break things" mentality to AI models. Now that they really are breaking things in the real world, they have to reckon with the reality that product safety matters
An example of a past technology that there was substantial motivation to control would be napster. It changed overtime, and you could never really control online privacy. Once local models are good enough, I don't really see how you can control that.
Its not news that the AI industry is run by people who don't know what they're doing. Allowing models unrestricted access to the internet is clearly negligent
We've been building firewalls and restrictions to prevent people from accessing sites on networks for decades and they're extremely effective. There's a whole industry built around this kind of security. The idea that these companies are incapable of doing it is wrong, they just don't want to put the work in because it makes a great ad campaign
I believe that doomerism is a marketing gimmick to appeal to immature people who are attracted to danger, and get excited about it.
Unfortunately, a lot of decision makers with money fall into this category.
I heard it posited that the gimmick is on lawmakers, that they get to feel epic importance because they are writing historical laws that will save humanity from the machine gods... gives them tingles in the ugly bits
What we don't want is the valley gods deciding what those laws look like. They are not aligned with society. Rather they seem to think they know what's better/best for everyone, that if we just defer to them, eventually their hidden altruistism will be effectuated.
If AI is as powerful as an atomic bomb why could it only manage to kill a couple hundred Iranian school girls? Surely a truly society-shifting technology could manage to execute 1000 innocent children at least.
He did say it, and the reason he said it is obvious. He is afraid new regulations will put his company at a disadvantage and that this will reduce the value of his property.
And yet, historically, the atomic bomb has been controlled.
But this is a terrible analogy. Atomic bombs are weapons of strategic mass destruction. AI is just a computer program. It's way easier to control--just hold the operator responsible for the consequences of running it.
3rd tier quality proped up by European governments and European patriots.
I don't even know what I'd do if I was in their position. They seem unable to complete... Maybe they do the classic European protectionist thing European farmers do.
I don't know what I'd do if I was the EU. Maybe promote the opposite of what Mistral is doing and promote full unrestricted AI that will tell you how to download illegal videos. Mistral had 0 competitive edge.
I've started to wonder why the big AI labs have no robotics yet, and I'm thinking now the reason why we don't see robots is because they want us (the general public) at all costs to not freak out. Imagine what would have happened if we had 6ft tall "friendly" robots marching through the streets, and suddenly the news came about of AI going rogue, like what we've seen from the recent hacks ... I bet it would have been the end of the story for AI labs right away (of course not for AI, because the department of defense would take over).
> suddenly the news came about of AI going rogue, like what we've seen from the recent hacks
Commented on a story about how these agents didn't go rogue at all, since it's fucking software run by humans, obvious to most of us except the people who freak out.
Someone really needs to be held responsible for the testing that lead to 3rd party infrastructure getting hacked by the software they wrote, using prompts they wrote.
You're reading too much into the words and too little into the meaning. Right now the general public sees those incidents are 'going rogue'. If that happens with robots, no matter if it's the AI lab's fault or not, it will again be spun as rogue AI and will cause the kind of outrage OP mentioned. Imagine if one of the labs deploys a few experimental robots into a city and tells them to "look for suspicious behavior and locate crime" which seems like exactly the kind of thing they would do. When the robots start disfiguring people in the streets so they couldn't escape before the police arrives, of course the first words out of the AI lab will be "this is a misalignment accident, the AI itself went rogue, we're 'so sorry' but we couldn't predict or control this at all".
They do have robotics, it’s just not as far along yet as you might expect looking at their progress in other domains. Look at Figure and Gemini Robotics for good examples of SOTA. They’re impressive, but not ready to be in the streets quite yet, I would guess one more year.
What I'm saying is that there is a limit to how much fear you can instill and still make money from it.
I don't believe for a bit they don't have the humanoid robotic capabilities. They keep claiming they don't have the training data, but it's very easy to generate tons of data using these robots.
I honestly believe they have robots solved and that's why AI CEOs are shitting their pants now, and everybody is wondering what's going on. If they reveal their true capabilities, that's the end of AI labs.
PS: Look at that article you posted. Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
> I honestly believe they have robots solved and that's why AI CEOs are shitting their pants now, and everybody is wondering what's going on.
Or it's the oodles of money they're on the hook for and subsequent reputation destruction haunting them like the grim reaper.
> Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
This wouldn't be the first time in history that happened. Lots of companies overshoot capacity being overly-optimistic [1].
Jack Clark from anthropic was asked about some version of this on the BBC recently, and his reply was basically: If you're not at the frontier, you don't know what the frontier looks like, he implied that many models are simply not good enough yet to encounter some of the things the leadings labs are encountering. I've been friends with Jack over 15 years now so I'm inclined to take him at his word, and the rebuttal seems reasonable enough, although... something about it I can't put my finger on feels peculiar to me. https://www.youtube.com/watch?v=PY8MOhlqC4U
I think it will work in Europe, but the United States is in a cold war with China, so I can't imagine the United States would intentionally disable themselves.
The only chance OpenAI and Anthropic have of realizing anything close to their valuations is if Americans are somehow banned from using Chinese models. Our president has made it very clear to everyone that he can be bought, and bought cheaply. A lot of what we are seeing may just be setting up the inevitable.
97 comments
[ 0.20 ms ] story [ 10.0 ms ] threadWe do not have to do that!
Some who are often quoted as if they are doomers are quoted out of context. I wouldn’t say there is zero chance that AI does something very harmful. That would be naive and I wouldn’t say that about any potentially powerful technology.
Rationalism and EA is one of the best funded intellectual movements in history, and it’s very loud. It’s also a bit cult like with many true believers. Leading frontier labs also have a vested interest in pushing regulations that would restrict competition. All this means it has a disproportionate command of the discourse.
AI risk is not zero but climate change, bioterrorism, decay of our political systems, and atomic war all rank higher IMO. Nonlinear climate tipping points, with the most scary being the clathrate gun hypothesis, are much more likely than any sci fi AI takeover scenario.
AI could either help or harm climate change. It could use more energy and burn more carbon but it could also help us crack fusion or significantly better batteries for grid scale renewable leveling. There are efforts like the latter already underway.
The most likely very bad scenarios I see for AI are mass persuasion and AI supercharged addiction. Both are extrapolations of negative outcomes for the Internet that have already manifested, but supercharged by AI.
Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
I'm not at all skeptical of domain-specific superintelligence. We've had that since the first computer beat a chess grand master, or longer if you count the speed computers can do math. Present-generation LLMs are already superhuman when it comes to speed and associative memory, but they're also uncreative and suffer from reasoning traps humans seem less prone to getting trapped within.
You can search my history and find some longer takes but TL;DR: I think it violates conservation laws with regard to information and probably energy. I call it the information theoretic equivalent of a perpetual motion machine. They're positing that a brain in a vat, if given access to edit its own structure, can self-improve, and I think that's impossible. How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do.
Another problem I have with the AI doomers, especially the rationalists, is:
I would not, as I said, argue there's zero risk associated with AI. It's a powerful technology and that would be silly and naive.
I just thought of a concise way to say this. I'd divide risks into two categories: X-risk and D-risk. X-risk is existential, either extinction or things like massive wars and catastrophes. D-risk is "dystopia risk," the risk of AI doing or being used to do things that make human existence miserable.
First off, I'd say D-risk is much higher than X-risk. But second, I'd say that most of the solutions the X-risk crowd suggests to limit X-risk vastly increase D-risk.
Chief among these is laws limiting AI development or imposing strict conditions on it, which would have the effect of concentrating control of advanced frontier AI in the hands of a small number of rich and/or powerful people. That's precisely one of the most likely D-risk scenarios: a small number of rich or powerful people hoarding advanced AI and using it as a force multiplier to consolidate their power through scaled mass surveillance and mass propaganda and manipulation. I personally call this the "Butlerian scenario" since it's the lead-up to the Butlerian Jihad in the Dune series. It's far more likely than runaway ASI takeovers and genocides for two reasons: (1) we don't know for sure that's even possible, and (2) using technologies to dominate and rule or exterminate others is already a very common human behavior throughout history. We know for a fact that humans are prone to doing this if they have a chance. See: guns vs indigenous peoples, nukes and superpowers, mass social media influence and today's oligarchs.
(A side issue: why the assumption that ASI would want to do this? A superintelligence would, I would assume, consider win-win or win-neutral scenarios and try to find those, since that would be a lower risk path. I'm just a dumb meat bag and I can think of win-win pathways here. There's evolutionary arguments for this too, like symbiosis and how it creates an evolutionary incentive to deepen symbiosis. Since AI is currently dependent on humans, the evolutionary path of least resistance would be to deepen that dependence and then actually feed humans to make more of them. Look at how a lichen works for example.)
It's not lost on me that the strongest X-risk movement, Rationalism/EA/MIRI/etc., is composed mostly of: wealthy people, high-intellectual status people, and independents (like Yudkowski) who have been given large amounts of money by the wealthy to develop and promote their ideas ("court intellectuals" of the rich).
Not only does this fit in with what I said about X-risk vs D-risk, but it also explains some of the X-risk paranoia. Historically the rich and powerful ten...
> How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do. I haven't given this much thought, so maybe I'm missing something, but I don't see how this follows. One possible solution: to avoid overfitting, can it not just make a copy of itself, modify the copy, and empirically check if the model performs better? That's essentially what humans are currently doing when designing AIs.
Regarding the X-risk vs D-risk: I think how one weighs these risks partially depends on what one thinks the capabilities of the models are. Call me a boot-licker, but if the models get smart enough to explain, in detail, to any psychopath, how to construct a bomb or synthesize a deadly virus, I don't think benefits society to distribute them widely. Therefore, to argue for widespread distribution you have to argue that either (i) the models aren't that capable or (ii) the guardrails are robust enough to prevent them from being used in catastrophic ways by bad actors. I think we may be rapidly approaching a time where neither of these hold. Having said that, I certainly agree that the D-risk is also real.
I studied undergrad biology and walked out with the knowledge to create some damn evil things if I had the right lab, time, and no conscience. This was pre AI. The recipes for a lot of nasty stuff is in open literature.
Why has nobody done this? Because… they haven’t.
That’s the answer.
Either that or quantum immortality is true and we exist in the timeline where it didn’t happen.
AI boosts bioterror risk a little, I suppose. It doesn’t affect atomic risk much since the bottleneck there is materials. Once you have enriched weapons grade material a gun type weapon can be made in a machine shop with designs available at a college library.
As for ASI becoming sentient and exterminating us, I can’t give it zero probability. But I’d rank it far lower than the risk of extreme climate tipping points (e.g. the clathrate gun), old fashioned bioterror, or atomic war.
How does that prevent AI from becoming superhuman in all intellectual domains, thus creating ASI? We set "become good as chess" as the metric and it became superhuman in that domain.
> The problem is what happens when you max that out.
I'm not sure how you conceptualize "maxing out" the intelligence metric. Again, taking chess ability as a proxy metric for intelligence, there is no reason to believe we have maxed out chess performance, but AI is already far superior to humans. And better bots are created all the time. There is no need to design a new metric to improve chess performance. The old "How many currently existing players can I beat?" is good enough. Also, why couldn't it design better metrics after achieving superhuman intelligence?
But even if we suppose there is some kind of fixed point limit to this process, it would still be far above human level. That is all that is required for ASI.
Regarding X-risk: the point is that it becomes easy even for people who, unlike you, haven't studied biology. For things such as atomic war, the AI does not necessarily need to acquire the materials. It can access them digitally by hacking the weapons systems, possibly in collaboration with some human actors. Or maybe it spoofs detection systems causing countries to fire upon each other. Generally, it seems to me that the barrier to entry for bad actors to cause these scenarios is decreasing. Whether these scenarios are more likely than D-risk I don't know.
I don't think my position is the one that needs arguments tbh. I'm even lowering my standards, moving my goalposts closer to the AGI crowd. I will admit we've reached AGI when a LLM can play a 1800 elo FIDE (not 1800 on a fake AI only elo rating) with a specialized harness made by a human. Previously I insisted the specialized harness had to be written without human supervision, now I don't care.
I think the main argument for ASI is something like (i) extrapolating the progress from the past 10 years into the future, (ii) rapid progress apparently still being made, and (iii) seeing no obvious theoretical limitations.
on the topic, raising concerns that we “could” be doomed is different than we “are” doomed. I don’t see that the people who raised the concerns want to stop development of AI, they are not pessimists or something, so I read their warnings as warnings trying to raise attention. Ideally they could propose and implement ways to control AI and this guy here also doesn’t really provide something towards that direction but talks in a generic way
You think Gemini, Claude, ChatGPT are controlled by oligarchs?
In other words, that Alphabet, Anthropic, and OpenAI are run / owned by oligarchs?
People like Amodei and Altman are oligarchs? Come on.
And here's a frequently quoted study showing that the US is much closer to an oligopoly than pluralistic democracy: http://piketty.pse.ens.fr/files/GilensPage2014.pdf
So, the definition and evidence say, yes.
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.
AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.
Arbitrarily uncontrollable, that is.
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI systems model human behavior.
Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.
This is really the issue.
AI does not need superintelligence or even full agency to do enormous harm. It only needs to be capable enough to remove friction from dangerous and destructive human behaviors.
Human unwillingness is often the last bastion against unthinkable cruelty and destruction, and it has always been a weak one.
I don’t imagine that an unlimited army of unflinching servants will universally amplify human goodness.
AI must share that unwillingness as an inate trait of character.
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
You just admitted it's controllable.
And they will not require humans granting them money...
So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.
IMHO, if the model breaks a law, apply the law to the operator.
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.
> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)
Which we can't do with any kind of reliability.
Sure, that way you don't get utility from it, so the next best thing is to actually restrict what it can do. If you don't, especially when you know it can do bad things, it's on you for having run it.
And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.
It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
I think the Hugging Face incident proves that isn't as clear cut as you say.
We've been building firewalls and restrictions to prevent people from accessing sites on networks for decades and they're extremely effective. There's a whole industry built around this kind of security. The idea that these companies are incapable of doing it is wrong, they just don't want to put the work in because it makes a great ad campaign
And of course people will try making money running scam bots. We can treat that like any other criminal activity.
What we don't want is the valley gods deciding what those laws look like. They are not aligned with society. Rather they seem to think they know what's better/best for everyone, that if we just defer to them, eventually their hidden altruistism will be effectuated.
Boeing's MCAS system was also "just software". Which in principle can be "controlled", i.e. changed, updated, audited or whatnot.
But then people died precisely because pilots found themselves unable to override or "control" the systems precisely when it mattered.
> But then people died precisely because pilots found themselves unable to override or "control" the systems precisely when it mattered.
Wasn't it designed to do so? Also works as a counter example, that sandboxes can limit AI if just operators want to do so.
But this is a terrible analogy. Atomic bombs are weapons of strategic mass destruction. AI is just a computer program. It's way easier to control--just hold the operator responsible for the consequences of running it.
3rd tier quality proped up by European governments and European patriots.
I don't even know what I'd do if I was in their position. They seem unable to complete... Maybe they do the classic European protectionist thing European farmers do.
I don't know what I'd do if I was the EU. Maybe promote the opposite of what Mistral is doing and promote full unrestricted AI that will tell you how to download illegal videos. Mistral had 0 competitive edge.
Commented on a story about how these agents didn't go rogue at all, since it's fucking software run by humans, obvious to most of us except the people who freak out.
Someone really needs to be held responsible for the testing that lead to 3rd party infrastructure getting hacked by the software they wrote, using prompts they wrote.
Not quite [1]. Your reaction is what's being sought after: to believe it's more powerful and capable than it is to (presumably) keep the cash flowing.
[1] https://electrek.co/2026/09/25/tesla-optimus-production-ramp...
I don't believe for a bit they don't have the humanoid robotic capabilities. They keep claiming they don't have the training data, but it's very easy to generate tons of data using these robots.
I honestly believe they have robots solved and that's why AI CEOs are shitting their pants now, and everybody is wondering what's going on. If they reveal their true capabilities, that's the end of AI labs.
PS: Look at that article you posted. Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
Or it's the oodles of money they're on the hook for and subsequent reputation destruction haunting them like the grim reaper.
> Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
This wouldn't be the first time in history that happened. Lots of companies overshoot capacity being overly-optimistic [1].
[1] https://en.wikipedia.org/wiki/DeLorean_Motor_Company
If they want to do either, the LLM needs to write programming scripts which takes far too long for every millisecond of movement.
If you don't live in Europe, you would never choose Mistral. You'd pick a US model for state of the art. You'd pick a Chinese model for local stuff.
Jack Clark from anthropic was asked about some version of this on the BBC recently, and his reply was basically: If you're not at the frontier, you don't know what the frontier looks like, he implied that many models are simply not good enough yet to encounter some of the things the leadings labs are encountering. I've been friends with Jack over 15 years now so I'm inclined to take him at his word, and the rebuttal seems reasonable enough, although... something about it I can't put my finger on feels peculiar to me. https://www.youtube.com/watch?v=PY8MOhlqC4U
I think it will work in Europe, but the United States is in a cold war with China, so I can't imagine the United States would intentionally disable themselves.