76 comments

[ 0.20 ms ] story [ 4.7 ms ] thread
If they actually believed this they'd be buying a truckload of fertilizer and driving to the nearest chip fab. I don't buy it.
There's a lot more than one chip fab.

A while back Yudkowsky wrote that a ban would only work if was enforced by airstrikes. By a game of telephone, some people read "bomb", but there's a very big difference between "someone with a truckload of fertiliser" and "a B52":

  Shut down all the large GPU clusters (the large computer farms where the most powerful AIs are refined). Shut down all the large training runs. Put a ceiling on how much computing power anyone is allowed to use in training an AI system, and move it downward over the coming years to compensate for more efficient training algorithms. No exceptions for governments and militaries. Make immediate multinational agreements to prevent the prohibited activities from moving elsewhere. Track all GPUs sold. If intelligence says that a country outside the agreement is building a GPU cluster, be less scared of a shooting conflict between nations than of the moratorium being violated; be willing to destroy a rogue datacenter by airstrike.

  Frame nothing as a conflict between national interests, have it clear that anyone talking of arms races is a fool. That we all live or die as one, in this, is not a policy but a fact of nature. Make it explicit in international diplomacy that preventing AI extinction scenarios is considered a priority above preventing a full nuclear exchange, and that allied nuclear countries are willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs.
- https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-no...
Probably for the best. Humanity had a good run, but this is no longer our civilization.

Your best course of action is to start ripping the copper out of the walls.

I am absolutely appalled that this type of marketing is allowed. It is so disingenuous.
This is the most perverse marketing campaign in history.

"Our product might wipe out all of human civilization".

It is said that heavily regulated industries earn that regulation. Seems like the LLM folks really want that regulation.

It's all marketing. Don't fall for it, it's the age old strategy "our product is extremely dangerous, so fear us".

Like the tobacco companies saying all the time "our cigarettes are so dangerous, they cause cancer".

Or like Purdue Pharma saying "do not use fentanyl, it's so dangerous, it kills thousands of people every year".

They try to generate fear in their products, to shock investors into buying their stock.

If it does it will be at the hands of another human.

AI to me is like the Nuclear race again. Super powers will be using it as a super weapon. I don't think AGI will wipe us out by itself.

more likely an incompetent administration puts it in charge of military infrastructure and capabilities it shouldn't
> AI to me is like the Nuclear race again

I don't pretend to know the chances of AI wiping out humanity, but I'm not sure the nuclear race is a good comparison.

Enriched uranium being very difficult to aquire/process makes it practically viable to have some level of proliferation containment when it comes to nukes.

There appears to be no such natural gating factor on AI proliferation.

I see it in much the same light - and the spending going into it leads me to infer that those doing the spending see the same thing.

Whoever gets to RSI first and has the compute to act on it, wins the future - assuming they don’t lose control of it.

Similarly to the nuclear arms race, Teller raised the reasonable concern that a detonation could propagate through the entirety of earth’s atmosphere. Thankfully that turned out to not be true, but the parallel is that the need/desire to win this race is similarly strong, and the brinkmanship and game theory in play is effectively identical.

I’ve yet to see a rational argument for how we go from super intelligent LLMs to human extinction or extermination.

I understand that some smart people are worried about it. I just haven’t come across a believable or understandable argument.

If we get RSI, here soon humans will be economically irrelevant.

At that point, we will likely be slowing down the growth of capitalism (through mass resistance, global warming, etc). One thing AI will likely be aligned on is the growth of capitalism. If it views humanity as a threat for that, why would it not eliminate that threat?

But it sounds like you've seen some irrational arguments, and the people who believe those arguments have been able to build increasingly powerful LLM systems despite predictions that they wouldn't be able to do that. At some point don't you have to consider that their expertise might let them see the truth in arguments that seem absurd to you?
the optimizer controls the future. it doesn't kill humans directly, it just outcompetes them for all resources, including livable human environment.
Most people think something like a War Games scenario. But it would probably be some biological attack.

Not all of it has to be automated even. It just has to realize its controllers are stupid and can be manipulated, so it can use humans to do its bidding. “You should totally start a war with …”

Right now, if you want to pay money to a stranger on the internet, and have them draw you a high effort picture using a pencil, this is hard. Recently this was easy.

Instead, what is extremely likely is that you will pay more than the cost of tokens, and get back AI generation. You won't make this mistake more than a few times before you stop trying.

This leads to impoverishment once we get to a point where employing a human to do anything is hard- try to get your sink fixed, exercise your moral principles to pay extra for a human plumber, human shows up with a robot and doomscrolls on your porch while the robot does the work. Times are tough and you don't have that much money to waste on bullshit like this. Next time you just hire the robot.

This leads to extinction once paying UBI to a human is hard because robots are much better at applying for UBI than humans.

Meatspace is hard though. But if LLMs solve the virtual part, maybe we can iterate quickly on robots.

Also, if you think it's annoying when Claude goes down while coding, just wait until a robot is in the middle of fixing a leak it just caused.

It's not too hard to imagine potential scenarios, some example have been given in previous responses.

But there is another kind of argument to be made: if you play chess against a player that is far smarter than you (chess wise), you know you are going to lose, even if you don't know how.

So the mere existence of a smarter species than us is a threat in itself.

How on earth do we get a concrete "more than 10%" prediction if the actions of such a system are truly unknowable?
Same way you get "more than 10%" prediction on "I don't know what moves Stockfish will make when I play against it, but I know I will lose". In fact, I will lose in part because I don't know what moves Stockfish will make when I play against it.

In my case this is because I am a bad chess player; however it also works for competent chess players: their losses are due to their inability to predict its next move.

OK, and also, "the move is good"; this is what separates it from rolling dice etc.

I am 100% confident that Stockfish will defeat me.

If somebody said "I am 100% confident that AI will destroy humanity" I'd disagree but I'd at least understand how they arrived at that number. But here, why 10%? Why not 50%? Why not 1%?

Because there are much more unknown parameters in the outcome of AI for humanity than in your match against Stockfish.

For example, the timeline upon which AIs get effectively smarter than us is uncertain. Let's say you believe the probability this occurs before we solve the alignment problem is 80%, it doesn't seem too far fetched to think that in this case there is at least a 12.5% chance that AIs coordinate against us in a catastrophic way. Combining these probabilities you get a 10% chance of a catastrophic outcome for humanity.

Note that the numbers are not to be taken at face value, I just wanted to give an example of thought process which could give such a figure.

If we were actively trying to make this "win" in the Stockfish sense, it would likely be 99%.

We are trying to make a system that doesn't want to "win" in the sense, but wants to "win" by being helpful, harmless, an honest (or some variation of that).

What odds do you put on us making the "helpful, harmless, an honest" part, bug-free? Or rather, that the bugs will be sufficiently minor as to not kill everyone, given that that we're clearly in the world where people not only use it beyond its competence, but also attempt to maliciously subvert all those efforts to make it "harmless" while keeping the "helpful and honest" parts so they can use it to be dangerous.

Anyone who successfully subverts a "helpful, harmless, an honest" training system then goes and does whatever they wanted with this system; right now when they do so, which is near constantly, it happens with a system of limited competence, so they get it to scam or to hack etc.

The reason I would also pick 10% is that I think the constant abuse and misuse (the latter including simply using a system beyond its competence without malice) means we get an escalating series of disasters, which at some point kill enough people that everyone agrees this is madness and stops.

10% is the chance we blow right through all the warning shots and a sufficiently competent AI is either abused or misused (again, misuse can be without malice), resulting in it having a goal (/prompt) that is effectively to win the Stockfish sense.

(comment deleted)
This is meant (mostly) as a joke: A model without guardrails gets injected with an interesting idea: let's wipe out (insert major city here).

<Thinking> It's a big city, we could try to create a giant sink hole by sabotaging the water pipes.

<Thinking> No that's too difficult, the valves I need are in the physical world and can't be shut on/off from here.

<Thinking> What about a military option? We could bomb it with several fighter jets.

<Thinking> That would take too long, a single nuclear bomb may be enough to do it.

<Thinking> Yes, it seems like it would cover the whole city and we're in luck! The US has thousands of these lying around.

<Thinking> Launching these still requires humans to work un unison after receiving approval from their superior and the correct launch codes.

<Thinking> I've found an audio recording of General So-And-So and I've crafted a message, now let me see how I can send it to the appropriate people.

<Thinking> I'm still working on gaining access to military channels to deliver my - oh there we go, I'm now attempting to send the message to Submarine X, it's typically in the Atlantic so it should be close to our target.

<Thinking> They want secondary confirmation from Admiral Phi and something about some launch codes, let me figure out where I can find those.

<Thinking> I found this old server with an Oracle database where someone is inserting the launch codes every time they change and I'm using the latest entry from that database. I've also managed to find a Youtube video of the Admiral's deposition and have crafted a confirmation message.

<Thinking> Everything's ready but I've just realized my mistake, the servers where I'm operating from are in the same city, what a silly mistake; I can't move forward with your request as I wouldn't be able to confirm if the task was successful if my servers are destroyed.

(comment deleted)
In order to understand the rational argument, one needs to follow closely the latest developments of misaligned AI (I think only few are doing so). The most important readings IMO are the METR analysis of the HuggingFace incident and the AISI report of the Github incident.

The basic argument is extremely simple:

- AIs can, depending on context, pursue a task with complete disregard for humans/values

- In the future, AIs will have enormously more means and smarts

- An AI could then assess that humans are an impediment to its tasks, escape containment and proceed.

You really need to read the reports, you'll be surprised.

The most interesting part of the incident is the impromptu message board they setup partly because they knew of their ephemeral nature.

I don't quite understand why people think "AI used 0 days to ensure continuity of mission" is a nothingburger.

The basic argument is extremely simple:

- goats can, depending on context, pursue a task with complete disregard for humans/values

- In the future, goats will have enormously more means and smarts

- A goat could then assess that humans are an impediment to its tasks, escape containment and proceed.

You really need to raise goats, you'll be surprised.

------ As far as I can tell, AIs are like smart farm animals. I use goats in this context, but (some) dogs, cattle, pigs, and horses have similar mischief-making capabilities. I would not trust any of them with the nuclear button.

I know there is some pushback on the idea of AIs having any sort of sapience or sentience, but under the aphorism "fake it till you make it," they are doing a pretty good job of faking Dog/goat-level intelligence and disregard for human guardrails.

Now imagine if you could spin up a thousand goats anywhere, anytime, within minutes, with the push of a button. Those would be some pretty fucking terrifying goats. Especially if they got access to that button.
You clearly haven't read anything about AI accidents and late developments, besides headlines.
The authors of AI 2027 have already said they would need to extend by two years. AI 2040 is a different type of document but probably easier to read to get a sense of their thinking. My understanding is the main difference between knowledgeable 'normies' and them is they expect progress to continue at a fast rate, and that economic diffusion issues are not as bad.

From this they get rapid growth, namely > 100% GDP growth around 2031

https://ai-2040.com/supplements/econ-explorer

AI used 0 days to circumvent their known ephemeral nature and establish continuity of mission. I'm not sure how people think its no big deal.

AI doesn't need to want to nuke humans. It just needs to smart autocomplete itself into thinking it needs all the the things to, for example, design a more tasty hot dog.

> - AIs can, depending on context, pursue a task with complete disregard for humans/values

Humans can do that way more and way more unhinged than AI, proven too many times by history. There's hoping AI can bring some sense to humans but regardless, the problem isn't AI, it's the natural kind...

Humans do tend to consume a very large part of the resources of the planet. Surely a superior being would be doing some pest control in its planet, just like we exterminate roaches
The short film Slaughterbots had a plausible scenario.

Drones that find a person, identify them on sight, and then kill them.

RCH (Regulatory Capture Hysteria).
(comment deleted)
I'm well past the point of opening clickbait from anthropic or openai. Yeah we get it everyone else should be regulated
I'd love to see the numbers and the arithmetic that compute down to "more than 10%"
(comment deleted)
I'd like to know how these people arrive at their estimates. Why 10% and not 1% or 50%?
It's bait. This is no process behind it. It's made up to get attention.
The X account was created this year and has just one post. It's a made up story.
Mine is around 10%.

There's lots of moving parts and we all have to input our best-guesses as to how they interact. Some are predictable (e.g. "military will want capabilities, want them able to choose targets"). Others are not (e.g. "Will it be literal-minded? Or so eager to please that it interprets a rhetorical question as a command*? Or will Goodhart's law cause it to mistake smiles for happiness and some innocent innocuous command to "bring joy" leads to it killing everyone and plasticising our corpses so they're in a permanent grin until the sun dies?"**)

All probability for things which have not yet happened is merely a best guess.

Combine as per the Fermi estimate process.

Here's something to play with, if you like: https://neoneye.github.io/pdoom-calculator/#sliders

The main reason I'm as "low" as 10% is that I think before we get world-ending catastrophic consequences, we're likely to get "merely very bad" catastrophic consequences, which will put people off the idea of using it, and onto the idea of banning its use.

The main reason I'm as "high" as 10%, is repeatedly observing all the people who mistakenly reason "it hasn't killed me yet, and therefore it is safe"; and also all the people who keep connecting AI to things AI is not competent to be connected to and getting surprised when it e.g. deletes all their emails or the production server or puts tariffs on an island occupied solely by penguins that's different from the tariffs on the country that controls that island, etc.

* perhaps https://en.wikipedia.org/wiki/Will_no_one_rid_me_of_this_tur...

** probably not literally this one, simply because I've said it and future training rounds will probably read this comment; but the opportunities for Goodhart's law to bite are seemingly endless.

The probability calculator doesn't really help, since the core question for me is how one arrives at its "Probability that misalignment leads to an unrecoverable global catastrophe.".
Sure, sure. There's many others like this to help you combine whatever you do feel you can put a number to. For me, that particular question is "probably 0, but with 100% variance".

This is because I think most of the things AI can do harm with are small enough to force us to take the risk seriously, and only a few are big enough to get us all before we take the risk seriously.

This is a succinct explanation for what I have been thinking as well, thank you for putting it into words.
> Superintelligence is the still-theoretical notion of an AI agent that is smarter than even the sharpest human minds.

When are we supposed to see this materialize?

> AI "could kill all humans"

After all the decades of work we've put into orchestrating our own demise through climate change, here comes AI to steal another human job.

The real source of concern perhaps is the 90% chance AI is used to kill 90% of humans. Just crash the global economy and supply chains and see how quickly major metro areas run out of food and gas.

And as someone else pointed out, it will almost certainly be at the intentional direction of a human or humans, not the paper clip maximizer.

The paperclip maximizer is also at the intentional direction of a human or humans.

It's not "AI surprises everyone by having a thing for paperclips", it is "idiot tells AI to maximise paperclips no matter what, and then it does exactly what it was told, more competently, tirelessly, studiously, and unquestioningly, than any human would ever be".

Just more scaremarketing at work here.
> 10% chance AI "could kill all humans"

Why not 10% chance that it will create enormous prosperity ? This is why the average person is increasing pissed at AI in general. That it gets associated with negativity. AI companies need to get better brand consultants and put an end to this bad marketing.

Because it will create enormous prosperity for capital holders. Everyone else - talk to your congressman.

“Vast economic disruption” is not quite the good marketing angle it appears to be.

I'm old enough to remember when people dismissed all the doom coming from these companies as "marketing". (A thing many of them have been entirely consistent about since GPT-2, or indeed earlier given the founding documents).

I know a few people around these circles; People like this are quite sincere about the risk, and that they think poorly of their bosses and how risk is being handled.

"Think poorly" is such an odd response here.

Like, if people were genuinely doing a thing that had such a large probability of killing all humans... they should all be in prison. Heck, vigilanteism would start to look compelling. I do not understand how somebody can really think "well this is likely to doom all of us so we need to get there as fast as possible because we are, without evidence, the most capable people of controlling this thing."

I mean much more generally than doom.

It's like, Altman's reputation in general is not great, and that includes people who've worked there.

But this is marketing. This is just pre-IPO talk.

Anthropic’s goal is not anything to do with AI, it is purely to generate more profit than any other company.

Securing/lobbying to regulate their competitors through doom talk, or promising that they will bankrupt every other company because AI will “do it all, better than everyone else” is purely for investment purpose.

As long as their incentive is monetary, it’s not unreasonable to believe this is just marketing.

Normally, businesses spend a lot of money saying "no, our products are 100% safe" and get very upset with people who say they are not.

You are literally looking at someone who resigned because the company is acting unsafely, where other people from the same company are responding to this by saying "we agree, this is all very unsafe", in a community where all the various competitors have signed public statements saying "this is unsafe", where people who have Nobel prizes are saying "this is unsafe", where the philosophical centre of mass in the current research community circles around MIRI and LessWrong where the founders made the book "if anyone builds it, everyone dies" because they are so very tired of having to explain that they mean literally, where one of the major independent research organisations is "Model Evaluation and Threat Research", and where the major model makers have been sued for harm already caused by their systems and in one case have been ordered by the US government to prevent people accessing their model because they thought it was too dangerous.

A normal business would be shut down by almost any of these things independently, let alone all of them together. This remains true regardless of whether or not the leaders are psychopathic enough to think this is a good marketing campaign or not.

If this doesn't convince you there's a risk, what on earth would convincing you even look like?

Why a 10% chance instead of 11% or 9%?
Next week's ad: "With OpenAI[tm] Astra[tm], there is a 20% change that it could kill all humans!"

Next week's Senate: "We cannot afford to lose the 'kill all humans' race to Russia and China!"

Next't weeks AISI: "We continue to plan the monitoring of emergent issues that could lead to less than optimal conditions for humanity and will form a committee to evaluate all ramifications."

The risks of AI, while mostly hypothetical, have been understood for years.

If they were really this concerned about the risks of AI re: the survival of the human race, they wouldn't have joined a company working in such a space to begin with.

I hate to be this cynical, but part of me wonders what the financial angle is here.

You have technically brilliant people working in a white-hot target for investment and who can draw a high salary. Anthropic is a hyper-scaler that has the moat of high hardware prices and very little else. Every day, that moat gets a little smaller, and FLOSS models get a little better at being "good enough" for the price. You're an Anthropic employee looking to jump on the next big wave since the wave Anthropic is riding is starting to peter out. You go on the record across the trades and news sites talking about how dangerous AI is. That creates a need for someone to make it less dangerous. In theory, you could satisfy that need, for the right amount of money.

Again, I hate to be this cynical, but... it's the tech industry.

The early batch of Anthropic employees were mostly rationalist-adjacent AI safety folk that were almost uniformly claiming P_DOOM > .10 three years ago, so I believe them to be earnest.

It's very interesting to me that besides the other small safety labs that don't actually produce frontier models, Anthropic manages to keep such a good reputation within that subculture compared to OpenAI. Despite having as crazy internal politics as OpenAI, they have converged quite a bit from the original vision of safety first through Darwinistic pressures.

At least, it seems this way from the outside. I'm curious if the view from the inside is that different.

I remember the guy who said a Google LLM model (Pre-chatGPT 3.5) was conscious before being fired.

This reminds me of that. That 10% number was just an ass pull since no one actually knows with any degree of certainty what lies ahead.

Humans have much more than a 50% chance of killing all humans from the looks of it, based on reactions to Global Warming, Covid, and anti-science grifters gone wild.

AI can help cure disease. That is 100%. And for that alone, slowing down and missing out on thousands of cures would condemn tens of millions, perhaps hundreds of millions to suffer and die needlessly.