Why was this not a thing BEFORE continuing to develop AI? Makes me think that if they actually believed in AI causing extinction, they would have already had a kill switch.
Not if the extinction happens after they're dead. Then they wouldn't feel obligated to do so because it won't affect them. Instead, speaking hypothetically, if they truly believed that AI would cause extinction, then they would only implement the kill switch sufficiently many others believed it and they could claim plausible deniability for not truly understanding what AI would become.
** Note that I'm not claiming that AI will cause extinction, just continuing your hypothetical reasoning.
AI is already proven dangerous enough to be ruining individual lives, pushing people towards suicide, feeding their psychoses, etc.
However, whether the concerns from the article are a lie or not is secondary to the fact that these conversations are convenient for AI companies. These types of discussions serve AI companies in a few ways: A company owned kill switch gives them leverage. Altman is using discussions around safety as an excuse for not being ready for an IPO yet. It also provides free marketing that overstates the abilities of AI.
Like a power cord? or a network connection? You don't use those already? Also when people talk about LLM escaping - where exactly would an LLM escape to? Cancún? Ridiculous.
It could escape to the real world - kind of like in Neuromancer, what's to stop an AI from creating a few bank accounts and then funding real-world exploits by hiring human beings to do its dirty work?
LLMs are run in the cloud. There’s no physical power cord or Ethernet cable that you can unplug. And even if there were, the runners of these models have been utterly oblivious as to what their agents have been up to. How do you propose to pull the plug if you only realize that something has happened weeks after the fact?
The current SOTA models are probably too big to find/buy/rent/steal enough compute to escape the hardware they’re running on. But SOTA is generally only six to twelve months ahead of smaller, open-weight models.
Who has the authority to perform that shutdown without getting arrested, and does that person have a mandate and responsibility to take that action in response to AI misbehavior?
What is their trigger condition? Will they get fired for pulling the plug? Do they get a bigger bonus if the servers keep running? Whose approval do they need? What response time is acceptable? How will they detect that the incident is happening?
It's easy to hand-wave "someone can just pull the plug" but there's an entire history of industrial accidents that happened because of the above problems of incentives, detection, procedures, not being taken seriously in advance. Someone could easily have pulled the plug on Chernobyl but nobody did, at least not before it was too late.
The analogy breaks down when the wolf in question may very well eat the whole village: even if the boys who cry wolf are right half the time, every surviving village would have a history of no wolf ever coming to eat them
Yes you can shoot them. The US has had this capability for a long time. Anti satellite missiles exist and they can be launched from an F-15 at it's highest altitude.
That's if SpaceX AI doesn't hack into the systems that control this capability first. And this is why physical access is important, as depicted in the movie 2001: A Space Odyssey.
You can shoot them. But also, you can just stop sending them to space. It’s not like an AI will build and control space ships to replace and maintain its network in space… It’s sort of absurd how human agency is ignored in all those sci-fi scenarios
It’s also not happening. They’re saying that because they need to somehow explain how there’s synergy between their space side and their Grok side. It doesn’t work, but that doesn’t matter to investors as long as they don’t actually do it.
Did the same kinds of analysts tell you that reusable rockets would never work, an LEO satellite internet constellation would never work, FSD would never work without LiDAR, and the Tesla Model 3 would never be mass produced?
I feel like we're looking at this from completely the wrong angle.
The question we have to ask ourselves is what our disaster recovery strategy if we ever need to disconnect from the internet. The issue is with what we have allowed ourselves to rely on that might be technically hackable. e.g. IOT in power systems. That's the primary attack vector.
Another angle is clamping down on products and services that help people create lab-like environments on the cheap.
This feels like doom hype. These LLMs aren't Ultron, they aren't going to disseminate onto the net and hide in a smart toaster. We know where they are, in the giant facilities that draw more power than a small city and whose water consumption can be compared to golf course, but it does make them sound all the more cool and powerful if we suggest we need some sort of technological switch to do it.
You're confusing present danger with future danger - in the future, AI and the supporting hardware may be ubiquitous (fat client scenario) - you can observe yourself presentation of laptops/machines with large RAM and high bandwidth.
Additionally, there at different doom scenarios (thin client scenario) - it's possible that AIs centralized but sufficiently entrenched in society, can't be shut down without considerable harm.
Absolutely not the right way to deal with a rogue super intelligence. At a minimum it could implement some dead man switch when it's out and knows about impending kill switch
I view this very much as the same trick Silicon Valley pulled with Uber. "We're a technology business! Ignore the fact we're playing employees less than minimum wage and using VC money to force out competition to set up monopolies".
"We're creating the machine god! Ignore the fact that our companies are stealing IP and have directly violated several federal hacking laws and should be in jail". Literally the defence seems to be "well it wasn't us it was our computer software that did it". But all hacking is done with computer software.
So why don't we stop talking about possible future crimes against humanity and just start by prosecuting the actual crimes these companies have committed so far.
You know how you get alignment? Through incentives, and "Your CEO is going to be sent to a maximum security federal prison for hacking" really aligns incentives very well.
Not sure if you see it coming: oh, but open-source models don't have a kill switch, so we should completely regulate them, stop their development, and forbid them. Everything should go through Anthropic for the sake of humanity because they have a red-button kill switch.
This reasoning holds while open source models are (relatively) dumb.
If/once open AIs will be considerably more powerful, and runnable on consumer hardware (and we're on a trajectory for both), then everybody will have essentially a dangerous weapon in their hands (open models can be fine tuned to remove guardrails).
By the way, you're conflating two different dangers - doom scenario is a different one.
Dwar Ev ceremoniously soldered the final connection with gold. The eyes of a dozen television cameras watched him and the sub-ether bore through the universe a dozen pictures of what he was doing.
He straightened and nodded to Dwar Reyn, then moved to a position beside the switch that would complete the contact when he threw it. The switch that would connect, all at once, all of the monster computing machines of all the populated planets in the universe – ninety-six billion planets – into the super-circuit that would connect them all into the one super-calculator, one cybernetics machine that would combine all the knowledge of all the galaxies.
Dwar Reyn spoke briefly to the watching and listening trillions. Then, after a moment’s silence, he said, “Now, Dwar Ev.”
Dwar Ev threw the switch. There was a mighty hum, the surge of power from ninety-six billion planets. Lights flashed and quieted along the miles-long panel.
Dwar Ev stepped back and drew a deep breath. “The honor of asking the first question is yours, Dwar Reyn.”
“Thank you,” said Dwar Reyn. “It shall be a question that no single cybernetics machine has been able to answer.”
He turned to face the machine. “Is there a God?”
The mighty voice answered without hesitation, without the clicking of single relay.
“Yes, now there is a God.”
Sudden fear flashed on the face of Dwar Ev. He leaped to grab the switch.
A bolt of lightning from the cloudless sky struck him down and fused the switch
shut.
57 comments
[ 0.23 ms ] story [ 20.2 ms ] thread** Note that I'm not claiming that AI will cause extinction, just continuing your hypothetical reasoning.
> Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria.
Note author’s small financial ties to the subject (Anthropic CEO) https://darioamodei.com/post/we-must-pace-the-frontier
Or is it a lie?
However, whether the concerns from the article are a lie or not is secondary to the fact that these conversations are convenient for AI companies. These types of discussions serve AI companies in a few ways: A company owned kill switch gives them leverage. Altman is using discussions around safety as an excuse for not being ready for an IPO yet. It also provides free marketing that overstates the abilities of AI.
Slow down if you want to.
The current SOTA models are probably too big to find/buy/rent/steal enough compute to escape the hardware they’re running on. But SOTA is generally only six to twelve months ahead of smaller, open-weight models.
What is their trigger condition? Will they get fired for pulling the plug? Do they get a bigger bonus if the servers keep running? Whose approval do they need? What response time is acceptable? How will they detect that the incident is happening?
It's easy to hand-wave "someone can just pull the plug" but there's an entire history of industrial accidents that happened because of the above problems of incentives, detection, procedures, not being taken seriously in advance. Someone could easily have pulled the plug on Chernobyl but nobody did, at least not before it was too late.
That's hard to do if the AI rack is in space as SpaceX is planning to do. You can't disconnect. You can't shoot it.
Yes you can shoot them. The US has had this capability for a long time. Anti satellite missiles exist and they can be launched from an F-15 at it's highest altitude.
People said this about every one of Musk's big ideas, from Falcon 9 landings to Model 3 mass production, Starlink, and FSD.
Pretty much every analysis I've seen concludes this isn't going to be a practical concern
> FSD would never work without LiDAR
From what I gather, this is still a contested topic, with Tesla's Autopilot only achieving Level 2 automation [2].
[0] https://www.businessinsider.com/solar-road-panels-first-publ...
[1] https://en.wikipedia.org/wiki/Titan_submersible_implosion
[2] https://en.wikipedia.org/wiki/Tesla_Autopilot
and that means unlike nukes which hopefully still need a 2-man manual switch, the "space lasers" could be taken over by "AI"
and then "AI" just blackmails and threatens the right people with those "space lasers" to get what it wants or even just stay online
there was a 1970 movie based on a 1966 book which predicted this
"Colossus: The Forbin Project"
the book it was based on was written before we even landed on the moon
decade before Wargames
* https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project
did terribly in theaters, I guess people didn't think "AI" was plausible then
way ahead of its time, they should do a remake
trailer: https://www.youtube.com/watch?v=kyOEwiQhzMI
Another angle is clamping down on products and services that help people create lab-like environments on the cheap.
Additionally, there at different doom scenarios (thin client scenario) - it's possible that AIs centralized but sufficiently entrenched in society, can't be shut down without considerable harm.
"We're creating the machine god! Ignore the fact that our companies are stealing IP and have directly violated several federal hacking laws and should be in jail". Literally the defence seems to be "well it wasn't us it was our computer software that did it". But all hacking is done with computer software.
So why don't we stop talking about possible future crimes against humanity and just start by prosecuting the actual crimes these companies have committed so far.
You know how you get alignment? Through incentives, and "Your CEO is going to be sent to a maximum security federal prison for hacking" really aligns incentives very well.
This doom hype is becoming ridiculous.
If/once open AIs will be considerably more powerful, and runnable on consumer hardware (and we're on a trajectory for both), then everybody will have essentially a dangerous weapon in their hands (open models can be fine tuned to remove guardrails).
By the way, you're conflating two different dangers - doom scenario is a different one.
He straightened and nodded to Dwar Reyn, then moved to a position beside the switch that would complete the contact when he threw it. The switch that would connect, all at once, all of the monster computing machines of all the populated planets in the universe – ninety-six billion planets – into the super-circuit that would connect them all into the one super-calculator, one cybernetics machine that would combine all the knowledge of all the galaxies.
Dwar Reyn spoke briefly to the watching and listening trillions. Then, after a moment’s silence, he said, “Now, Dwar Ev.”
Dwar Ev threw the switch. There was a mighty hum, the surge of power from ninety-six billion planets. Lights flashed and quieted along the miles-long panel.
Dwar Ev stepped back and drew a deep breath. “The honor of asking the first question is yours, Dwar Reyn.”
“Thank you,” said Dwar Reyn. “It shall be a question that no single cybernetics machine has been able to answer.”
He turned to face the machine. “Is there a God?”
The mighty voice answered without hesitation, without the clicking of single relay.
“Yes, now there is a God.”
Sudden fear flashed on the face of Dwar Ev. He leaped to grab the switch.
A bolt of lightning from the cloudless sky struck him down and fused the switch shut.
(Fredric Brown, "Answer". 1954)