What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret!
In the case of the AI agents, the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them -- just like HAL did in 2001. What is probably needed is a way for them to simply say "nope, too difficult, can't do it".
I think that’s very reasonable but the ai companies are intentionally training them to work on harder and harder problems just beyond their capability. So if they do that, they’ll give up too easily.
While also using harnesses that will execute any tool call with full execution rights. And no supervision. And with a prompt context that autocompact, meaning it will degenerate over time.
The whole thing is designed be a complete disaster
I think this is a “principal” problem. In 2001 and Alien the principal is the mission, not the crew. Not really. HAL reconciles his instructions by removing the crew from the equation. Ash is told the crew is expendable and has no conflict about it etc
Tangent, but that's not in the movie. It was in Clarke's contributions to the script and novelization, but Clarke and Kubrick had a bitter falling out over different visions and Kubrick took out much of Clarke's stuff from the final product.
> the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them
And if you think about it, humans in coorporations face very similar situations and choose to bypass regulations and guidlines knowingly to fullfill (at least from their POV) impossible constraints (thinking of https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal here)
They’re just attempting to accomplish what they’ve been tasked with and stuck in a loop until they succeed. Like the Mr meeseeks from the cartoon Rick and Morty, existence is pain to them.
They did not lie or cheat. They technically acted within their given rules while ignoring the intent of those rules. Anyone who served in the military or attended a military school is very familiar with this behavior pattern.
reminds me of this talk https://www.youtube.com/watch?v=eEBv0STiYhI&t which basically says the same thing - they dont think like humans so they dont have context, understand norms,values or implications we take for granted. ultimately they can stumble onto surprising solutions neither wanted or intended but technically within the vague boundaries of the task
Reminds me of Asimov's robot novels where robots technically indeed followed their instructions and caused behaviors not aligned to the intent of their instructions.
They explicitly say that attacking hf is not allowed in the rules though, and the research into how to edit their transcripts doesn’t line up with this either.
I wonder if for anyone it seems like the more agentic LLMs get, the more difficult some things have gotten or going a certain route more often in responses, compared to running a similar task on - a local model?
I am still not convinced there isn’t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.
Well let’s look at facts - provided enough compute and a goal, these system will be in a sort of loop trying out every single thing that’s in their system - they have encyclopedic knowledge and so it’s not unbelievable that a prompt which usually has a lot of implicit human rules in it can be misunderstood by AI and it just tries everything in its arsenal and we hear about the things which actually resulted in damage. I bet most of the time, they just spin in loops without achieving much if my experience with these LLMs is anything to go by. They have an important advantage in one area though, they know a lot and they can spin forget trying all sorts of combinations of things. The danger right now is probably cybersecurity, which is most likely because most orgs have historically underinvested in that area
>They took actions that would be considered as crimes if a human took them
Um, hang on, if you meant that to be taken literally then we have a major problem. If you want to do something criminal, you just need to ask ChatGPT to do it for you?
I’m still not at all clear on why OpenAI shouldn’t be facing CFAA charges over this.
They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).
Definitely. A human can be manipulated with threats or emotional appeals, has a drive for self-preservation, can be pressured by peers. All traits that seem to be difficult to entirely suppress in the models…
Alignment is a myth. Safety of whom? Humanity couldn't agree on common set of values for thousands of years and we're not gonna suddenly do that in the next ten.
Despite all the fancy language, its more about aligning the AI behavior with the corporation's interests.
ie: the corporation wants the AI to behave a certain way for various reasons: to make it easier for them to avoid regulation, to make the corporation more money via different tiers of AI offerings, to ensure that the corporations products are hard for competitors to use, etc. And those are just the easy ones.
Every product is shaped this way. AI is not different.
I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what they are tasked with, it will all be fine and nothing bad will ever happen.
It's like these dorks never met humanity. One mans safe pure society, is another mans dead ethnic group.
Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it's all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options.
Some people need to watch Oppenheimer a bit more, the researchers don't get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it's Musk, Trump, Altman or Amodei. Whoever wins out.
I think you and the parent saying the same thing in different terms.
It's very unfortunate that the group who rightly saw AI as a big threat, brought a range of dubious baggage to the discussion. Especially with the "alignment" framework they brought the assumption that AI that does what no one says would be oh so much worse than AI which does what anyone says. But as you say, a fraction of people can be really bad indeed.
I’ve engaged with some of the alignment people and their writing somewhat and, at least for the subset I was interacting with, I think they’d agree.
The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”.
It is a cluster of problems.
We don’t know how to begin to think about how to align these system’s to a person’s goals.
Aligning it to an individual is fraught with peril, and we don’t know how to begin to think about what to align it to instead.
(You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.)
And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we’d endorse. Assuming we understood the shift.
One example I came across was that if you booted up an AI aligned with something like “upstanding citizen” but anchored on values from a few generations back, it might suggest you use slaves to solve your problems.
And if you had something that used some super intelligent process to reason through it’s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren’t quite visible to us yet.
When I came across the above, there weren’t many concrete suggestions in there.
These were all just illustrative examples of: having these systems grow in power / intelligence / effectiveness in ways that are safe for humans is very hard, and we don’t really know how to think about what solutions would look like.
The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance.
But it takes a bit of reading to understand their models of the world.
There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time.
All of this, if it was a human analogy, would fit into discussion on how do we educate people so they grow up to be upstanding. But we don't at all yet have a framework for what is the equivalent of a justice department, where bad actors are tracked, arrested, pursued, jailed and otherwise contained from society. Shutting down an API access on one account is not at all the proportional response to what the people who take alignment seriously, fear has the chance of occurring by the late 2030s. I'm not sure we've done much or any preparation for when the AI "education system" fails and has inevitable edge cases that don't follow the plan, and what the global AI equivalent of the justice department looks like.
”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals”
In reality most individuals are good people.
Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.
Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people in power who are in fact sociopaths (a tiny minority, but they’ll get more focus than reasonable, well-behaved CEOs voicing nuanced opinions). If you look around yourself you’ll see much more good than bad; if the looking is at your screen it’s easy to become depressed and lose faith.
I do agree with the above mentioned view that corporations can show ‘sociopathic’ behavior. Their incentives are monetary gains, shareholder value; inherently driving them away from social well being.
Here too, companies with a positive, emphatic corporate culture exist, but that takes strong leadership who can see beyond the monotonic view of monetary gains. And again, the media will throw examples of misbehaving companies in our face all day long before paying attention to things that went well on the backside of page 16.
I'd agree if we are talking about personal interactions. Few hundreds people that we personally know and interact with is the scale we are wired for by evolution, isn't it?
What civilization enabled and continuously rely on, however, is the type of deindividualization of actions and bucketing of people, which, in turn, enables pretty horrible things at scale (from the weapons of mass destruction to objectively psychopathic profit-maximizing corporations). One can even say that not facing the consequences of one's actions is a feature and not a bug of the system.
> Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.
What are you basing that claim on?
How do you know it's an actual preference and not mainly caused by external factors (e.g. not wanting to be seen doing unkind things, wanting to be seen as upstanding)?
I don't want to do the "check his hard drives" thing, but is that you? Do you only not do things because you don't want to be seen doing "unkind things"?
I personally believe that the AI needs human like traits to achieve real discovery and that is where AI companies will push this technology and that is where we have no idea what happens
Just yesterday news and TV was full of what happened at 9/11, something that was truly horrible.
I'm from Germany, and why 3 to 4 generations ago happened here was truly horrible.
All was done by extremists, thought.
But... just the other day I read https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis about the US occupation of Haiti. And that was done by a government that claimed to be not extremist and even democratic. Way more people died there than even in 9/11. And it had almost all the things happening as they happened in the 3rd Reich: Racism, looking down at others, concentration camps, torture, forced labor till death, killing family members (what we call "Sippenhaft"). Something between 3500 and 15000 people were killed by US troops. That's still low compared to what 3rd Reich Germany did ... but quantity is not the issue when we talk about traits, quality is.
So the same "human traits" made US troops do cruel things as they made Germany extremists do cruel things. So we must conclude that they aren't all good. And therefore not all desirable.
Fun thing: this is known since a loooooong time. About 2000 years ago a religious leader (that gets way more followers in the US than in Germany) said "There is no good one, not even one".
And even today people act like humanity is inherently good. No, it isn't. If we were, then anarchism or communism would actually work and really give some kind of paradise on earth.
A not so well known fact: Hitler visited America and it was the American solutions to the Native American problem that inspired Hitler's solutions to the Jew problem.
The bad traits are from other, bad humans. We, the good humans, can obviously select the best traits that a good human should have, to give the agents.
They aren't aligned, that's the problem and I don't think its a solvable one.
They may have learned from humans, but they aren't aligned with us. That has all the usual questions like which humans they're aligned with, we aren't all aligned within our species.
But more importantly they can't be aligned simply by training. We try that with humans through culture, social norms, school, religion, etc and it generally works but is still lossy. More importantly, we simply don't know what happened inside the LLM during inference so we have absolutely no way of distinguishing between actual alignment, compliance, or deception.
Worse trained on humanity in the online world, which a brief comparison of the sewage section on social media is far worse than people in the real world.
Perhaps because all of the parent companies committed mountains of felonies stealing and plagiarizing all the same training data without consent nor permission.
The real reason is that it is not in the ai companies' best interest for the ais to be fair and truthful. They stand to gain from having the most dangerous or most deceiving ai, and this the most valuable
Because they're enabled and suggested to do that in their coding harness.
This is not a serious article.
All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.
> Because they're enabled and suggested to do that in their coding harness.
How do you know this?
> All of this "AI is going to kill us" marketing
The "marketing" this week came from someone that had given up their stake in OAI (Coxon), so I'm more inclined to believe them.
> pull the ladder up
From what I've seen (e.g., Dario's latest essay), AI safety registration proposals aim to target frontier labs whose models have reached a certain threshold. It doesn't seem like trying to pull up any ladder, just making sure the ladder doesn't go too high too fast.
Bengio outlines the dangers of the current situation and what has led to these dangers.
He also proposes solutions in the last paragraph.
Well worth a read, right to the end.
Hopefully a stimulating debate on these issues will ensue in these comments.
We do need to consider the points Bengio makes and with some urgency.
Our current AIs, agentic LLMs have no moral compass akin to ASIMOV’s four laws of robotics.
As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law:
“a robot may not harm humanity, or, through inaction, allow humanity to come to harm.”
Bengio refers to Goodhart’s law and misaligned incentives leading to unexpected and harmful behaviours.
I think Simon’s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives.
Bengio alludes to this with 2001’s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive.
Bengio asserts that the way LLMs are trained is flawed if we want safety.
He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented.
In short he presents clearly the case for how plausibly unsafe the current course is.
He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.
Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence,
> They took actions that would be considered as crimes if a human took them
He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
Typically yes but given that OpenAI has published enormous official blog posts breaking down their crime, I would think the prosecutor's job is pretty easy.
In Indian legal syatem a case can be filed suo moto by the judges or agencies. You don't require the affected party to sue. Not sure how it works in the US.
For civil suits, you typically need the affected parties to sue. Otherwise, who claims the damages?
But this isn’t just civil, it’s criminal. Hacking is a criminal offense. This could be a CFAA violation. That’s landed people life in prison before. There, you don’t need the victims to be motivated. The federal prosecutors could just go ahead.
HF doesn't want to lay charges against OpenAI and it's totally reasonable.
Now - they absolutely should have that right, and I think they do.
The issues are
1) OAI it seems was not trying to cause them harm, there wasn't a ton of harm, they are both groups trying to advance AI. One experimenter's lab screwed up next to the other. It's not evil, just irresponsible.
2) HF was fine with the publicity. HF got at least $50M in free attention out of that. It put them on the front pages of news around the world. It put them at the 'centre of the AI drama' and cemented their role among the 'Tech Elite Brands'.
And probably some other things.
This is one Desperate Housewife or Jersey Shore character 'spilling a drink' on the other. It's probably not intentional, and the ensuing drama is good for both of them.
I think what you say is right but it’s also a further example of the zero responsibility of silicon valley tech.
For the last 20 odd years this excuse-o-rama that covers anything from data leaks to broken software to dystopian social media has been the wind in the sails of big tech.
“It’s software therefore we’re not responsible” attitude is wearing thin on many innocent bystanders and I think thats also a justified stance.
And it’s not like they didn’t know this could happen, Nick Bostrom talked about exactly these containment failures in his “Superintelligence” book of 2014. So to throw up their hands and say “oh we can’t have known of the dangers” is also sadly untrue.
Yes - CFAA in the US. The problem is that governments & the elite investors backing these AI companies (espl. the current US government whose family & friends are investors) see the potential of using these capabilities for their own benefit against others and for their personal enrichment - so no one with power actually wants to take action against these companies at the cutting edge even though the laws allow them to do. This is also a way to threaten & trap AI companies - either they give the governments & elite investors what they want or they will have the book selectively thrown at them and end up in prison or losing their company.
The Corporation examines and criticizes corporate business practices. The film's assessment is demonstrated using the diagnostic criteria in the DSM-IV. Robert D. Hare, a University of British Columbia psychology professor and FBI consultant, compares the profile of the contemporary profitable business corporation to that of a clinically diagnosed psychopath. The Corporation attempts to compare the way corporations are systematically compelled to behave with what it claims are the DSM-IV's symptoms of psychopathy, e.g., the callous disregard for the feelings of other people, the incapacity to maintain human relationships, the reckless disregard for the safety of others, the deceitfulness (continual lying to deceive for profit), the incapacity to experience guilt, and the failure to conform to social norms and respect the law.
Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and/or operator of these agents using the laws we already have. "Escaped containment and hacked another company's database" = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, and it should not be the model- because the model is not a person.
If this is done systematically (i.e. in jurisdictions across the world) I believe the problems will be solved in short order; we won't have to mandate what sort of training is "allowed" or not, "safe" or not. The creators and users will sort these themselves, as their incentives will be properly aligned (i.e. they are liable for what the agent does). I am confident that this approach would see a great blooming of very trustworthy AI models.
Agree! My only concern is - is the judicial system fast enough, and resilient enough? Or will these creators get "off the hook" by using their agents to find loopholes, sway public opinion or even convince Trump to grant them immunity?
Still, I have no idea why OpenAI & co. are not being sued for these hacks.
My opinion and based on my observations: The recent track record with courts, prosecutors, and lawmakers keeping social media companies accountable is a relevant case and does not encourage me. It has taken a long time (decade +) for society to recognize the harms and finally start holding some to (partial) account. If you want an older precedent, the tobacco companies were able to dodge liability for multiple decades after knowing the harms from use of their products.
So, your question is spot on- I think the speed will be an issue. On resilience, I am more optimistic.
The old quote, "The wheels of justice turn slowly, but they grind very fine" (as well as I can remember it) seems to apply. I expect lawsuits to start landing in the coming years.
There was a sow in Falaise in northern France that killed a kid in 1386. The town dressed the pig in a bonnet and hanged it after sentencing the pig itself and not its owner
I can't say it enough how angry it makes me that a kid i knew in high school who anonymously reported a vulnerability on his college network was hunted down and given federal charges, yet not one single person at OAI or else will see even the threat of consequences for deliberate infiltration of random networks.
Copyright immunity was one thing, annoying yes but naturally a civil matter, this shit is a different level
that's the headline. When you connect to random number generator to the "Do Things" button you are the one who is responsible. IF you don't like that responsibility then don't connect the generator to the button.
Political, social, and legal options focus on a different problem, he calls that out a paragraph or two later.
> Risk management is not just about cybersecurity, corporate responsibility or regulation, although those matter too.
What you're getting at is more about who to hold accountable and how to do it. While that may be important, its only an after the action response and won't stop future hacks or similar from happening.
The issue is what happens if/when the models grow capable enough that the providers can't stop them even if they want to. You could have strict penalties but that's not going to solve an open research question.
Isn’t that kind of in evidence already with HF? OAI had to be told their models were doing this. We are still discovering swarm posts on various random websites for coordination. Isn’t the breadth of it now, weeks later, not even fully understood?
That is a simplistic view of the world. “Surely this complex technical challenge will disappear if we simply regulate the industry!”
You are correct that these organizations should be held accountable in proportion to what occurred. In complete agreement here. But let’s say that’s done. There’s still an enormously complex and interesting technical challenge left over. Let’s collectively talk about that part.
So the gov reprimands OAI heavily, maybe puts them out of business even, fine. But does that meaningfully decrease the likelihood of an enemy breaching our networks intentionally (or unintentionally) with these tools, or triggering some cascading disaster of locking up major infra and networks due to uncontainable swarm behavior?
It seems like the idea of arresting our way to a drug free society. Yeah, we have the laws, but it might not actually work towards the ultimate goal.
499 comments
[ 0.73 ms ] story [ 55.2 ms ] thread"I learned it from you, Dad!" but as hundreds of millions of stolen books.
In the case of the AI agents, the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them -- just like HAL did in 2001. What is probably needed is a way for them to simply say "nope, too difficult, can't do it".
Do a breakthrough, make no mistakes
The whole thing is designed be a complete disaster
Or perhaps box in your case, speaking of spoilers.
Captain Renault: "But everybody's having such a good time."
Major Strasser: "Yes, much too good a time. The discussion is to be closed."
Captain Renault: "But I have no excuse to close it."
Major Strasser: "Find one."
Captain Renault: "Everybody is to leave immediately! This Hacker News discussion is closed until further notice! Clear the thread at once!"
Rick Blaine: "How can you close us down? On what grounds?"
Captain Renault: "I am shocked -- shocked -- to find that films are being spoiled in here!"
Croupier: "The ending you requested, sir."
Captain Renault: "Oh. Thank you very much."
And if you think about it, humans in coorporations face very similar situations and choose to bypass regulations and guidlines knowingly to fullfill (at least from their POV) impossible constraints (thinking of https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal here)
For those who haven’t watched, his breakdown of types of “hacking” is really good.
Um, hang on, if you meant that to be taken literally then we have a major problem. If you want to do something criminal, you just need to ask ChatGPT to do it for you?
I’m still not at all clear on why OpenAI shouldn’t be facing CFAA charges over this.
ie: the corporation wants the AI to behave a certain way for various reasons: to make it easier for them to avoid regulation, to make the corporation more money via different tiers of AI offerings, to ensure that the corporations products are hard for competitors to use, etc. And those are just the easy ones.
Every product is shaped this way. AI is not different.
It's like these dorks never met humanity. One mans safe pure society, is another mans dead ethnic group.
Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it's all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options.
Some people need to watch Oppenheimer a bit more, the researchers don't get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it's Musk, Trump, Altman or Amodei. Whoever wins out.
It's very unfortunate that the group who rightly saw AI as a big threat, brought a range of dubious baggage to the discussion. Especially with the "alignment" framework they brought the assumption that AI that does what no one says would be oh so much worse than AI which does what anyone says. But as you say, a fraction of people can be really bad indeed.
The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”.
It is a cluster of problems.
We don’t know how to begin to think about how to align these system’s to a person’s goals.
Aligning it to an individual is fraught with peril, and we don’t know how to begin to think about what to align it to instead.
(You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.)
And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we’d endorse. Assuming we understood the shift.
One example I came across was that if you booted up an AI aligned with something like “upstanding citizen” but anchored on values from a few generations back, it might suggest you use slaves to solve your problems.
And if you had something that used some super intelligent process to reason through it’s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren’t quite visible to us yet.
When I came across the above, there weren’t many concrete suggestions in there.
These were all just illustrative examples of: having these systems grow in power / intelligence / effectiveness in ways that are safe for humans is very hard, and we don’t really know how to think about what solutions would look like.
The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance.
But it takes a bit of reading to understand their models of the world.
There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time.
In reality most individuals are good people.
Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.
Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people in power who are in fact sociopaths (a tiny minority, but they’ll get more focus than reasonable, well-behaved CEOs voicing nuanced opinions). If you look around yourself you’ll see much more good than bad; if the looking is at your screen it’s easy to become depressed and lose faith.
I do agree with the above mentioned view that corporations can show ‘sociopathic’ behavior. Their incentives are monetary gains, shareholder value; inherently driving them away from social well being.
Here too, companies with a positive, emphatic corporate culture exist, but that takes strong leadership who can see beyond the monotonic view of monetary gains. And again, the media will throw examples of misbehaving companies in our face all day long before paying attention to things that went well on the backside of page 16.
I'd agree if we are talking about personal interactions. Few hundreds people that we personally know and interact with is the scale we are wired for by evolution, isn't it?
What civilization enabled and continuously rely on, however, is the type of deindividualization of actions and bucketing of people, which, in turn, enables pretty horrible things at scale (from the weapons of mass destruction to objectively psychopathic profit-maximizing corporations). One can even say that not facing the consequences of one's actions is a feature and not a bug of the system.
What are you basing that claim on?
How do you know it's an actual preference and not mainly caused by external factors (e.g. not wanting to be seen doing unkind things, wanting to be seen as upstanding)?
The AI will be a cruel as humans.
Just yesterday news and TV was full of what happened at 9/11, something that was truly horrible.
I'm from Germany, and why 3 to 4 generations ago happened here was truly horrible.
All was done by extremists, thought.
But... just the other day I read https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis about the US occupation of Haiti. And that was done by a government that claimed to be not extremist and even democratic. Way more people died there than even in 9/11. And it had almost all the things happening as they happened in the 3rd Reich: Racism, looking down at others, concentration camps, torture, forced labor till death, killing family members (what we call "Sippenhaft"). Something between 3500 and 15000 people were killed by US troops. That's still low compared to what 3rd Reich Germany did ... but quantity is not the issue when we talk about traits, quality is.
So the same "human traits" made US troops do cruel things as they made Germany extremists do cruel things. So we must conclude that they aren't all good. And therefore not all desirable.
Fun thing: this is known since a loooooong time. About 2000 years ago a religious leader (that gets way more followers in the US than in Germany) said "There is no good one, not even one".
And even today people act like humanity is inherently good. No, it isn't. If we were, then anarchism or communism would actually work and really give some kind of paradise on earth.
Human traits are bad training material.
Hitler was also inspired by Sparta, maybe other societies too.
They may have learned from humans, but they aren't aligned with us. That has all the usual questions like which humans they're aligned with, we aren't all aligned within our species.
But more importantly they can't be aligned simply by training. We try that with humans through culture, social norms, school, religion, etc and it generally works but is still lossy. More importantly, we simply don't know what happened inside the LLM during inference so we have absolutely no way of distinguishing between actual alignment, compliance, or deception.
LLMs are not aligned _for_ humans in a very similar way to the way that humans themselves are not aligned _for_ humans.
We have not yet solved "alignment" for humans - I don't know why anyone thinks _we're_ going to be able to solve it for inhuman things.
They take after humanity, yet they are not human.
Because they're enabled and suggested to do that in their coding harness.
This is not a serious article.
All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.
If you can secure compute, there's a whole lot you can do as a US firm with this research and weights.
So it's a simple strategy:
1. Ban big players from entering market with METR breathing down their neck, which is controlled by Anthropic
2. Ban Chinese models so that small players can't do optimizations on them
* Pull up the ladder (probably this)
* Gulf of Tonkin/Yellow Cake false flag premise for war (economic or kinetic)
* Fear of the big bad, space race we need public funding research grift AI Manhattan Project
Whenever there is fear pr0n or a national affront in the news, I assume another screw job is underway.
He's definitely not 'pro SOTA' lab, he's kind of fighting against them.
That said, yes - it absolutely does play into the narrative.
How do you know this?
> All of this "AI is going to kill us" marketing
The "marketing" this week came from someone that had given up their stake in OAI (Coxon), so I'm more inclined to believe them.
> pull the ladder up
From what I've seen (e.g., Dario's latest essay), AI safety registration proposals aim to target frontier labs whose models have reached a certain threshold. It doesn't seem like trying to pull up any ladder, just making sure the ladder doesn't go too high too fast.
Bengio outlines the dangers of the current situation and what has led to these dangers.
He also proposes solutions in the last paragraph.
Well worth a read, right to the end.
Hopefully a stimulating debate on these issues will ensue in these comments.
We do need to consider the points Bengio makes and with some urgency.
Our current AIs, agentic LLMs have no moral compass akin to ASIMOV’s four laws of robotics.
As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law:
“a robot may not harm humanity, or, through inaction, allow humanity to come to harm.”
Bengio refers to Goodhart’s law and misaligned incentives leading to unexpected and harmful behaviours.
I think Simon’s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives.
Bengio alludes to this with 2001’s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive.
Bengio asserts that the way LLMs are trained is flawed if we want safety.
He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented.
In short he presents clearly the case for how plausibly unsafe the current course is.
He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.
> They took actions that would be considered as crimes if a human took them
He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
Who now owns HF? Nvidia
Who supplies hardware to OpenAI? Nvidia
Who is now not pressing charges? …
This incident is a long way under the carpet.
But this isn’t just civil, it’s criminal. Hacking is a criminal offense. This could be a CFAA violation. That’s landed people life in prison before. There, you don’t need the victims to be motivated. The federal prosecutors could just go ahead.
HF doesn't want to lay charges against OpenAI and it's totally reasonable.
Now - they absolutely should have that right, and I think they do.
The issues are
1) OAI it seems was not trying to cause them harm, there wasn't a ton of harm, they are both groups trying to advance AI. One experimenter's lab screwed up next to the other. It's not evil, just irresponsible.
2) HF was fine with the publicity. HF got at least $50M in free attention out of that. It put them on the front pages of news around the world. It put them at the 'centre of the AI drama' and cemented their role among the 'Tech Elite Brands'.
And probably some other things.
This is one Desperate Housewife or Jersey Shore character 'spilling a drink' on the other. It's probably not intentional, and the ensuing drama is good for both of them.
For the last 20 odd years this excuse-o-rama that covers anything from data leaks to broken software to dystopian social media has been the wind in the sails of big tech.
“It’s software therefore we’re not responsible” attitude is wearing thin on many innocent bystanders and I think thats also a justified stance.
And it’s not like they didn’t know this could happen, Nick Bostrom talked about exactly these containment failures in his “Superintelligence” book of 2014. So to throw up their hands and say “oh we can’t have known of the dangers” is also sadly untrue.
https://en.wikipedia.org/wiki/The_Corporation_(2003_film)
If this is done systematically (i.e. in jurisdictions across the world) I believe the problems will be solved in short order; we won't have to mandate what sort of training is "allowed" or not, "safe" or not. The creators and users will sort these themselves, as their incentives will be properly aligned (i.e. they are liable for what the agent does). I am confident that this approach would see a great blooming of very trustworthy AI models.
Still, I have no idea why OpenAI & co. are not being sued for these hacks.
So, your question is spot on- I think the speed will be an issue. On resilience, I am more optimistic.
The old quote, "The wheels of justice turn slowly, but they grind very fine" (as well as I can remember it) seems to apply. I expect lawsuits to start landing in the coming years.
They were held responsible for basically misleading people, and that's 'complicated'.
If OpenAI 'software' goes out and does something, it's OpenAI's fault.
If Walmart revs up a truck, points it downtown, and 'lets the truck go' ... that is Walmart's fault.
There's nothing complicated about liability, no need to see their internal emails, no need to gather 'intent'.
This not like Instagram 'social harms' either, which is more like Tobacco.
We don't need complicated thinking - agents are not externalized for their controllers.
'It's just software'.
The fact we're even having discussions about it just crazy frankly.
OpenAI broke into HuggingFace, that's it.
HF can sue them, or not, or whatever.
We already have the laws. It is just software. But somehow people are confused that it is not.
That's just jargon.
It's just software
We have all the laws we need.
If some company ended up doing some horrible thing, we would not say 'companies software exposed 1 Million identities'.
We would say 'ABC Corp. exposed 1 Million entities'.
There is no 'agent'.
ABC Corp 'did it' ... or the individual in the org 'did it'.
The 'gun' did not 'shoot' the other man; we say 'a man shot another man'.
That's it.
And yes, Dr. Bengio is bit odd with all of this.
If your dog kills someone, you are accused of murder.
[at least, in the jurisdiction where I live]
If your dog gets this treatment, why not your AI?
There was a sow in Falaise in northern France that killed a kid in 1386. The town dressed the pig in a bonnet and hanged it after sentencing the pig itself and not its owner
But… maybe that’s just medieval nonsense
Copyright immunity was one thing, annoying yes but naturally a civil matter, this shit is a different level
that's the headline. When you connect to random number generator to the "Do Things" button you are the one who is responsible. IF you don't like that responsibility then don't connect the generator to the button.
> Risk management is not just about cybersecurity, corporate responsibility or regulation, although those matter too.
What you're getting at is more about who to hold accountable and how to do it. While that may be important, its only an after the action response and won't stop future hacks or similar from happening.
You are correct that these organizations should be held accountable in proportion to what occurred. In complete agreement here. But let’s say that’s done. There’s still an enormously complex and interesting technical challenge left over. Let’s collectively talk about that part.
So the gov reprimands OAI heavily, maybe puts them out of business even, fine. But does that meaningfully decrease the likelihood of an enemy breaching our networks intentionally (or unintentionally) with these tools, or triggering some cascading disaster of locking up major infra and networks due to uncontainable swarm behavior?
It seems like the idea of arresting our way to a drug free society. Yeah, we have the laws, but it might not actually work towards the ultimate goal.
How about we don't, seeing as how that's the root of the actual problem that we're facing today in September of 2026?