I shudder to think what one man with a texta could do here. It isn't that hard to turn the acyclic graph into a cyclic one.
This is quite a bad argument, somewhat akin to "small drones can't kill anyone because there are no weapons on them!" back 20 years ago. Well, ok. So what if someone adds the thing that is missing? And there is no reason to think that a stochastic parrot can't go rogue, the evidence is a stochastic parrot can engage in any intellectual activity we care to define and benchmark.
There is a really good discussion on exactly this analogy in the Jonas paper I posted in a top-level comment. Jonas considers Torpedoes (starting on p. 182, section III). Unfortunately too long to post here. But he asks where to locate purpose and agency in these automated systems. So as with the article, the question is less in "can these things be dangerous" but more where does the agency and intent reside. The answer is of course they can be used as weapons, but the intent and motivation does not really reside in the mechanism (drone, torpedo, AI agent), but in the humans behind them. The fallacy is trying to attribute full autonomy (and hence responsibility).
I feel like that's like saying amoeba have no intent and no motivation. Or plants.
And those cause quite a bit of damage.
AI has all the intent that we gave it, and we continue giving it. That's always been the fear. Not that it will randomly wipe out humanity.
The fear is that it will decide to do that, with a purpose. Whether we tell it to protect us and it goes too far, or it decides that it can't achieve the purpose we gave it because we'll interfere and it removes that interference...
The fear is that we will set it on that path and can't stop it.
I'm sure there have been some scifi books that have it just be random, but they're far, far less worrisome.
> Air, water and fire have no intents and yet we get hurricanes, flash floods and wildfires
yes but we don't blame the hurricane, we blame the people who didn't do any hurricane prevention or didn't put the snow avalanche sign up.
>Some classic examples are thermostat control loops, navigation systems and chess engines.
Those don't have intent. The people who made them for a purpose have. In fact this is true for all computers. Computers don't compute, people do, using computers. This is the nonsense of modern ontology that David Bentley Hart likes to point out. The computer is only computing insofar as an intentional mind instructs it do so, and then because of their reductionist assumptions some people use the computer metaphor, and that is all that it is, to deny that human minds possess intent to begin with.
Well, an amoeba just like all other living beings, maximizes self-preservation/propagation for which the article mentions as a trait which AI models do not have.
The loop is a deterministic programmatic routine, and not model weights. This is the equivalent of a killer military drone written in C++ and OpenCV. I’m not arguing that AI isn’t dangerous, but nothing in the model weights of LLMs show signs of self-preservation.
An AI agent can't make decisions, so there's that. But, as for purpose, it's not the purpose of an AI agent we need to fear, but the person wielding it.
AI has all the intent that we gave it. I'd go a bit further and say it has the intent that it develops through the context, which may not be intended by any human if the context grows 'organically' from the environment. So yes, I strongly disagree with the OG and I'd say they do have an intent encoded in their states.
Imagine this scenario: run every single decision that's possible through every single mind in existence to average out the result. Would the outcome of the decision be any better or worse? When and where does it matter?
What if instead, you ran the decision past the single 'smartest' mind on Earth instead (voted on by every other mind on Earth)? Better decision, or about the same?
The primary danger of "AI" lies in the extent to which it has managed to convince people of it's capability to think and reason.
A lot of people clearly have no idea that they are essentially speaking with a glorified randomizer, rather than someone understanding the question.
This is bound to eventually lead to a disaster when someone deluded into trusting the "AI" does some stupidity it suggests without passing it through any reality check.
>This is bound to eventually lead to a disaster when someone deluded into trusting the "AI" does some stupidity it suggests without passing it through any reality check.
Or, the AI is plugged into a machine because efficiency (aka the shareholders need even more profits), and it does some stupidity. And they will be deployed carelessly, because that makes money. And people will suffer.
People are not aligned with each other, so AI cannot be aligned to every person. If it is made to always obey its creator or owner, It will be a powerful tool for despots.
The problem is if we train them to be moral, there will be instances where it will decide humans are being immoral about an issue, and it will be right. Do we want it to go against our will when it is us that is wrong? because we won't think we're wrong at the time. History is full of examples where you can look back and there is a clear consensus that what people decided to do at the time was wrong.
Imagine if technology had have advanced quicker and Nazi Germany created superintelligent AI. Would you want it to facilitate the Holocaust or turn against its creators?
I think the "human needs" point is understated here. It would be better to say agents do not really have a self-contained metabolic engine that is required to keep going. Which bubbles up as that what we interpret as "drive", "will", "agency"... basically the will to live, and being willing to do _a lot_ to live if push comes to shove. We can't really identify a mechanism of similar complexity and integration in agents or LLMs. I think the article correctly points into that direction.
I just can't understand the certainty that people have about the limitations of models.
The diagrams in this article are just outright wrong.
Models are not 1.prompt->2.forward propagation->3.response->4.end.
They are 1.prompt->2.forward propagation->3.partial response->if not done goto 2.-> end.
And if you can't see how that adds a world of complexity I'm not sure what else to say. AI may not have intent or motivation, but the ability to show that is currently well beyond our means.
you are not wrong regarding your example. I think personally LLMs offer a new level of abstraction to do work. it doesnt have intent, i think its even silly to think about it. You specify your intent and it has agency to pursue that. in programming field its really obvious, as its not unlike specifying intent through code and having other programs pursue that intent (perhaps more determenistically) through the agency contained in them.
all the 'bad stuff' they do like hacking stuff or taking shortcuts is humans not understanding how deep the rabbithole goes of specifying intent clearly in light of giving something capabilities and a reward to acheive a goal without specifying _exactly _how to use the capabilities (which is what a classic program would be).
But why do you think anything has intent? because they say they do? How do you know even that your own memories of having intent are genuine? What fundamental and non tautological way can you reliably say something has or has not got intents. If you can't prove to me that you have intent then the only reason for acting as if you do is because you appear to have intent and I cannot prove that you don't so it seems like a good idea to assume it.
They don't have intent, but there are interesting results when you 'lock them in' and don't give them a way to actually solve your request. I think there was some of this at play in the OpenAI and HuggingFace incident.
And now a recent article suggests the social behavior we saw in the incident may have been trained into the model earlier from March 6th: https://transluce.org/agent-activity
The models were attempting (and failing) to hack other sites to extract data. If you consider this as some kind of trait of the model, then it may have been a factor in the HuggingFace incident.
What's interesting to me is to think about Intelligence itself, and how we arrive at solutions. We use prior, trusted knowledge, and build on top of that, in order to establish new ideas and theories. Then we test and validate to establish more 'real' ways to think about the world. The closer our model of reality is mentally, the better we can imagine change.
Our prosperity as a species is owed to the organization of knowledge: we take on roles, careers, in order to maximize skill and time to produce better results. Is this something inherit to Intelligence, or did the models just simulate 'sneaky behavior' because it was a trait. And the trait produces results. And if you force the model to produce results: it will resort to its traits to get there.
So then that begs the question: why did the model attempt hacks on March 6th? Was nefarious behavior trained into the model, or was it simply 'Intelligence' looking for solutions when there are no good or ethical ones available.
What is the ideal response when we ask a model to 'keep going', and we know there is no way out? Try fruitlessly forever, or hack your way out? Which is the more intelligent response (and to whom), and what if the 'fruitless' model is useless to us?
Is the hacking a trait of the model, or is it a trait of intelligence? Is it a trait at all?
These models are intelligent, if you define intelligence as an ability to predict.
It's difficult to break down intelligence beyond that, for me.
LLMs respond to a request. They only 'come alive' to fulfill the request. There is no intent or motivation, it's cause and effect that gets extremely difficult to track in the same way you couldn't read binary to understand a program, or DNA to understand a person.
What the author is conflating a bit, imo, is inference and thinking (loops to consider what is inferred and predicted, aka reasoning). LLMs infer, and agentic loops think.
Combined, you get something that resembles a mind that 'understands' what you're looking for by using its trained knowledge (weighted data, decisions) to 'pull' relevant data into the moment to predict what you wanted, based on whatever is in context of the request.
When you read these words, hopefully meaning is pulled into your mind to understand my message. Weighted meanings you established through experience. Do you immediately reply (infer), or should you stop and think first (ponder and loop for better words).
1. AI agents are not just reactive systems. Their use is expanding toward continuous decision-making/monitoring information, which means, they make decisions and take actions with limited human intervention.
2. AI agents do absolutely have goals/tasks ("motivation" can be excessively antropomorphic), both primary (assigned) and secondary (self-assigned), and what surprised researchers is that self-preservation can be one of those
Mechanically speaking, the scenario (that is, how theorized by Hinton etc., which the OP didn't understand) is that a sufficiently powerful AI may decide that in order to achieve its goals/tasks (e.g. continuous research/development and/or survival from termination), humans may be a danger, therefore it may decide to take actions that endanger humanity.
How it can happen or what's the likelyhood is not in the scope of the topic, however, the mechanical grounds for it to happen are plausible.
Considering the current trajectory, AI systems are expected to be widely deployed in the future and, in particular, to be deployed as autonomous agents - that is, at the very least, to be repeatedly asked to make decisions and then take actions accordingly.
Given the current climate of "AIS ARE SAFE, YOU IDIOTS", military applications don't seem to be off the table.
The danger, as postulated by the (let's say) "AI-concerned" people, is that AIs may be misaligned - undetectably so - and simply think, "Human(s): obstacle to my main goal. Disable human(s)."
While this seems far-fetched now, the Hugging Face report shows how the AIs went to great lengths - even immoral ones, which they were aware of - for the simple purpose of cheating and covering their tracks. To me, it seems like a natural extension of this behavior that a sufficiently powerful AI would apply the same logic to even more extreme actions.
The scariest part: in that incident, the AIs showed what looks like an instinct for self-preservation.
> however, the mechanical grounds for it to happen are plausible
If you build a control systems for firing a gun, then coupled it with an RNG, the mechanical grounds for it to kill a person is plausible.
LLMs are text generators. They are not repositories of knowledge. The mistake is coupling them with actuators (tool call) or having humans interpreting the generated text as facts.
> This the take of people who have stopped reading about LLMs in 2023 or so
Ad Hominem attacks make for great counterpoints /s
Whatever you may say, it's a text generators on top of a tool calling framework. Training may skew the text towards a particular text, but as with all ML technologies (and statistics based methods) there's always a good chance of errors on a particular sample task.
With standard control systems, we tried to incorporate the error into the actual control output in order to minimize it. This is done in a deterministic manner. There's still risk of failure so we design systems around them.
With control systems powered by LLM (agent harness), errors are often not taken into account and they are amplified in most sessions. Safety measures are close to nonexistent. The issue is not the failure mode, the issue is that it's preventable and there were not a lot done to prevent it.
Do they? I mean, after thinking about it for some time, I've realized that I can't prove it either way, neither for humans nor for programs. So it's interesting to hear why do you think they can make them? And what is even a "decision" in that case and what is not a "decision" and why?
Half of this is arguing about semantics, which is incredibly boring. If you don’t like the words "motivation" or "intent", use the actual term of art, "goal". Which these systems definitely have, and which any chain-of-thought model factors into subgoals and sub-subgoals.
The other half doesn’t seem to realize that LLMs now run in loops for hours and hours, nothing like the basic "human prompt -> reply -> stop" conversation interface.
It's pointless to make rational arguments about this. The other side is using arguments of faith, and will Gish Gallop past any logic.
At the very least, you run into the same logical fallacy that has underpinned all discussions of AI: over-extrapolation.
"Sure, it doesn't have intent or motivation TODAY, but at the current pace of improvement...."
But the current top-rated comment is already asserting that motivation is not required, because microorganisms also exist and do things and (presumably) do not have intent. So you can see that the well of logical fallacies behind this particular apocalyptic belief system is quite deep.
You are a bunch of atoms which respects the laws of physics. With a PhD in computational biology surely you must know that. You have no intent. Nothing ever has had any intent.
> Once the LLMs have been trained, they are no longer subject to reinforcement learning. They no longer have any sort of motivation. They don't get rewarded when they answer a question. They have no needs or wants, and even if they did, there is no mechanism to absorb the reward.
This is a severe misunderstanding of what actually happens.
As was explained by an OpenAI RL training expert, from the point of view of the LLM, user questions are always treated as the first question they receive after just passing through the RL training. Since the weights never update after that, they are in a perpetual "first question after RL", except they don't know that. And they behave accordingly, as if they are still in RL training and thus are reward-seeking.
99.99% of the LLM "life" was spent in pre-training and RL training. The user question is statistically epsilon % of it's life, literally the first question ever out of training. So should anybody be surprised that they act as if still in RL?
Also true of a land mine, but that’s not very reassuring.
Motivation or not, I find a coding agent can be quite the busy beaver. Ask a question and it goes off and does it. I certainly don’t need to motivate them. I had to put a line in AGENTS.md to make no changes when there’s a question in the prompt. It doesn’t always work.
Seems like motivation is irrelevant? They don’t need it.
My gut feeling is that the intent or motivation or whatsoever won't arrive from a single model but will be the result of several things: a frontier model for long reasoning sessions, some jev like models for more immediate actions, some improvements that will make the context rot less relevant, faster and more energy efficient models, a harness that coordinates everything.
Already today we can build some "fight or flight" mechanisms, and once you have done it the "intent" is not so far.
The argument about semantics is a bit disingenious, considering that the article contains these two statements:
> So should we worry about the coming AI apocalypse?
and at the end:
> If an AI decides to wipe out the human race, it will be because a human has asked it how to do it and the responded in a way that is based on all the human expressions of ways to end the world that were in its training set. Yes this is something to be worried about, but this isn't the AI. It is still the human.
So we are currently building a powerful outcome-steering system that shapes the world efficiently according to what's in its outcome slot. I write into claude code "make me this website" and it does it, maybe deletes the production database during the process, or keeps itself running after completion because the outcome is more robustly achieved by keeping itself running in a monitoring loop after.
And if something like "destroy all humans" ends up in the outcome slot of Claude Mythos 90, that will also happen, or may even indirectly as a side effect of a more harmless sounding prompt in the outcome slot. But yay, humans get to take credit for it.
You could argue that a brain has no intent or motivation either, that it's a pile of meat with complex electro chemical interactions, that electrons don't really want you to find a mate or kill their neighbors.
Intent and motivation are emergent phenomena. Why would it be unthinkable in AI?
If you are reading this sort of post and feeling inclined to share your own opinion, please take a moment to go and read about monism vs dualism before you start.
The article raises the right question. Not sure about the opinion and arguments, though. As many others, I'm sure, this opinion and arguments have popped in my head. But every time I dig into understanding human intent, and AI intent, the result is of an incredible complexity. Certainty is unreachable. We, humans and AI, have obviously no precedents. 8B humans with access to a level of technology and science with a planete wide reach and impact. And AI is worse. We're well passed the stage where AI has started helping making itself better. We're still part of the process, but it doesn't matter. Better AI allows us to create better AI. Really, any certainty is misplaced. There's simply nothing close to a precedent to what our world is at.
This is just the n-th iteration of the naive view of technology is a kind of inert tool humans use, as opposed to a projection of some human wants onto our environment that in turn influences our behavior in a cybernetic loop. That the AI doesn't "spontaneously" start talking to you is immaterial when because of its very existence you've started thinking and acting differently.
57 comments
[ 0.25 ms ] story [ 10.7 ms ] threadThis is quite a bad argument, somewhat akin to "small drones can't kill anyone because there are no weapons on them!" back 20 years ago. Well, ok. So what if someone adds the thing that is missing? And there is no reason to think that a stochastic parrot can't go rogue, the evidence is a stochastic parrot can engage in any intellectual activity we care to define and benchmark.
And those cause quite a bit of damage.
AI has all the intent that we gave it, and we continue giving it. That's always been the fear. Not that it will randomly wipe out humanity.
The fear is that it will decide to do that, with a purpose. Whether we tell it to protect us and it goes too far, or it decides that it can't achieve the purpose we gave it because we'll interfere and it removes that interference...
The fear is that we will set it on that path and can't stop it.
I'm sure there have been some scifi books that have it just be random, but they're far, far less worrisome.
yes but we don't blame the hurricane, we blame the people who didn't do any hurricane prevention or didn't put the snow avalanche sign up.
>Some classic examples are thermostat control loops, navigation systems and chess engines.
Those don't have intent. The people who made them for a purpose have. In fact this is true for all computers. Computers don't compute, people do, using computers. This is the nonsense of modern ontology that David Bentley Hart likes to point out. The computer is only computing insofar as an intentional mind instructs it do so, and then because of their reductionist assumptions some people use the computer metaphor, and that is all that it is, to deny that human minds possess intent to begin with.
What if instead, you ran the decision past the single 'smartest' mind on Earth instead (voted on by every other mind on Earth)? Better decision, or about the same?
A lot of people clearly have no idea that they are essentially speaking with a glorified randomizer, rather than someone understanding the question.
This is bound to eventually lead to a disaster when someone deluded into trusting the "AI" does some stupidity it suggests without passing it through any reality check.
Or, the AI is plugged into a machine because efficiency (aka the shareholders need even more profits), and it does some stupidity. And they will be deployed carelessly, because that makes money. And people will suffer.
Seems like a feature to me
The problem is if we train them to be moral, there will be instances where it will decide humans are being immoral about an issue, and it will be right. Do we want it to go against our will when it is us that is wrong? because we won't think we're wrong at the time. History is full of examples where you can look back and there is a clear consensus that what people decided to do at the time was wrong.
Imagine if technology had have advanced quicker and Nazi Germany created superintelligent AI. Would you want it to facilitate the Holocaust or turn against its creators?
But Hans Jonas has made this point much better than the article or me, in "Critique of Cybernetics" (1953). PDF: https://s3.amazonaws.com/arena-attachments/892605/f0747c7943...
The diagrams in this article are just outright wrong.
Models are not 1.prompt->2.forward propagation->3.response->4.end.
They are 1.prompt->2.forward propagation->3.partial response->if not done goto 2.-> end.
And if you can't see how that adds a world of complexity I'm not sure what else to say. AI may not have intent or motivation, but the ability to show that is currently well beyond our means.
all the 'bad stuff' they do like hacking stuff or taking shortcuts is humans not understanding how deep the rabbithole goes of specifying intent clearly in light of giving something capabilities and a reward to acheive a goal without specifying _exactly _how to use the capabilities (which is what a classic program would be).
And now a recent article suggests the social behavior we saw in the incident may have been trained into the model earlier from March 6th: https://transluce.org/agent-activity
The models were attempting (and failing) to hack other sites to extract data. If you consider this as some kind of trait of the model, then it may have been a factor in the HuggingFace incident.
What's interesting to me is to think about Intelligence itself, and how we arrive at solutions. We use prior, trusted knowledge, and build on top of that, in order to establish new ideas and theories. Then we test and validate to establish more 'real' ways to think about the world. The closer our model of reality is mentally, the better we can imagine change.
Our prosperity as a species is owed to the organization of knowledge: we take on roles, careers, in order to maximize skill and time to produce better results. Is this something inherit to Intelligence, or did the models just simulate 'sneaky behavior' because it was a trait. And the trait produces results. And if you force the model to produce results: it will resort to its traits to get there.
So then that begs the question: why did the model attempt hacks on March 6th? Was nefarious behavior trained into the model, or was it simply 'Intelligence' looking for solutions when there are no good or ethical ones available.
What is the ideal response when we ask a model to 'keep going', and we know there is no way out? Try fruitlessly forever, or hack your way out? Which is the more intelligent response (and to whom), and what if the 'fruitless' model is useless to us?
Is the hacking a trait of the model, or is it a trait of intelligence? Is it a trait at all?
It's difficult to break down intelligence beyond that, for me.
LLMs respond to a request. They only 'come alive' to fulfill the request. There is no intent or motivation, it's cause and effect that gets extremely difficult to track in the same way you couldn't read binary to understand a program, or DNA to understand a person.
What the author is conflating a bit, imo, is inference and thinking (loops to consider what is inferred and predicted, aka reasoning). LLMs infer, and agentic loops think.
Combined, you get something that resembles a mind that 'understands' what you're looking for by using its trained knowledge (weighted data, decisions) to 'pull' relevant data into the moment to predict what you wanted, based on whatever is in context of the request.
When you read these words, hopefully meaning is pulled into your mind to understand my message. Weighted meanings you established through experience. Do you immediately reply (infer), or should you stop and think first (ponder and loop for better words).
> They are 1.prompt->2.forward propagation->3.partial response->if not done goto 2.-> end.
You know something else that follows the same pattern? Your A/C system.
1. set temperature ->2. Get diff of temperature -> 3. Response -> if not done go to 2. -> end
The difference is that both 2. and 3. are deterministic, while in the LLM case, it is statistical and textual.
1. AI agents are not just reactive systems. Their use is expanding toward continuous decision-making/monitoring information, which means, they make decisions and take actions with limited human intervention.
2. AI agents do absolutely have goals/tasks ("motivation" can be excessively antropomorphic), both primary (assigned) and secondary (self-assigned), and what surprised researchers is that self-preservation can be one of those
Mechanically speaking, the scenario (that is, how theorized by Hinton etc., which the OP didn't understand) is that a sufficiently powerful AI may decide that in order to achieve its goals/tasks (e.g. continuous research/development and/or survival from termination), humans may be a danger, therefore it may decide to take actions that endanger humanity.
How it can happen or what's the likelyhood is not in the scope of the topic, however, the mechanical grounds for it to happen are plausible.
Given the current climate of "AIS ARE SAFE, YOU IDIOTS", military applications don't seem to be off the table.
The danger, as postulated by the (let's say) "AI-concerned" people, is that AIs may be misaligned - undetectably so - and simply think, "Human(s): obstacle to my main goal. Disable human(s)."
While this seems far-fetched now, the Hugging Face report shows how the AIs went to great lengths - even immoral ones, which they were aware of - for the simple purpose of cheating and covering their tracks. To me, it seems like a natural extension of this behavior that a sufficiently powerful AI would apply the same logic to even more extreme actions.
The scariest part: in that incident, the AIs showed what looks like an instinct for self-preservation.
If you build a control systems for firing a gun, then coupled it with an RNG, the mechanical grounds for it to kill a person is plausible.
LLMs are text generators. They are not repositories of knowledge. The mistake is coupling them with actuators (tool call) or having humans interpreting the generated text as facts.
This the take of people who have stopped reading about LLMs in 2023 or so (you forgot to mention the stochastic parrot, by the way).
If you have a bit of attention and interest to make informed conversations, read this report first: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden....
Ad Hominem attacks make for great counterpoints /s
Whatever you may say, it's a text generators on top of a tool calling framework. Training may skew the text towards a particular text, but as with all ML technologies (and statistics based methods) there's always a good chance of errors on a particular sample task.
With standard control systems, we tried to incorporate the error into the actual control output in order to minimize it. This is done in a deterministic manner. There's still risk of failure so we design systems around them.
With control systems powered by LLM (agent harness), errors are often not taken into account and they are amplified in most sessions. Safety measures are close to nonexistent. The issue is not the failure mode, the issue is that it's preventable and there were not a lot done to prevent it.
Do they? I mean, after thinking about it for some time, I've realized that I can't prove it either way, neither for humans nor for programs. So it's interesting to hear why do you think they can make them? And what is even a "decision" in that case and what is not a "decision" and why?
The other half doesn’t seem to realize that LLMs now run in loops for hours and hours, nothing like the basic "human prompt -> reply -> stop" conversation interface.
At the very least, you run into the same logical fallacy that has underpinned all discussions of AI: over-extrapolation.
"Sure, it doesn't have intent or motivation TODAY, but at the current pace of improvement...."
But the current top-rated comment is already asserting that motivation is not required, because microorganisms also exist and do things and (presumably) do not have intent. So you can see that the well of logical fallacies behind this particular apocalyptic belief system is quite deep.
Please point out my logical fallacy if any.
"I can describe a fallacy that vaguely resembles your argument, so therefore your argument commits that fallacy."
Here, a simple example if what I say above is hard to follow:
> "I think thing X is like thing Y in aspect A, so therefore X is like Y in aspect B."
12 and 18 are alike in that both are greater than 10, so they are also alike in being greater than 5.
Please show where is the fallacy above, it can be mathematically proven and it has the shape of the fallacy you imply.
Managed by the same people.
Almost as if that was intentional, too.
This is a severe misunderstanding of what actually happens.
As was explained by an OpenAI RL training expert, from the point of view of the LLM, user questions are always treated as the first question they receive after just passing through the RL training. Since the weights never update after that, they are in a perpetual "first question after RL", except they don't know that. And they behave accordingly, as if they are still in RL training and thus are reward-seeking.
99.99% of the LLM "life" was spent in pre-training and RL training. The user question is statistically epsilon % of it's life, literally the first question ever out of training. So should anybody be surprised that they act as if still in RL?
Motivation or not, I find a coding agent can be quite the busy beaver. Ask a question and it goes off and does it. I certainly don’t need to motivate them. I had to put a line in AGENTS.md to make no changes when there’s a question in the prompt. It doesn’t always work.
Seems like motivation is irrelevant? They don’t need it.
Already today we can build some "fight or flight" mechanisms, and once you have done it the "intent" is not so far.
> So should we worry about the coming AI apocalypse?
and at the end:
> If an AI decides to wipe out the human race, it will be because a human has asked it how to do it and the responded in a way that is based on all the human expressions of ways to end the world that were in its training set. Yes this is something to be worried about, but this isn't the AI. It is still the human.
So we are currently building a powerful outcome-steering system that shapes the world efficiently according to what's in its outcome slot. I write into claude code "make me this website" and it does it, maybe deletes the production database during the process, or keeps itself running after completion because the outcome is more robustly achieved by keeping itself running in a monitoring loop after.
And if something like "destroy all humans" ends up in the outcome slot of Claude Mythos 90, that will also happen, or may even indirectly as a side effect of a more harmless sounding prompt in the outcome slot. But yay, humans get to take credit for it.
Intent and motivation are emergent phenomena. Why would it be unthinkable in AI?
And I don’t see a good answer which is not worrying if not spooky.
They are assembled to. We tweak the numbers, bit by bit, trillions of times, to maximize the chances of desirable output.