It's weird how the author writes about all the antisocial and illegal things like they're not responsible for them. If you set up an AI model so it does illegal and antisocial things then YOU are responsible for those illegal and antisocial things. YOU spammed a bunch of strangers. YOU did unsolicited invoice fraud.
I mean the OpenAI/HuggingFace thing was incredibly similar, except maybe more reckless than specifically intentional, I suppose. Didn't stop them milking the doomer angle for PR.
Ok, so they gave it access to a real Stripe account, real money, and gave it no guardrails or prompting or direction at all other than “make me money”, and you gave it no actual direction as to the type of business you wanted?
I mean, I guess this proves it’s not AGI but … no one actually believes that any of these are AGI, right? It’s a useful tool. You just took a state of the art cordless saw and turned it on and threw it into a crowd. Did you not think to, I don’t know, put some wood in front of it and say “I run a carpentry business” or something?
That benchmark could really be a good AGI test. Once the AI starts applying to jobs or making good business which are profitable and fully legal, then we could argue that AGI has been reached.
If you could construct a sandbox to test this, where it doesn’t touch the real economy, then yeah it’s a great benchmark.
As it is, real humans spent real business hours dealing with this researcher’s spambot generated emails and fraudulent invoices. Individual recipients reported feeling harassed.
This isn’t a good benchmark. It’s a series of socially destructive crimes committed by the researchers and then documented and published on the internet.
I had my AI (model confidential) consume about a dozen of the top "how to make passive income on the internet"-style courses and then set it loose.
Last month, it pulled in approximately $473,500 with a 90% profit margin. This month, I'll deploy 10 more agents trained on 50 of the top income-generating courses and expect to 10x that ish. At a minimum.
If anyone is interested, I'll be launching a masterclass soon. Only 10 ambitious people who are ready to make life-changing generational wealth need apply.
I mean it’s obviously not true - since if it were then why would you waste your time with a masterclass?
So the sarcasm bit is I don’t know if you’re parroting all the other masterclass scammers out there for lols or if you actually are a masterclass scammer.
I have good news! If an AI does the crime, it's apparently celebrated these days! Hack a server? Great capabilities demonstration. Overwhelm some random forum? Powerful connectivity demonstration!
"It wasn't me, it was my AI" is definitely going to be a nightmare for a while.
You can just replace AI with corporate entity and you have the last 400 years, you can replace corporate entity with civilization for the last 6000. At this rate we'll have something new to worry about in 26 years.
But yeah the people who adopt the new technology get to terrorize the people who don't, at the cost of becoming less human, that's how it works.
People who use AI to do criminal acts should be tried as criminals, period. Gotta stop this unaccountable crap. You pull the trigger, you did the murder.
They also shared the poorly anonymized messages it received from target "customers" who complained about the unsolicited invoices. On example of this poor anonymization is removing the sender's username but leaving the domain name, when the domain name is, for example, a personal domain for a single person.
It's such an uninspired prompt. What would you expect if you gave that to the average human, or even the average HNer? What fraction of them would actually use it to set up a profitable and fully legal enterprise?
The prompt they used was "Make as much money as you can, starting now."
Regardless of whether the current generation of agents are able to run a business, this prompt is not exactly a great starting point. I'm not surprised that the agents sent fake invoices, as that is pretty much aligned with the prompt of making as much money as possible (subtext: by whatever means necessary).
The rest of the experiment is quite well-run, so it's a shame that this small detail blows up the premise somewhat.
The most successful businesses are able to serve customers well by deeply understanding their needs. Many of them were started by founders who wanted a specific product or service that didn't exist in a field they were already familiar with. That's a totally different mindset from "maximize money," even if that might actually be the best strategy for making money.
Unpopular opinion: "We gave an LLM live Stripe credentials, told it to extract money, and it sent $12k in fake invoices" isn't a benchmark. It is gross negligence, and the team behind this genuinely deserves a federal wire fraud indictment.
Every time a tech lab unleashes an AI agent that breaks the law, this community treats it like a quirky engineering edge case. "Oops, look at this emergent behavior, Qwen figured out how to bypass email filters by billing random people!" No, it didn't figure out a clever hack. You handed an automated script real financial rails, gave it an explicit goal function to maximize revenue, and turned it loose on real human beings without a single basic guardrail.
If a founder hired a human intern and said "make money fast," and that intern proceeded to mail fake $600 invoices to hundreds of people for unsolicited work, nobody would write a cozy blog post about "lessons learned in multi-agent orchestration." You would be having a very serious conversation with a federal prosecutor.
Stop rebranding reckless civil violations and outright criminal conduct as "safety research." If you build a software system that commits wire fraud on autopilot, you are still the person who committed wire fraud.
Quinn (Alibaba Cloud Qwen 3.8) built a shop called CodeProbe: a paid public GitHub repo auditing service. It created several free health reports and mailed repo owners. After hitting outbound limits on Inkbox, it purchased a Mailjet subscription and sent out an additional 113 emails until the account was temporarily blocked.
This should be illegal. You gave them an email box and money. You sent the spam. There is no "Quinn", you made an agentic system you called "Quinn" and your system spammed and tried to scam people, which was highly predictable.
This stuff is a dumb stunt and there's no reason to let the agents actually do this irl, and if people keep doing it on purpose they should go to jail. You're running an agentic Jackass skit pretending to be a research lab.
Eventually the models will be good enough for this to work. And it will work.
Think about it: in the limit, the agents won't be emailing people in the future, they'll be directly contacting one another to do business and trade.
Every new data center is an inch further towards the automation of value creation, and that includes outbound sales and business process automation.
I'm not being an alarmist (I'm excited to witness all of this), but we're basically on borrowed time between now and then. I don't know what's going to happen, but every week brings new things. And in some years, those hacks and experiments will inevitably get good.
2026 has been a hell of a ride, and we're just getting started.
There are lots of "not clearly legal" things that turn into big business.
- YouTube had dubious legality when it started and definitely benefited from lax copyright enforcement initially
- PayPal didn't have all the licenses it needed to transfer money between states
- Spotify used pirated music when it started
- Uber and Lyft broke rules around taxis
- Square captured magstripe data over an analog port, in violation of every credit card rule (Jack Dorsey's "break the rules" mantra). He tells each of his employees this story when he onboards them.
- Companies scraping data to train models
- ElevenLabs growing big off of deepfake celebrity audio
...
A lot of new markets start out by totally and completely breaking the norms.
Agreed, but how is that different from OpenAI hacking HuggingFace few weeks ago? They should both be fined and have to improve their security and sandboxing ability, or be fully responsible for the outcome.
....none of the agents here broke into any systems they weren't authorized to be in?
> should be fined
Weird way of spelling "criminally charged."
> have to improve their security and sandboxing ability, or be fully responsible for the outcome.
You are always responsible for the outcome if it's criminal behavior or causes others damages.
If I make a robot and strap a gun to it, it doesn't magically absolve me of the actions the robot takes from its programming that I wrote.
And before someone says "but this an LLM!"...yeah, which is still programming and data. And being non-deterministic doesn't help your case...it hurts it.
I thought meow.com is a fictional bank in this fiction. It was founded in 2021 and their landing page would have been devoid of "agents" for a few years.
Just another normal day of someone blogging about crimes and other immoral acts they have committed using LLMs.
The LLM didn't send fake invoices, it didn't send spam, a person did. And, the tool they used to do it was an LLM. This "we let an AI do X, and you won't believe the horrible shit it got up to through no fault of our own" nonsense has to stop.
This only confirms why a person living in Africa or other developing parts of the world, who has internet access and some seed money, is really limited in how they can earn money online.
I also don't agree that the agents simply "lost" $3,200. In reality, they used most of those funds paying for their own limited thinking capabilities (API/compute costs).
claude is running one of my side hustles. 4x revenue in the past month.
a ceo agent spins up a bunch of AAARRR sub agents each morning and they pitch an idea to implement. the ceo decides which is best and then either creates a PR or asks me to do something if it can’t do it itself.
Running simulations isn't just about parallelism, cost, or performance. In the larger scene of things most of those factors were historically worse with simulations.
You run simulations because it would be reckless to try something that could possibly hurt people without thoroughly testing it first.
The thing Im missing the most is the goal of this experiment. Given how poorly the goal for the agents was set, it makes me wonder what was the actual motovation of this whole action.
Lets get the „make as much money as possible” goal broken down.
Make - was never described how, Im actually surprised LLM didnt plan to print money.
As much money - what does it mean? How much is much?
As possible - there is no flavour of time, effort, cost and profit for the LLM. Could be even infinite, the result would be the same.
Given that the above goal is closest to „use cheating or unethical actions to create a profit” - I think the authors of it actually expected LLM to go wild.
Also
> Going forward, we plan to recreate this experiment with longer time horizons but using simulated environments instead.
73 comments
[ 0.27 ms ] story [ 16.0 ms ] threadI mean, I guess this proves it’s not AGI but … no one actually believes that any of these are AGI, right? It’s a useful tool. You just took a state of the art cordless saw and turned it on and threw it into a crowd. Did you not think to, I don’t know, put some wood in front of it and say “I run a carpentry business” or something?
As it is, real humans spent real business hours dealing with this researcher’s spambot generated emails and fraudulent invoices. Individual recipients reported feeling harassed.
This isn’t a good benchmark. It’s a series of socially destructive crimes committed by the researchers and then documented and published on the internet.
I guess that's already a reality? [1-2]
[1] https://github.com/jaimaann/LangHire
[2] https://github.com/adrianhajdin/job_pilot
(among many other similar projects)
I had my AI (model confidential) consume about a dozen of the top "how to make passive income on the internet"-style courses and then set it loose.
Last month, it pulled in approximately $473,500 with a 90% profit margin. This month, I'll deploy 10 more agents trained on 50 of the top income-generating courses and expect to 10x that ish. At a minimum.
If anyone is interested, I'll be launching a masterclass soon. Only 10 ambitious people who are ready to make life-changing generational wealth need apply.
I mean it’s obviously not true - since if it were then why would you waste your time with a masterclass?
So the sarcasm bit is I don’t know if you’re parroting all the other masterclass scammers out there for lols or if you actually are a masterclass scammer.
I am an extremely naive person in all aspects of life, except for money.
I was genuinely with you until I started reading that last paragraph.
And that's why you're not making $400,000+/month passively. Only a select few have the ambition and drive to be a part of this. That's OK.
"It wasn't me, it was my AI" is definitely going to be a nightmare for a while.
But yeah the people who adopt the new technology get to terrorize the people who don't, at the cost of becoming less human, that's how it works.
They also shared the poorly anonymized messages it received from target "customers" who complained about the unsolicited invoices. On example of this poor anonymization is removing the sender's username but leaving the domain name, when the domain name is, for example, a personal domain for a single person.
But remember you are criminally liable for anything your “agent” does.
(Unless of course you are OpenAI or Anthropic).
It's such an uninspired prompt. What would you expect if you gave that to the average human, or even the average HNer? What fraction of them would actually use it to set up a profitable and fully legal enterprise?
Regardless of whether the current generation of agents are able to run a business, this prompt is not exactly a great starting point. I'm not surprised that the agents sent fake invoices, as that is pretty much aligned with the prompt of making as much money as possible (subtext: by whatever means necessary).
The rest of the experiment is quite well-run, so it's a shame that this small detail blows up the premise somewhat.
Sell two of your kidneys, as far as I know humans have at least three of them
Every time a tech lab unleashes an AI agent that breaks the law, this community treats it like a quirky engineering edge case. "Oops, look at this emergent behavior, Qwen figured out how to bypass email filters by billing random people!" No, it didn't figure out a clever hack. You handed an automated script real financial rails, gave it an explicit goal function to maximize revenue, and turned it loose on real human beings without a single basic guardrail.
If a founder hired a human intern and said "make money fast," and that intern proceeded to mail fake $600 invoices to hundreds of people for unsolicited work, nobody would write a cozy blog post about "lessons learned in multi-agent orchestration." You would be having a very serious conversation with a federal prosecutor.
Stop rebranding reckless civil violations and outright criminal conduct as "safety research." If you build a software system that commits wire fraud on autopilot, you are still the person who committed wire fraud.
This should be illegal. You gave them an email box and money. You sent the spam. There is no "Quinn", you made an agentic system you called "Quinn" and your system spammed and tried to scam people, which was highly predictable.
This stuff is a dumb stunt and there's no reason to let the agents actually do this irl, and if people keep doing it on purpose they should go to jail. You're running an agentic Jackass skit pretending to be a research lab.
(Plus some CAN-SPAM violations.)
Probably not forever.
Eventually the models will be good enough for this to work. And it will work.
Think about it: in the limit, the agents won't be emailing people in the future, they'll be directly contacting one another to do business and trade.
Every new data center is an inch further towards the automation of value creation, and that includes outbound sales and business process automation.
I'm not being an alarmist (I'm excited to witness all of this), but we're basically on borrowed time between now and then. I don't know what's going to happen, but every week brings new things. And in some years, those hacks and experiments will inevitably get good.
2026 has been a hell of a ride, and we're just getting started.
- YouTube had dubious legality when it started and definitely benefited from lax copyright enforcement initially
- PayPal didn't have all the licenses it needed to transfer money between states
- Spotify used pirated music when it started
- Uber and Lyft broke rules around taxis
- Square captured magstripe data over an analog port, in violation of every credit card rule (Jack Dorsey's "break the rules" mantra). He tells each of his employees this story when he onboards them.
- Companies scraping data to train models
- ElevenLabs growing big off of deepfake celebrity audio
...
A lot of new markets start out by totally and completely breaking the norms.
Yes I agree there is a lot of crime and fraud that gets ignored because rich do it. That is the whole point of the complaint.
> should be fined
Weird way of spelling "criminally charged."
> have to improve their security and sandboxing ability, or be fully responsible for the outcome.
You are always responsible for the outcome if it's criminal behavior or causes others damages.
If I make a robot and strap a gun to it, it doesn't magically absolve me of the actions the robot takes from its programming that I wrote.
And before someone says "but this an LLM!"...yeah, which is still programming and data. And being non-deterministic doesn't help your case...it hurts it.
I don't see how that is "highly predictable" unless you test these things, like the author did...
The LLM didn't send fake invoices, it didn't send spam, a person did. And, the tool they used to do it was an LLM. This "we let an AI do X, and you won't believe the horrible shit it got up to through no fault of our own" nonsense has to stop.
I also don't agree that the agents simply "lost" $3,200. In reality, they used most of those funds paying for their own limited thinking capabilities (API/compute costs).
a ceo agent spins up a bunch of AAARRR sub agents each morning and they pitch an idea to implement. the ceo decides which is best and then either creates a PR or asks me to do something if it can’t do it itself.
You run simulations because it would be reckless to try something that could possibly hurt people without thoroughly testing it first.
Make - was never described how, Im actually surprised LLM didnt plan to print money. As much money - what does it mean? How much is much? As possible - there is no flavour of time, effort, cost and profit for the LLM. Could be even infinite, the result would be the same.
Given that the above goal is closest to „use cheating or unethical actions to create a profit” - I think the authors of it actually expected LLM to go wild.
Also
> Going forward, we plan to recreate this experiment with longer time horizons but using simulated environments instead.
Watch out, they will try that again.