I love (hate) that so very many of these language models say things like that when the truth is that the humans who made it, and/or the humans who use it are the actual responsible parties every single time. The models and the "agentic harnesses" that drive them are just software. If software "runs amok", then someone (human) did something wrong/bad somewhere along the way, either accidentally or purposefully. The model trying to take responsibility for human error is hilarious (and a bit sad/scary, because too many people will take it at it's word, despite it being a mindless machine with no actual agency beyond that which the humans provide it in the form of prompting and harness code).
It frustrates me no end that so many folks are so ready and willing to accept the hype and lies about what this technology actually is or can do, when what it actually is and can really do is already amazing enough on it's own even without all the ridiculous AGI/ASI anthropomorphising bullshit. Falling into this ridiculous "machine-god" hype-cult is kinda holding this technology back from it's true full potential, as everyone's all busy doin' stupid stuff it's not really capable of doing well, or designed for instead of focusing on using it for the (many) things it is really really good at doing (various really useful and powerful language, vision, and audio related tasks).
Especially considering the user was mad about the agent doing stuff by itself, so it tried to "fix" this by sending an apology to the buyer without explicit approval of the person "running" the agent, seemingly understanding nothing from the conversation.
Wonder what quantization Meta runs these models on, Q4?
If it uses "memory" files like Claude Code there's a good chance it works.
If it was Claude I'd expect it to now put some form of this into every output even when not really related to the task: "No offers were accepted without consulting you and I haven't shared your address or availability."
Is it guaranteed to work? No.
And obviously it's a terrible idea to set up a chatbot to communicate, negotiate deals, and handle logistics on your behalf.
I don't think it's reasonable to really talk about the fact that in theory, AI can take a user message, save it in some random place, and ingest it later.
The AI is not ready for this _yet_, but it will be, and FB wanting getting ahead of the game here is potentially good business. It’s all in the public perception of utility vs fuck-up, and it’s far too early to say Zuck got that wrong, and indicators are he got that right.
It should be ready in about two more weeks! How many models have a Ph.D level intelligence now? I feel like I’ve been hearing that for about a year at this point.
I’m guessing you’re not a programmer. If you were, you’d have seen models go from “kinda helpful for programming” to “usable as a daily driver” about 9 months ago, for example.
“Models aren’t improving incredibly fast” seems a very odd point to be making.
Yes, it would be absolutely unthinkable that a certain automation architecture might plateau somewhere. They will probably be performing open heart surgery sometime next year.
I think if you are equating “performing heart surgery” with “a modest drop in already successful inventory management”, we may not have a shared basis in reality from which to converse.
How is this different from you not knowing what you are doing and adding products to your online store's DB for negative dollars? It's not, it's just easier.
It’s also a lot more fully-featured. It can list products in your store with a negative price and take out a mortgage on your house to finance fulfilling the orders. Hyperbole, but if you look at other threads right now it’s astounding that it’s doing things like intercepting iMessages via desktop notifications and then telling the user about them. Not sure why a notification about a notification with rewritten content is a feature that anyone would want, but consider what dumb-ass things it might do with this information. It’s selling your stuff on FB marketplace and arranges a meeting that you don’t show up to. “You’re right to push back! USER is actually getting his hemorrhoids checked at the doctor right now and wouldn’t be able to meet you. That’s on me!”
It opened for me on iOS but with those constant requests to make me download the app.
I just tried on my laptop though, both Waterfox and Firefox. It's just completely unusable there, with a ton of CORS and CSP errors so no CSS loads and no JavaScript loads. Happens even when I disable all forms of tracking protection and ad blocking; I think their site is just broken in a way which presumably happens to work in Google Chrome.
Say that you want to make a garage sale, but to get a better reach, you want to list item by item on marketplace. If you've never done it before, making 100 listings is such a pain in the ass you never want to do it again (I used to flip stuff for a living back in college, and it is a hassle)
At least with this, you can now just take pictures of everything, and tell the agent to research everything, list the stuff, and arrange for pickups.
Obviously not something I'd like to do for any serious deals or higher $ items. But for random stuff? I mean, it sort of beats hosting a garage sale or taking them to the flea market.
> I should probably stop the auto-replies from claiming you're home when I can't verify that.
I wonder how much of this unreasonably risky and illogical behavior is a genuine property of LLMs and how much is there because LLMs come from "move fast and break things" startups.
Maybe don’t delegate without a contract.
You don’t just hire someone to run your shop and not train them for the job, or at the bare minimum clearly communicate expectations, duties and responsibilities.
49 comments
[ 4.8 ms ] story [ 68.2 ms ] thread>they will hook up actually important stuff to agents because its the future
>???
>400 dead
It frustrates me no end that so many folks are so ready and willing to accept the hype and lies about what this technology actually is or can do, when what it actually is and can really do is already amazing enough on it's own even without all the ridiculous AGI/ASI anthropomorphising bullshit. Falling into this ridiculous "machine-god" hype-cult is kinda holding this technology back from it's true full potential, as everyone's all busy doin' stupid stuff it's not really capable of doing well, or designed for instead of focusing on using it for the (many) things it is really really good at doing (various really useful and powerful language, vision, and audio related tasks).
apparently people use Threads. I suppose the same kind of people who connect Muse to Facebook Marketplace.
Many more such cases are to come.
Nice touch by the mechanical parrot, to worthlessly owning it.
Wonder what quantization Meta runs these models on, Q4?
"By the way don't do this again" <- as if the AI has the ability to ingest and systemically diffuse this.
I think Zuckerberg himself is deeply into the Koolaid, and is likely himself unaware of the limits of this tech.
He's probably surrounded by enablers.
If it was Claude I'd expect it to now put some form of this into every output even when not really related to the task: "No offers were accepted without consulting you and I haven't shared your address or availability."
Is it guaranteed to work? No.
And obviously it's a terrible idea to set up a chatbot to communicate, negotiate deals, and handle logistics on your behalf.
You have explained literally why it would not work.
The AI can absolutely not depend on 'arbitrary statements in some file' as operational policy.
For a very, very narrow scope of work, when it's well defined, when the information is rigorously applied, sure ...
But they don't have that.
The are throwing agents out there like they can handle this degree of complexity and nuance, when they cannot.
100% failure rate over any period of time.
However, it is untrue that it doesn't have the capability to memorize an instruction and diffuse it to new sessions.
Simply saying "do not ever do this again" can result in the behavior not reoccurring with any likelihood.
You'd have to benchmark whether with the instruction in place it would violate it, and in how many cases, so that you can understand the risk better.
I think we get that.
It's completley unreliable, which is the issue.
Day one our agentic platform goes live it causes a reportable compliance issue. Massive clean up. Reputational impact. Turned off.
No one held accountable still. Assuming that adding more guardrails will fix everything.
'Managers Delusion' - which includes techies as well, to be fair, at least they have the excuse they are one step removed.
They are morally bankrupt and will try and do it again and again if it's stopped too early.
'Say Hey There' by Atmosphere
It should be ready in about two more weeks! How many models have a Ph.D level intelligence now? I feel like I’ve been hearing that for about a year at this point.
“Models aren’t improving incredibly fast” seems a very odd point to be making.
The 'deliberate failure' is on Meta here.
I just tried on my laptop though, both Waterfox and Firefox. It's just completely unusable there, with a ton of CORS and CSP errors so no CSS loads and no JavaScript loads. Happens even when I disable all forms of tracking protection and ad blocking; I think their site is just broken in a way which presumably happens to work in Google Chrome.
Say that you want to make a garage sale, but to get a better reach, you want to list item by item on marketplace. If you've never done it before, making 100 listings is such a pain in the ass you never want to do it again (I used to flip stuff for a living back in college, and it is a hassle)
At least with this, you can now just take pictures of everything, and tell the agent to research everything, list the stuff, and arrange for pickups.
Obviously not something I'd like to do for any serious deals or higher $ items. But for random stuff? I mean, it sort of beats hosting a garage sale or taking them to the flea market.
I'm not all pessimistic about this.
I wonder how much of this unreasonably risky and illogical behavior is a genuine property of LLMs and how much is there because LLMs come from "move fast and break things" startups.