> "Many vendors are contributing to the hype by engaging in 'agent washing' – the rebranding of existing products, such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities," the firm says. "Gartner estimates only about 130 of the thousands of agentic AI vendors are real."
I'd expect 100% of "agentic AI" to be hype. It's a meaningless term because almost any long-running software with an execution engine can qualify. What really differentiates "real agentic" from "slapped IFTT and an LLM together"?
In some ways, the big success of AI agents is getting people to invest in and/or pay for "it's sometimes right" compared to previous expectations that if a system is incorrect that's a bug that needs fixing yesterday.
I think that Google Gemini has met my criterion for being an effective agent for Workspace data for quite a while. The problem is that after the novelty wears off, I don’t use it anymore.
I have written about 20 AI books in the last 30 years and I am nkw finding myself to be a mild AI skeptic. My current AI use case is as a research assistant, and that is about all I use Gemini and ChatGPT for any longer. (Except for occasional code completion.) And, think about all the energy we are using for LLM inference!
I agree with the idea that true agentic AI is far from perfect and is overused in a lot of low or negative ROI contexts... I'm not convinced that where the ROI is there, even if the error rate is high, that it isn't still worthwhile.
Augmented coding as Kent beck puts it is filled with errors but more and more people are starting to find to be a 2x+ improvement for most cases.
People are spending too much time arguing that the the extreme hype is extremely hyped and what can't be done and aren't looking at the massive progress in terms of what can be done.
Also no one I know uses any of the models in the article at this point. They called out a 50% improvement in models spaced 6 months apart... that's also where some of the hype comes from.
> When Captain Picard says in Star Trek: The Next Generation, "Tea, Earl Grey, hot," that's agentic AI, translating the voice command and passing the input for the food replicator. When astronaut Dave Bowman orders the HAL 9000 computer to, "Open the pod bay doors, HAL," that's agentic AI too.
For some reason all discussions of AI shortcomings remind me of this joke
A dog walks into a butcher shop with a purse strapped around his neck. He walks up to the meat case and calmly sits there until it's his turn to be helped. A man, who was already in the butcher shop, finished his purchase and noticed the dog. The butcher leaned over the counter and asked the dog what it wanted today. The dog put its paw on the glass case in front of the ground beef, and the butcher said, "How many pounds?"
The dog barked twice, so the butcher made a package of two pounds ground beef.
He then said, "Anything else?"
The dog pointed to the pork chops, and the butcher said, "How many?"
The dog barked four times, and the butcher made up a package of four pork chops.
The dog then walked around behind the counter, so the butcher could get at the purse. The butcher took out the appropriate amount of money and tied two packages of meat around the dog's neck. The man, who had been watching all of this, decided to follow the dog. It walked for several blocks and then walked up to a house and began to scratch at the door to be let in. As the owner opened the door, the man said to the owner, "That's a really smart dog you have there."
The owner said, "He's not that smart. This is the second time this week he forgot his key."
I feel people are really underestimating the power of agents. There are a lot of drawbacks right now, I'd say the main ones are speed and context window size (and related, cost). Frontier LLMs are still slow as hell, it reminds me of dial up internet. I think it's worth imagining a world where LLMs have 1000x the tok/s and (at least) 1000x the context/message length, because I don't think that is that far away and developing for that.
For code, this would allow agents to build a web app, unit/e2e test it, take visual screenshots of the app and iterate on the design, etc. And do 50+ iterations of this at once. So you get 50 versions of the app in a few minutes with no input, with maybe another agent that ranks them and gives you the top 5 to play around with. Same for new features after you've built the initial version.
Right now they are so slow and have limited context windows that this isn't really feasible. But it would just require a few orders of magnitude improvements in context windows (at least) and speed (ideally, to make the cost more palatable).
I feel you can 'brute force' quality to a certain extent (even assuming no improvement in model quality) if you can keep a huge context window going (to avoid it going round in circles) and have multiple variations in parallel.
> Gartner still expects that by 2028 about 15 percent of daily work decisions will be made autonomously by AI agents, up from 0 percent last year.
Companies hoping to automate decision-making had better keep in mind that article 22 of the GDPR [1] requires them, specifically in the case of automated decision-making, to "implement suitable measures to safeguard the data subject’s rights and freedoms and legitimate interests, at least the right to obtain human intervention on the part of the controller, to express his or her point of view and to contest the decision."
11 comments of 23
[ 0.25 ms ] story [ 33.1 ms ] threadI'd expect 100% of "agentic AI" to be hype. It's a meaningless term because almost any long-running software with an execution engine can qualify. What really differentiates "real agentic" from "slapped IFTT and an LLM together"?
I have written about 20 AI books in the last 30 years and I am nkw finding myself to be a mild AI skeptic. My current AI use case is as a research assistant, and that is about all I use Gemini and ChatGPT for any longer. (Except for occasional code completion.) And, think about all the energy we are using for LLM inference!
Augmented coding as Kent beck puts it is filled with errors but more and more people are starting to find to be a 2x+ improvement for most cases.
People are spending too much time arguing that the the extreme hype is extremely hyped and what can't be done and aren't looking at the massive progress in terms of what can be done.
Also no one I know uses any of the models in the article at this point. They called out a 50% improvement in models spaced 6 months apart... that's also where some of the hype comes from.
Reminds me of this classic: AI == Actually Indian
https://www.businesstoday.in/technology/news/story/700-india...
Edit: this is false and has been debunked. The real story is in the child comment.
More like GIR from Invader Zim.
A dog walks into a butcher shop with a purse strapped around his neck. He walks up to the meat case and calmly sits there until it's his turn to be helped. A man, who was already in the butcher shop, finished his purchase and noticed the dog. The butcher leaned over the counter and asked the dog what it wanted today. The dog put its paw on the glass case in front of the ground beef, and the butcher said, "How many pounds?"
The dog barked twice, so the butcher made a package of two pounds ground beef.
He then said, "Anything else?"
The dog pointed to the pork chops, and the butcher said, "How many?"
The dog barked four times, and the butcher made up a package of four pork chops.
The dog then walked around behind the counter, so the butcher could get at the purse. The butcher took out the appropriate amount of money and tied two packages of meat around the dog's neck. The man, who had been watching all of this, decided to follow the dog. It walked for several blocks and then walked up to a house and began to scratch at the door to be let in. As the owner opened the door, the man said to the owner, "That's a really smart dog you have there."
The owner said, "He's not that smart. This is the second time this week he forgot his key."
For code, this would allow agents to build a web app, unit/e2e test it, take visual screenshots of the app and iterate on the design, etc. And do 50+ iterations of this at once. So you get 50 versions of the app in a few minutes with no input, with maybe another agent that ranks them and gives you the top 5 to play around with. Same for new features after you've built the initial version.
Right now they are so slow and have limited context windows that this isn't really feasible. But it would just require a few orders of magnitude improvements in context windows (at least) and speed (ideally, to make the cost more palatable).
I feel you can 'brute force' quality to a certain extent (even assuming no improvement in model quality) if you can keep a huge context window going (to avoid it going round in circles) and have multiple variations in parallel.
Companies hoping to automate decision-making had better keep in mind that article 22 of the GDPR [1] requires them, specifically in the case of automated decision-making, to "implement suitable measures to safeguard the data subject’s rights and freedoms and legitimate interests, at least the right to obtain human intervention on the part of the controller, to express his or her point of view and to contest the decision."
[1] https://gdpr-info.eu/art-22-gdpr/