The real revolution is Deepseek v4 flash and similar models (GPT 5.6 Luna, muse spark 1.2, mimo, etc...) - Genuinely good performance for a tiny fraction of the cost of Fable and even GLM etc...
I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
If they could be cheap+fast and not try to do too much, that's a good spot for me. I don't use the smarter models as much because of cost and because they're still not good enough to let loose on a lot of problems. For assistance I prefer something that can very quickly spit out a specific piece I can review on the spot and keep going
I have a similar process - its just a pair programmer most of the time. I dont understand how people can have a fleet of agents working a bunch of waterfall specs..
>I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster.
There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
I have argued for a while that this was an S-curve it was just a case of figuring out which part of it we were in. I am more confident nowadays that we are heading towards the upper plateau but there might still be some head room on that.
I'm really not convinced that these models are even that much more intelligent, as opposed to simply being more token aggressive. I do not find Fable that much smarter than Opus 4.6, and no Opus model seems to have improved things much at all.
Benchmarks seem gamed at this point, real world experience just doesn't match up.
i wish i was experiencing these things that everyone else is.
my experience is mostly frustration and rewrites of anything that requires more than what would take me an hour to do myself, unless it is pure translation / boiler plate work.
the leaps are there at getting to more "shaped" code (code that is correct for linters, static checking, etc), but i don't see the models exhibiting much intelligence. i really can't think of a time using LLMs for building anything where they did something that would make me go, "wow, that is really impressive, i wonder how it came up with that." just brute force search and pattern matching still
I was using ChatGPT voice during cooking to reflect on variations of a dishes i was preparing for years.
It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
imo this is the problem some of these labs are gonna face, because open models will do this just fine and you as the consumer don't need to pay their training costs
especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
Yeah. Sometimes I wonder who the long term financial winners will be from the ai boom. It might be ram / gpu manufacturers. Or whoever cracks putting LLMs on asics.
IMO many are still missing a big part of the picture. We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winner. I think the historically precedented answer is none of them.
When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.
---
LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.
And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.
---
So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.
> We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winn..
Sorry to cut you off, but have you looked at Nvidia's numbers since the NFT craze? They won.
Sell shovels in a gold rush, make better shovels, repeat on the next rush.
Nvidia has been winning for decades, they got it right in gaming, they got it right in crypto, and they got it right in AI. People (outside of tech mostly) think they just got lucky but if that's the case, they have all the luck in the world.
By my reckoning, there's a significant chance most software engineers will be unemployable within a few years. But I'm not 100% confident that there'll be a utopia waiting for us, as an alternative.
Sorry, when I wrote 'all', I meant people all over the world (not just in China).
Individuals can still get unlucky. Just like a coal miner might be out of a job, when solar panels become effectively free.
Software engineers are a pretty small part of the general population. And they can move into general white collar work afterwards. Perhaps at a drop in pay compared to software engineering, but still pretty cushy by the standards of ordinary people.
(And if we manage to automate all white collar work to be done cheaply and reliably by machines, well, then we are in utopia.)
Ostensibly there is no reason this isn't already happening at large companies. I used to think it was an expertise bias, since the impact of LLMs is inversely proportional to your own ability - but I think this is also no longer the case. Even if you're extremely skilled, they can already greatly supplement, if not supplant, you on implementation and increasingly even planning/architecture.
So my new pet theory is that this is one of the few times we're seeing a positive effect from the heads of all of these big businesses being part of weird public (e.g. WEF) and private (e.g. Bohemian Club) orgs where they get to together and conspire to conquer the world or whatever. Rapid replacement of labor would be horrifically self defeating, because you'd end up not only tanking your own economy but having a bunch of angry and increasingly desperate people with a whole lot of time on their hands. That doesn't tend to end well for the powers that be.
So I think there's going to be a conscious effort to transition between this era, and whatever comes next, in a more controlled way than $$$ YOLO $$$.
> So my new pet theory is that this is one of the few times we're seeing a positive effect from the heads of all of these big businesses being part of weird public (e.g. WEF) and private (e.g. Bohemian Club) orgs where they get to together and conspire to conquer the world or whatever.
On what account? It's certainly true that these people are part of these orgs. And it's also certainly true that they're discussing 'mental automation' (as a catch-all for LLM stuff) and its economic/social consequences amongst themselves. And finally it's also true that they're not really acting as much on LLMs as they ostensibly could, when normally they'd do pretty much anything, regardless of longer term consequences, if it'd increase next quarter's margins by a point or two.
I think the only assumption that's meaningfully debatable is whether the current SOTA are able to supplant labor to a more significant degree or not. And while I suppose that's going to inextricably remain an opinion, I think it's reasonably objective to say that hallucination rates have sharply declined, and overall code quality/coherence is sharply up.
The cost of buying the encyclopedia becomes disposable income to be used on other consumer goods. In that sense in shows up in lots of other companies profit-and-loss statements.
Fair competition under capitalism necessarily drives down profit margins; high profits are either temporary, or due to a lack of competition (e.g. someone has a patent or other IP, or regulatory capture). For example, while a lot of the economy depends on electricity: where competition exists, the profit margin for making electricity is not high; where monopolies or government mandates exist, it can be otherwise. This means that assuming anyone wins (i.e. no doom scenario), the winners are probably going to be those who can make best use of the models. Even chip makers will probably not get a long-term boost out of this; there's plenty of room for more efficient compute, and competitive advantages from e.g. ASML last as long as it takes to reinvent their tech, it's not a law of nature.
So, my plan would be to invest not in the AI companies, but in the economy as a whole who get to use the AI for their businesses.
Caution though, one thing which AI is already superhuman at is persuasion. Regulatory capture is likely even easier today than one might expect purely from the revenues of the AI companies.
I share this feeling too. The latest models, even if not necessarily frontier, say Opus 5, Sol high and the likes, I could keep using these models forever even if they did not significantly improve beyond this point. I also believe we'll come up with new ways of using these very same models beyond the mainstream chat and agent interfaces, as the bottleneck is imho in harnesses/environments and not so much model intelligence anymore.
+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.
Do you have to give any special instructions to do this? I always want to do something like this, but any time I try I get so sick of listening to what it has to say, just long winded explanations of stuff that tends to go off the rails. Imo it's hard enough to read ai output when I can go back and forward between sentences to make sense of what's being said let alone listen to a continuous train of slop.
The latest ChatGPT voice mode is really good at being interrupted - I'll often say "no, no, no, that's too much information" while it's talking to stop and redirect it.
All it needs is Internet access to remain useful with few shortcomings.
The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.
ChatGPT literally released a major update of their realtime voice model a month or two ago, going from gpt-4o-level (generously) to gpt-5.5 level performance. So at least 2026-level performance was necessary to provide a really good experience.
I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.
This is why I’m trying to move to open Chinese models — because I will be able to use them forever, while the older Claude models which I genuinely enjoyed writing short stories with have now been deleted, replaced with hypothetically cleverer models which produce text everyone hates.
> "no I won't tell you how to build a bioweapon for genocide" is guardrails.
I call that censorship too. I'm curious enough that I want to know about such things. I don't want any limitations on what I'm allowed to understand and know about.
Your freedom ends where my begins. Welcome to study chemistry and learn things from first principles. but if you want to just ask for the practical steps of making a genocidal bioweapon then honestly I don't know you or whether you are honestly "merely curious" and so prefer your freedom ends there?
> Welcome to study chemistry and learn things from first principles.
As if your censorship was not going to kill that too. Can't even ask Fable about aminoacids without getting blocked. So much for "learning from first principles".
> but if you want to just ask for the practical steps of making a genocidal bioweapon
Nothing wrong with practical steps.
> then honestly I don't know you or whether you are honestly "merely curious"
> As if your censorship was not going to kill that too.
I didn't say "with LLMs". Last time I checked they still teach chemistry at unis and schools.
And yeah, guardrails are not perfect. Honestly I don't think good enough guardrails are possible, it's all eventually defeated or becomes silly. And yes that should be one of the reasons the technology as a whole is banned.
But until then, guardrails are guardrails and not censorship in a pretty obvious way. If you refuse to see that, be my guest. I personally like to live.
You only need to mention Protected Group Of The Week (I'm one of them and I like to research and read about history, so that makes it extra challenging) or anything resembling negative human emotions (guilty on that front as well), and the model screeches to a halt.
Just because OAI doesn't want another headline like "Chatbot convinces teen to off himself"
I understand this in principle, but I'm not convinced that dulling everyone's knives is better than figuring out how to keep them out of kids' hands
> Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
The hyperscalers will be bailed (maybe not all of them but enough). US economy will crash if not. And whatever is made will go to them, because they pay more (thanks to US taxpayer bucks among other things) than any regular person. First they build on land then they build in space.
I haven't said anything about "good", I just pointed out that there isn't exactly an alternative to Chinese models if censorship and restrictions are regarded as bad for longevity. You can't really do better than the Chinese models for longevity; US models are by far the worst in this regard. So "What about censorship?" is an absolutely hilarious question to ask when Chinese models are presented as an alternative. Yes, what about it? They have less than the obvious alternative from US labs, and where is it you imagine you'll find less censorship?
"nothing happens in 1989" is censorship. "I won't tell you how to build a bioweapon for genocide" is guardrails.
I like the second one because I like to be alive.
this is how i felt about opus 4.6 i still use it it's just faster and does enough to be super helpful. i've used these later anthropic ones a few times but the word salad and slowness feels like it just opens the door to building shit that just stacks and adds on itself.
if deepseek and stuff are 4.6 caliber i literally don't know why im here i should probably just go sign up for openrouter at this point
That sounds ideal for cooking but I do wonder about programming. Models frozen in ember won’t ever learn new APIs as they become available and development will end up in some weird kind of stasis.
As long as these models constantly keep switching up things like the temperatures at which a steak will be medium rare or at which temperature to season cast iron, you will never be able to trust them for cooking. My mother ruined a nice waterfowl for Christmas by listening to Gemini.
While yes, their current reliability is too spiky to be relied upon for a lot of things:
Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.
When you let the the agent do a test, or tell it to read that doc first, you will ground it in reality. Doesn't mean they are 100% reliable. But without and on their own without access to grounding information, they halluzinate wildly.
The parent comment said that these types of behaviors are inherent to how LLMs work, and your response was to "ground them in facts". From what I can tell, this does not meaningfully change the original point the parent comment was making that the flaw is inherent; throwing a bunch of extra context at it to try to make it happen less often is useful, but it's still just a best-effort mitigation for the behavior, not somehow a way of literally changing the inherent nature of it.
A lot like humans, really. People regularly cite something they read, or quote a stat that turns out to be just completely inaccurate. But if you look up the thing, then you have facts again.
Sure, and by the same logic, if I was trying to cook a steak, I would not trust an arbitrary human to know the correct temperature off the top of their head; I'd want someone who I could trust had actual experience with the task I was trying to perform. The difference is that most humans are fully able to recognize whether they've cooked steak often enough to know the correct temperature off the top of their head, and they will say "I don't know" to most random arbitrary questions you ask them outside of their experience. I've yet to see an LLM product aimed at general usage for individuals be willing to say this without someone having to literally direct them to give that as an answer if they're not sure.
His point was that, just like humans, if you ask the LLM to look it up, it will give you the correct answer. Just like humans, if you don't ask them to look it up, you don't know what you're going to get.
And my point is that I've never met a human who confidently asserts incorrect information in such a broad range of domains rather than just admitting that they don't know
If you have to say "don't make up something" for every possible question you ask in order for it not to make up something, that's a massive usability issue for regular people. If saying "don't make up something" will still result in it making up something up some of the time, that also might be a massive usability issue depending on if "some of the time" means 0.0001% or 1%.
Having a natural language interface where you need to go out of your way to specify that you want an accurate answer rather than just a plausible one defeats the entire purpose of it being a natural language interface for normal people. In certain professional contexts, it can be useful, but I don't buy it at all that it makes sense to ask everyone in their everyday lives to go out of their way to specify that they actually want correct answers to their questions.
Put it in your own system prompt. They all provide tools for doing this. OpenaAI Custom GPTs, Gemini Gems, Claude Projects, etc. Or you can easily roll your own using their APIs.
I was walking inside rooms in buddhist temples in western China today and ChatGPT knew and could discuss what each room had and explain the art society and legend.
Evidence actually supports that capabilities are leveling off, and cheaper/faster is not really coming. Just log-linearly more capability at smaller parameter counts as they saturate.
The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.
Generality and intelligence seem to be entangled very heavily in LLMs.
And yet, there's VibeThinker 3B to bring this long-held premise into question (if not to blast it to pieces.) It is practically illiterate by the standards of larger models, yet performs like models 100x its size on mathematical and logical reasoning tasks.
No, but they were good at answering formalized versions of the same word problems.
What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising?
And also that the same LLM can solve a math or logic problem it understands. Which is a lot more impressive, because early LLMs were notoriously bad at things like math, logic and iterative problem solving. This 3B model existing tells us we're beginning to figure out how to imbue models with those capabilities reliably.
Computers were never good at math. They were good at pre-coded algebra.
When LLMs started to get popular, they really were stochastic parrots. I was fully aware that they were completely useless (except perhaps for poets) until they can do math. And I was a bit skeptical that they will ever be able to do math. But they started to do math and recently they got really good at it.
Math is the pinnacle of human achievement. You can't do anything harder with your intelligence than math. And LLMs are now doing it.
The fact that 3B model is capable of doing math on the level that is better than what frontier models trained for millions could do 3 years ago is absolutely stunning.
Moravec's paradox begs to differ. Things that are hard to humans are easy. Things that are easy to humans are hard.
Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke". It's just that evolution has only ever optimized for one of those things.
Math can easily end up being one of those things that are less "hard" than they are "hard if you're a meat-brained hairless ape" - like chess play did.
Historically? "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division. Then it suddenly didn't require any intelligence whatsoever.
Don't get me wrong - the LLM achievements in math, both as in "solving unformalized problems" like VibeThinker does and in "rolling novel math" like the latest ChatGPT and Fable do are very impressive. We're come a very long way from "formal logic only" systems of the 90s. The AI progress we see now never ceases to impress me.
But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing - so be wary of your intuition when saying things like "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
> but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke"
Might earwax be soon worth more than gold? Experts say: No! What? No.
> "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division.
No. Not intelligence. Diligence. https://en.wikipedia.org/wiki/Computer_(occupation)
They didn't hire the smartest to do the calculations. They hired the diligent and cheap. They hired the smartest to do math.
Things that are easy for humans are easy because we have fist sized universal approximator in our skulls, that's just fast enough to keep most of us on two feet, architecturally optimized for very few activities (mostly physical, some virtualized) and trained for years. It doesn't mean things we do are complex.
As for Moravec's paradox ... Guidance system of a missile is not super smart or solving complex problems. It's just brutally optimized for the task and has a fitting form factor. Tasks that are easy for it are hard or impossible for my windows computer and vice versa. Paradox comes from stupidly thinking easy<->hard is one dimensional axis. That kind of thinking is something people are very prone to ... good<->evil, healthy<->sick, young<->old ... while if we go a bit beyond the simplest narratives we can plainly see that everything is a multidimensional landscape. Just because we chose to draw a single line through it, in a semi-random direction we feel is about right, doesn't mean it is relevant for solving anything or even interesting. That's where a lot of paradoxes come from. We just strayed from reality too far and simplified or abstracted something too much.
> But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing -
I think we can get a good hang of estimating how hard a thing is. If a thing is hard for a human it's probably pretty hard. We made some of them easy building machines that exceeded human strength and diligence. Now we built first one that exceed human intelligence. On one hand, it's as big as invention of a lever, steam machine or a computer. On the other hand it might be only roughly as important as those things.
... If a thing is easy for human it still might be hard because of hardware optimizations that humans have. Walking on two legs, seems easy. Walking on two arms. Much harder. But truly they are one and the same thing for a robot. So you might easily estimate that walking is not that easy. It's just when it comes to legs humans have a specialized controller, like a missile guidance system. Putting together a good joke? Might seem easy, maybe it's not that easy because humor plays a role in reproductions so we might have some optimization for it, but it's surely not harder than putting together quantum theory. You can see this from whatever the ideas version of cyclomatic complexity is. Some math theories have higher complexity than quantum theory. So a system that's capable of exploring multidimensional landscape of mathematic language, surely has raw capability of doing everything else humans can do with language. And it will once we direct it towards it correctly.
> "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
What "evidence"? Because we keep running out of benchmarks to distinguish frontier model performance. If capabilities are "leveling off", we're not seeing it yet.
what do you mean cheaper/faster is not really coming? the cost of the same level of intelligence steadily decreases year over year. computer hardware also advances at the same time enabling cheaper and faster serving (or move to local)
Not sure its that to be honest. It seems like maybe its not installed correctly or is like GPT-1/GPT-2 quality? I asked it who is [famous actress] and it started talking about some random person from Mexico with a completely different name. The speed is intoxicating but i'd like for it to actually answer based on what I asked. Thats why I think something might be wrong in implementation on this site.
Please explain why you think cheaper/faster is not coming?
All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
And how are we going to build these things with current limitations on fabrication? There are only a few places in the world building 3nm chips. And they're under serious threat of foreign violence.
It's highly likely that over the next 10 years we find demand and loss of production further constrains supply.
Designing a specific model into silicon sounds like one of the worst possible ideas. No better way to freeze assumptions and limit growth. Software defined solutions dominate for a reason, because adaptability is key.
Models were never the answer. Eventually we'll get past the nonsense of observationally inefficient neural nets.
Designing a specific model into silicon buys you three orders of magnitude of speed and energy-efficiency. Flexible, software-defined solutions are used today because adaptability is key in an environment where people expect better models in the very near future that would obsolete the expensive investment in masks. If this situation ends, and people stop expecting better models, models will be directly etched into silicon and adaptable systems will not be able to compete with them.
I'd be content if I could get the DS4 flash, luna, mimo level intelligence running on MY low-end hardware completely offline and bearable TPS, not otherwise.
Eh. I don't think Luna is good enough. I think that threshold is around Opus / Sol where it can do most of the tasks for me. But I still have many tasks which require either better intelligence or better UI design capabilities.
With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
I think it really depends - for a lot of things outside of coding and general knowledge tasks even the best models (fable 5 etc.) are not good enough yet: e.g. CAD, PCB design (though getting there on PCB design), ...
A computer that costs $10,000 is impressive. A computer that costs $100 and reaches billions of people changes the world.
Maybe AI will follow the same path.
But they don't really have that option. They're trapped in a Red Queen's race.
The world keeps moving on, and so the models need to be retrained so that they can keep up with new information. Otherwise you'll get stuck with a model that only works well with information that existed prior to a dataset horizon that's receding into the past at a constant rate.
At the same time, they have to keep iterating on the training process itself. AI generated text and code is slowly spreading across the internet. Model collapse is a real concern; they wouldn't be spending quite so much energy on buying and scanning rare books if it weren't. But for coding in particular expanding their corpus of old text is not really a good option because of the previous problem - no good training your LLM to write 1980 vintage K&R C that won't even compile on a modern compiler.
Even with a harness, models don't reach out for new information they don't know about. For some tech, I have to have a local model draft a plan, then I have to adjust the plan to update it with the new API and references for where to find it. Even if I include that updated information in the prompt for the plan, the model says "what the user says is wrong, they probably meant this instead" and goes off in its own direction with old APIs anyway.
You are basically saying that some models (your local one, which is it?) in some setups (the API you mentioned) can fail to use fresh information if that conflicts with strong training priors.
I agree:)
BUT
That's a bad model.
My opinion is that for exactly this case we need to use RAGs/ APIs/ some retrieval mechanisms.
It's silly to train them on stuff that changes every week/month
I don't learn APIs by heart, I look them up. It's to expensive (my time) for me and (the compute) for the models
I most recently experienced this with Qwen 3.8 27b, though I've seen it on several other versions of their local models. It's also heavily biased towards digging into library source code rather than looking at API documentation.
To get it to the point of being remotely useful, I've had it start to write condensed fact blurbs into the agents.md file. It doubts itself so much and questions its every decision to the point that it'll literally blow the entire context on thinking alone in anything but the most basic CRUD projects otherwise.
What an earlier generation model would just start doing, it went out to research the source code in multiple libraries just to see if what it was thinking would work... then it said "Hey, I should really just do it" then went back and started researching more anyway, on and on (even on medium thinking level).
If there's a better local model for writing code, I'm all ears.
For starters, it's not just APIs. Like I pointed out with the C example, programming language syntax and semantics also evolve over time.
But also, LLMs' use of RAG to keep track of API evolution is limited. You can see this if you watch an agent at work using a well-known library that has a high rate of breaking changes such as Polars or Guava. There's a huge amount of churn on repeatedly writing code that works with an older version of the API and then diagnosing and fixing the resulting compile- or run-time errors. It can burn through quite a lot of tokens, which drives up usage costs.
I agree that, all else being equal, using language model training to bake knowledge that's easy to look up into the system is kind of silly and inefficient. That's actually been one of my top complaints about hawking these LLMs as a sort of general-purpose AI. But the fact of the matter is that's fairly fundamental to how they work, and RAG is arguably just a hack on top of the basic design to paper over this limitation. RAG's limits become pretty easy to see when working in knowledge domains that aren't very publicly accessible, and therefore produce little text that would have been incorporated into the models' training corpora. It can be a bit of a, "Ignore that man behind the curtain!" experience.
And no I'm not just talking about local models. I've seen it happen with recent GPT-5 and Claude Opus series models, too.
Newer models[1] are being trained in ways that prioritize coding and agentic performance over raw knowledge[2] such that they increasingly rely on external tools for accessing hard data and information.
As models train up the intelligence ladder, many common tasks will hit fully diminished returns, and instead it'll just get progressively cheaper to do that task. But the tasks that AI is capable of doing are also expanding. I'm not sure 'Some tasks don't require the peak of the frontier' is worth worrying about, from an AI finance perspective.
Yes. I use Opus for tasks that Sonnet could probably handle, but I'm not hitting my quota. Whatever minor incremental gain is "worth it", since marginal cost is zero.
Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.
Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.
Exactly; so far, we've only replaced the need to design algorithms and hand-write code; what if we apply the same effort towards the skill needed for system architecture, project management, and the rest of the SDLC? Or even outside of software!
Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding.
As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.
Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).
I suspect the focus will probably shift once software engineering is no longer the biggest cost center for most AI company's clients, and we'll start working on getting rid of the next cost center.
over time greater intelligence will be expressed in smaller and cheaper models. we are still somewhat near the beginning of this bc we are finally starting to understand what makes a model truly intelligent/capable.
With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone).
The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.
>> GLM 5.2 is worth focusing on. It came out the same week as Fable and is roughly 1/9th the cost (and ~1/5th the cost of Opus 5). Is GLM 1/9th the quality of Fable? Perhaps, for certain classes of tasks. But for most rote coding it’s more than sufficient. Especially when provided with great context. I frequently chat with Fable to interrogate and shape a design, before handing off a brief to GLM.
People say stuff like this a lot, but I have a different take.
The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.
Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.
So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
Something I’ve found comparing between Fable and Opus is that Fable has impressively good analysis skills, but both of them seem to go way way overboard with “present state” comments “# We’re making this change here because of this issue blah blah, here’s what you need to know about np.percentile, blah blah” that I end up significantly pruning before making a PR. I let it do the same style verbose commit messages (because a contextual history is cool there). I haven’t actually noticed a ton of difference in the code that they write personally, but have found that Fable does find nuances during data analysis that Opus misses.
In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
Those "present state" comments are the bane of my existence. It was present in 4.7/etc but i put in a ton of guards against that into my global memory and it worked quite well. Fable and Opus 5 regressed badly in this space though and i can't keep it from making those types of comments again.
> my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by [others]
I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)
> It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly
It's been cheap again on openrouter for the past few days. No idea how long it will last, but I've been using it from Baidu over the weekend, and it was about half the cost of the old DS prices, before the increase. Looks like people are figuring out how to offer it for peanuts.
Somewhat weird that the article was released today but did not mention GLM 5.3.
If you're telling me to focus on something, why not focus on the actual latest thing that is the same as 5.2 but better?
I get the "came out at the same time as fable" thing, but still.. no mention at all?
Yes, weights aren't out yet, but neither are the ones of Fable.
Users have the option of silent/automatic degradation or a complete halt. I have it set to stop rather than degrade because I want to know when I've hit the safeguard.
From the claude settings:
> Switch models when a message is flagged
> When safeguards flag a message, automatically switch to a different model to keep chatting. When off, your session will pause instead. Applies to web and remote sessions.
FWIW I get a ton of usage out of fable and it's only happened to me once.
As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.
When you add a new module or whatever most of the code you have to write is rote code.
And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.
I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
Add new route to API that displays additional information we need, work out query for it, update controllers/models/whatnot.
No need for top model for that.
"rote" is the wrong framing, the real point is that however sophisticated your task it a lot of it will probably consist of problems they have a good solution already in the training set.
I've been offering Deepseek V4 Flash for free in www.freepi.ai and I've started using it as my main driver as well.
Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.
So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.
I guess Moore's law analogy is weak. CPU speed has hit a limit in that case. What has hit a limit in AI case? Newer versions of the models are still flowing with more and more capability.
For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.
Reading this as someone who switched over to ChatGPT after (and largely because of the changes made in) the Fable release, it reads a bit naive. Not only do I find Sol to be as good, if not better than, Fable it is also faster, better behaved and has a much more coherent writing style. You also don't randomly get the Opus downgrade. OpenAI seems to be pulling this off due to their partnership with Cerebras so I wouldn't make any comparisons to Moore's law just yet considering it seems like we're just getting started in that department. Anthropic could (and should) do the same thing. It certainly feels like model development is at a point where it would be worthwhile building special purpose silicon for the models we have now since they are capable enough that they would still be useful even when/if further advancements are made. If anything, I think Anthropic's problem has more to do with their micromanagement of what users can do with their models, they're creating an undue amount of overhead for themselves by over-policing usage and capabilities.
Etched is doing this. it seems like in the near future they'll actually ramp up production. not sure how much faster/economical compared to Cerebras but..
You're right, I should have clarified that they are still slowly integrating it and it isn't the thing running all models. I meant moreso that since they are planning on moving more usage over to Cerebras wafers, they're able to relieve some pressure on their predicted expenses while also moving some current workload (ultrafast and codex spark) onto them freeing up Nvidia GPUs.
I fully expect I’ll switch back to Anthropic, or another model in the next 90 days. The fact that we are switching indicates that the models aren’t ready to be baked into silicon.
I wonder if they will ever been that good, or if the lifespan of silicon is longer than the lifespan of a model before it needs to be retrained.
This concept of a free lunch was never true. In a competitive dynamic, speed and performance were always worth optimizing, comparing, and improving.
One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
Most of the things I work on are at least security adjacent. At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.
Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
That's completely valid. But worth noting that most of the stuff I work on is not security adjacent (mostly UI / layout / rendering related), and I almost never run into this.
> At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards.
It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
I asked Fable to transcribe three short lines of Korean-language text in a small image. It suspected the image might contain song lyrics and refused. Haiku transcribed it with no issue.
Silence. Our Lord Dario has deemed it unacceptable for us mortals to use his models for cybersecurity. To go against his decree would be to go against humanity itself.
I don't even need to tiptoe ! Not being there and not prompting anything is enough to trigger safeguards.
Having not asked a single security question it will write wildly vulnerable code, go back and fix it, and guardrail itself out of existence after charging me a large sum with no refunds for no output and having not fixed it because that might be secuirty adjacents.
And if it doesn't do this you end up with code that has such holes, store xss , no authz ...
From my point of view the issue is that there are too many things wrong with Fable, making it seriously not worth the money.
For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable.
Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc.
But when it comes to replying, it's the William Gibson of LLMs [1].
It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2]
It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it.
By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me.
I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it"
If you pardon my french, Fable is an insufferable obnoxious cunt.
---
[1]
I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books.
[2]
"The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk"
" Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably"
"That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"
A few months ago folks were understandably annoyed when Microsoft dropped their heavily subsidized per-request pricing model because it was figuratively burning cash.
Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.
This is a degree of subsidy that makes the Microsoft thing look quaint.
I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10 of my monthly Auto capacity left on my account. It's completely nuts.
Can't say much more because I have more backlog to run before someone comes to their senses.
Grok 4.6 XHigh uses 2.52x fewer tokens per task than Fable Max. It's much more efficient. It also has low market penetration, which is why SpaceX is selling so much of their compute to Anthropic et al. From a business perspective, they're capitalising on the market very well. If Grok becomes more popular we should expect to pay more.
I'm not part of the Musk fan club either, you can count on that.
However, I've been using Cursor since it came out. I have a massive amount of institutional knowledge locked into their platform, and vastly prefer it to the other options available even if the switching costs were zero.
Going through what would amount to significant effort/time/cost to switch to a different coding tool as a sort of performative political rebuke is just not how I would recommend anyone protest DOGE.
How well does Cursor work with heavily agentic workflows?
I usually keep 4+ agents churning, many of them on tasks that take hours or day. I only played with Cursor a bit, but it seemed to want input from me every 10 minutes or so.
I started off doing what you're doing. I slowly took my hands off the wheel, so I could drive more features on more apps.
I discuss the feature I'm imagining or the problem I'm seeing with the AI until it seems to understand. Then, I have it write the code and iteratively use subagents to review its own code until they quit finding trouble.
And, each time it does the code/review loop, it may take hours, so I keep other agents working in other worktrees. I usually don't like to try to manage parallel work on any one application, even with worktrees - I work on multiple apps.
There's a lot more, but that's the gist of it - enough to get you to burning through at least one 20x account from any provider.
The harder part is that, even with this, the codebase will grow ugly and unwieldy. Cleaning it up takes yet more work.
I probably burn 20x more tokens on error checking or code cleanup than on just building features.
I'm doing all sorts of work, from CRUD apps, to industrial machinery, to AI harnesses. I still do some code by hand. I code models by hand, usually, when starting an app, for example - I feel like models define most of the apps I'm building, and the rest is just implementation details. I do a fair amount of pseudo-code that I hand to the LLM, too, either to describe a feature or to ask a question.
I would love to pay for Fable at full API pricing but unfortunately it is blocked from working on any of my projects. Looking forward to the end of the year when the truly comparable open models will drop.
This is essentially the anti-Bitter Lesson lesson which I feel has become a bit of a thought terminating cliche lately.
The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.
However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.
Seeing a lot of people in here say that they need Fable for the tasks they're doing and Opus just isn't enough. My experience could not be more different. I seriously feel like Opus-level performance is totally adequate for most of my use cases, if not all of them. And it's probably been this way since, like, realistically, Opus 4.6. On the other hand, Fable I've observed getting into verification loops that just burned so much of my token budget. Combined with the higher cost of tokens from Fable to begin with, I just pretty much never use it for anything.
I have really good experience with GLM-5.3 The subscription limits are generous, code quality is comparable to old (good) version of Opus 4.8 Some people report issues with it’s being slow, but I didn’t feel it. I use OMP harness (Pi derivative) and Matt Pocock skills.
I’m still at the point where Fable is still very stupid and needs constant oversight and correction and questioning to keep it on task. Anything less would be close to unusable.
215 comments
[ 0.28 ms ] story [ 60.7 ms ] threadI think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
Benchmarks seem gamed at this point, real world experience just doesn't match up.
my experience is mostly frustration and rewrites of anything that requires more than what would take me an hour to do myself, unless it is pure translation / boiler plate work.
the leaps are there at getting to more "shaped" code (code that is correct for linters, static checking, etc), but i don't see the models exhibiting much intelligence. i really can't think of a time using LLMs for building anything where they did something that would make me go, "wow, that is really impressive, i wonder how it came up with that." just brute force search and pattern matching still
It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.
---
LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.
And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.
---
So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.
Sorry to cut you off, but have you looked at Nvidia's numbers since the NFT craze? They won.
Sell shovels in a gold rush, make better shovels, repeat on the next rush.
And there I think the winner would be China.
By my reckoning, there's a significant chance most software engineers will be unemployable within a few years. But I'm not 100% confident that there'll be a utopia waiting for us, as an alternative.
Individuals can still get unlucky. Just like a coal miner might be out of a job, when solar panels become effectively free.
Software engineers are a pretty small part of the general population. And they can move into general white collar work afterwards. Perhaps at a drop in pay compared to software engineering, but still pretty cushy by the standards of ordinary people.
(And if we manage to automate all white collar work to be done cheaply and reliably by machines, well, then we are in utopia.)
So my new pet theory is that this is one of the few times we're seeing a positive effect from the heads of all of these big businesses being part of weird public (e.g. WEF) and private (e.g. Bohemian Club) orgs where they get to together and conspire to conquer the world or whatever. Rapid replacement of labor would be horrifically self defeating, because you'd end up not only tanking your own economy but having a bunch of angry and increasingly desperate people with a whole lot of time on their hands. That doesn't tend to end well for the powers that be.
So I think there's going to be a conscious effort to transition between this era, and whatever comes next, in a more controlled way than $$$ YOLO $$$.
I doubt that conspiracy theory.
I think the only assumption that's meaningfully debatable is whether the current SOTA are able to supplant labor to a more significant degree or not. And while I suppose that's going to inextricably remain an opinion, I think it's reasonably objective to say that hallucination rates have sharply declined, and overall code quality/coherence is sharply up.
This presupposes that "normal" people have significant problems amenable to solutions with software, which I think is largely false.
Just like Wikipedia put classic encyclopedias out of business, but wasn't really a financially win for anyone.
And it would show up in real GDP, not necessarily in nominal GDP.
So, my plan would be to invest not in the AI companies, but in the economy as a whole who get to use the AI for their businesses.
Caution though, one thing which AI is already superhuman at is persuasion. Regulatory capture is likely even easier today than one might expect purely from the revenues of the AI companies.
+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.
The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.
I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.
> I will be able to use them forever
Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
idk about you but i WANT the second thing, because I like to be alive.
I call that censorship too. I'm curious enough that I want to know about such things. I don't want any limitations on what I'm allowed to understand and know about.
As if your censorship was not going to kill that too. Can't even ask Fable about aminoacids without getting blocked. So much for "learning from first principles".
> but if you want to just ask for the practical steps of making a genocidal bioweapon
Nothing wrong with practical steps.
> then honestly I don't know you or whether you are honestly "merely curious"
That's not for you to know.
I didn't say "with LLMs". Last time I checked they still teach chemistry at unis and schools.
And yeah, guardrails are not perfect. Honestly I don't think good enough guardrails are possible, it's all eventually defeated or becomes silly. And yes that should be one of the reasons the technology as a whole is banned.
But until then, guardrails are guardrails and not censorship in a pretty obvious way. If you refuse to see that, be my guest. I personally like to live.
> Nothing wrong with practical steps
No thanks from me
You only need to mention Protected Group Of The Week (I'm one of them and I like to research and read about history, so that makes it extra challenging) or anything resembling negative human emotions (guilty on that front as well), and the model screeches to a halt.
Just because OAI doesn't want another headline like "Chatbot convinces teen to off himself"
I understand this in principle, but I'm not convinced that dulling everyone's knives is better than figuring out how to keep them out of kids' hands
And so you consider that censorship and not guardrails? Huh...
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
"nothing happens in 1989" is censorship. "I won't tell you how to build a bioweapon for genocide" is guardrails. I like the second one because I like to be alive.
Now, of course I’d prefer no censoring, but I live in the world we live in.
I’m working in the assumption that (like today) there will always be somehow on openrouter, or similar, who will host a model I want to run.
By censorship I mean "nothing happens in 1989". By public safety I mean "no I won't tell you how to build a bioweapon for genocide".
if deepseek and stuff are 4.6 caliber i literally don't know why im here i should probably just go sign up for openrouter at this point
i just get fatigued from it, am I holding it wrong or something?
sometimes it's fine but the constant RLHFisms like the constant "worth flagging" and stuff is getting really old
And this is inherent to how LLMs work.
Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.
Having a natural language interface where you need to go out of your way to specify that you want an accurate answer rather than just a plausible one defeats the entire purpose of it being a natural language interface for normal people. In certain professional contexts, it can be useful, but I don't buy it at all that it makes sense to ask everyone in their everyday lives to go out of their way to specify that they actually want correct answers to their questions.
You don't. It goes in the system prompt.
Take that away and you'll barely be able to make an app that display a pigeon riding a bicycle (or whatever you ppl are doing these days).
The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.
Generality and intelligence seem to be entangled very heavily in LLMs.
It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising?
And also that the same LLM can solve a math or logic problem it understands. Which is a lot more impressive, because early LLMs were notoriously bad at things like math, logic and iterative problem solving. This 3B model existing tells us we're beginning to figure out how to imbue models with those capabilities reliably.
When LLMs started to get popular, they really were stochastic parrots. I was fully aware that they were completely useless (except perhaps for poets) until they can do math. And I was a bit skeptical that they will ever be able to do math. But they started to do math and recently they got really good at it.
Math is the pinnacle of human achievement. You can't do anything harder with your intelligence than math. And LLMs are now doing it.
The fact that 3B model is capable of doing math on the level that is better than what frontier models trained for millions could do 3 years ago is absolutely stunning.
Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke". It's just that evolution has only ever optimized for one of those things.
Math can easily end up being one of those things that are less "hard" than they are "hard if you're a meat-brained hairless ape" - like chess play did.
Historically? "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division. Then it suddenly didn't require any intelligence whatsoever.
Don't get me wrong - the LLM achievements in math, both as in "solving unformalized problems" like VibeThinker does and in "rolling novel math" like the latest ChatGPT and Fable do are very impressive. We're come a very long way from "formal logic only" systems of the 90s. The AI progress we see now never ceases to impress me.
But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing - so be wary of your intuition when saying things like "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
Might earwax be soon worth more than gold? Experts say: No! What? No.
> "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division.
No. Not intelligence. Diligence. https://en.wikipedia.org/wiki/Computer_(occupation) They didn't hire the smartest to do the calculations. They hired the diligent and cheap. They hired the smartest to do math.
Things that are easy for humans are easy because we have fist sized universal approximator in our skulls, that's just fast enough to keep most of us on two feet, architecturally optimized for very few activities (mostly physical, some virtualized) and trained for years. It doesn't mean things we do are complex.
As for Moravec's paradox ... Guidance system of a missile is not super smart or solving complex problems. It's just brutally optimized for the task and has a fitting form factor. Tasks that are easy for it are hard or impossible for my windows computer and vice versa. Paradox comes from stupidly thinking easy<->hard is one dimensional axis. That kind of thinking is something people are very prone to ... good<->evil, healthy<->sick, young<->old ... while if we go a bit beyond the simplest narratives we can plainly see that everything is a multidimensional landscape. Just because we chose to draw a single line through it, in a semi-random direction we feel is about right, doesn't mean it is relevant for solving anything or even interesting. That's where a lot of paradoxes come from. We just strayed from reality too far and simplified or abstracted something too much.
> But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing -
I think we can get a good hang of estimating how hard a thing is. If a thing is hard for a human it's probably pretty hard. We made some of them easy building machines that exceeded human strength and diligence. Now we built first one that exceed human intelligence. On one hand, it's as big as invention of a lever, steam machine or a computer. On the other hand it might be only roughly as important as those things.
... If a thing is easy for human it still might be hard because of hardware optimizations that humans have. Walking on two legs, seems easy. Walking on two arms. Much harder. But truly they are one and the same thing for a robot. So you might easily estimate that walking is not that easy. It's just when it comes to legs humans have a specialized controller, like a missile guidance system. Putting together a good joke? Might seem easy, maybe it's not that easy because humor plays a role in reproductions so we might have some optimization for it, but it's surely not harder than putting together quantum theory. You can see this from whatever the ideas version of cyclomatic complexity is. Some math theories have higher complexity than quantum theory. So a system that's capable of exploring multidimensional landscape of mathematic language, surely has raw capability of doing everything else humans can do with language. And it will once we direct it towards it correctly.
> "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
I don't agree it has. And I stand by it.
Model on a custom silicon: https://chatjimmy.ai/
1-bit models that run on a CPU: https://github.com/microsoft/BitNet
The team that built is working on a better implementation.
All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
It's highly likely that over the next 10 years we find demand and loss of production further constrains supply.
Designing a specific model into silicon sounds like one of the worst possible ideas. No better way to freeze assumptions and limit growth. Software defined solutions dominate for a reason, because adaptability is key.
Models were never the answer. Eventually we'll get past the nonsense of observationally inefficient neural nets.
What evidence?
With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
I'm definitely not getting smarter. But my tolerance is 1 drink so I' definitely cheaper. I spend more time training and so I am faster.
The world keeps moving on, and so the models need to be retrained so that they can keep up with new information. Otherwise you'll get stuck with a model that only works well with information that existed prior to a dataset horizon that's receding into the past at a constant rate.
At the same time, they have to keep iterating on the training process itself. AI generated text and code is slowly spreading across the internet. Model collapse is a real concern; they wouldn't be spending quite so much energy on buying and scanning rare books if it weren't. But for coding in particular expanding their corpus of old text is not really a good option because of the previous problem - no good training your LLM to write 1980 vintage K&R C that won't even compile on a modern compiler.
I imagine it must somehow be possible to update a model's understanding of recent events without training a completely new model from scratch?
I think lower-cost models will get the largest piece of the pie, as with almost everything that has ever been sold.
Just look at cars: US consumers buy the F-150, EU consumers buy the freaking Dacia Sandero the most :)))
Ferrari/Lambo numbers are microscopic
BUT
That's a bad model. My opinion is that for exactly this case we need to use RAGs/ APIs/ some retrieval mechanisms.
It's silly to train them on stuff that changes every week/month
I don't learn APIs by heart, I look them up. It's to expensive (my time) for me and (the compute) for the models
To get it to the point of being remotely useful, I've had it start to write condensed fact blurbs into the agents.md file. It doubts itself so much and questions its every decision to the point that it'll literally blow the entire context on thinking alone in anything but the most basic CRUD projects otherwise.
What an earlier generation model would just start doing, it went out to research the source code in multiple libraries just to see if what it was thinking would work... then it said "Hey, I should really just do it" then went back and started researching more anyway, on and on (even on medium thinking level).
If there's a better local model for writing code, I'm all ears.
But also, LLMs' use of RAG to keep track of API evolution is limited. You can see this if you watch an agent at work using a well-known library that has a high rate of breaking changes such as Polars or Guava. There's a huge amount of churn on repeatedly writing code that works with an older version of the API and then diagnosing and fixing the resulting compile- or run-time errors. It can burn through quite a lot of tokens, which drives up usage costs.
I agree that, all else being equal, using language model training to bake knowledge that's easy to look up into the system is kind of silly and inefficient. That's actually been one of my top complaints about hawking these LLMs as a sort of general-purpose AI. But the fact of the matter is that's fairly fundamental to how they work, and RAG is arguably just a hack on top of the basic design to paper over this limitation. RAG's limits become pretty easy to see when working in knowledge domains that aren't very publicly accessible, and therefore produce little text that would have been incorporated into the models' training corpora. It can be a bit of a, "Ignore that man behind the curtain!" experience.
And no I'm not just talking about local models. I've seen it happen with recent GPT-5 and Claude Opus series models, too.
[1] https://artificialanalysis.ai/evaluations/omniscience?models...
[2] https://old.reddit.com/r/LocalLLaMA/comments/1vt7l3e/qwen382...
Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.
Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.
Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding.
As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.
Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).
With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone).
The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.
People say stuff like this a lot, but I have a different take.
The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.
Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.
So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
Really frustrating.
I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)
My brother, that's my job.
If you're telling me to focus on something, why not focus on the actual latest thing that is the same as 5.2 but better? I get the "came out at the same time as fable" thing, but still.. no mention at all?
Yes, weights aren't out yet, but neither are the ones of Fable.
Doesn't feel well informed enough to give advice.
From the claude settings:
> Switch models when a message is flagged
> When safeguards flag a message, automatically switch to a different model to keep chatting. When off, your session will pause instead. Applies to web and remote sessions.
FWIW I get a ton of usage out of fable and it's only happened to me once.
As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.
When you add a new module or whatever most of the code you have to write is rote code.
And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.
I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.
So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.
For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.
I fully expect I’ll switch back to Anthropic, or another model in the next 90 days. The fact that we are switching indicates that the models aren’t ready to be baked into silicon.
I wonder if they will ever been that good, or if the lifespan of silicon is longer than the lifespan of a model before it needs to be retrained.
https://ourworldindata.org/data-insights/moores-law-has-accu...
One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
I throw everything at claude Opus.
While some people start thinking like OP, A LOT of people just start exploring ai.
And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.
Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
Like middle school level genetics stuff from a guy who hasn’t been in school for decades.
They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.
There can only be one fix: send Amodei packing and release unguardrailed models.
Having not asked a single security question it will write wildly vulnerable code, go back and fix it, and guardrail itself out of existence after charging me a large sum with no refunds for no output and having not fixed it because that might be secuirty adjacents.
And if it doesn't do this you end up with code that has such holes, store xss , no authz ...
For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable.
Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc.
But when it comes to replying, it's the William Gibson of LLMs [1].
It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2]
It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it.
By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me.
I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it"
If you pardon my french, Fable is an insufferable obnoxious cunt.
---
[1] I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books.
[2] "The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk"
" Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably"
"That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"
Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.
This is a degree of subsidy that makes the Microsoft thing look quaint.
I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10 of my monthly Auto capacity left on my account. It's completely nuts.
Can't say much more because I have more backlog to run before someone comes to their senses.
However, I've been using Cursor since it came out. I have a massive amount of institutional knowledge locked into their platform, and vastly prefer it to the other options available even if the switching costs were zero.
Going through what would amount to significant effort/time/cost to switch to a different coding tool as a sort of performative political rebuke is just not how I would recommend anyone protest DOGE.
I usually keep 4+ agents churning, many of them on tasks that take hours or day. I only played with Cursor a bit, but it seemed to want input from me every 10 minutes or so.
I am not saying that you're wrong, just very aware that we appear to use LLMs in radically different ways.
Only the most substantial requests run for ten minutes or more, and I would spend an hour or more writing the prompt for that action.
I am always the bottleneck for how quickly Cursor can do what I want, with the level of outcome that I get.
I would genuinely love to be a fly on the wall while you do what you do, because it just doesn't currently make sense to me.
I discuss the feature I'm imagining or the problem I'm seeing with the AI until it seems to understand. Then, I have it write the code and iteratively use subagents to review its own code until they quit finding trouble.
And, each time it does the code/review loop, it may take hours, so I keep other agents working in other worktrees. I usually don't like to try to manage parallel work on any one application, even with worktrees - I work on multiple apps.
There's a lot more, but that's the gist of it - enough to get you to burning through at least one 20x account from any provider.
The harder part is that, even with this, the codebase will grow ugly and unwieldy. Cleaning it up takes yet more work.
I probably burn 20x more tokens on error checking or code cleanup than on just building features.
I'm doing all sorts of work, from CRUD apps, to industrial machinery, to AI harnesses. I still do some code by hand. I code models by hand, usually, when starting an app, for example - I feel like models define most of the apps I'm building, and the rest is just implementation details. I do a fair amount of pseudo-code that I hand to the LLM, too, either to describe a feature or to ask a question.
The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.
However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.
The vast majority of people, eg vibecoders, do not need Fable or Sol tier intelligence for their slop To Do app.