I like the 2x2 grid that describes when to fine-tune a model, when to use a frontier model, etc.
From the article it’s not clear how the scorer grades every episode - was it a frontier model that assigned the grade? How does that continue to work as the model that is being fine-tuned becomes better at the task than the frontier model?
What I really started to notice is that SOTA models are really good at putting themselves out of the job.
We can see this already with GPT how luna can do 90% of what sol is used for. The only reason why china still bothers 'distilling' models is accurate training data generation, something that oai and anthropic had to spend years collecting while trying to dodge legal challenges.
The more intelligent models get, the more people will offramp to cheaper solutions that get the job done. There's no real benefit to using a sota model when the accuracy is already 99% and I think that is the biggest danger to US labs.
This is a general theme with technology and the 'S-curve'. Let's not even get into whether the improvement for AI reasoning ability has slowed - for practical purposes of writing a React frontend, it has.
But other tech is like that - I don't even remember when I bought my LCD TV - 2018 I think? I have no inclination of buying a new one.
Technology has a tendency to replace new technology, or intrude into vacant areas, but its very rare for technology to replace non-technology (like human interaction).
Most of the recreation humans do in front of screens is tending to (para)social relationships.
I didn't read the Ramp article but this reads like a post hoc fallacy. Companies with 2x revenue have money to spend on AI. Companies with 1.15x revenue don't.
Share of the maximum achievable score our GRPO-trained 9B open-source model reached on catalog review, vs 76.9% for the best frontier configuration: a 13.5% relative improvement over the frontier, and 36% over its own untrained base (64.2%). The five frontier models, even with optimized prompts, plateaued within a tenth of a point of each other; the trained specialist cleared that ceiling."
_______
This is hard for me to believe. I have a lot of skepticism that frontier models like GPT 5.5 that are likely 2T+ parameters in size only got about 12% more accurate than an untrained 9b parameter LLM.
This continuous cycle of fine-tuned open models beating frontier on (often vaguely labeled/defined) benchmarks doesn't provide an accurate comparison to the expanding generalized capabilities of the SoTA, which makes them effectively meaningless.
If we were to take these at face value, why is it that the frontier labs' models are making legitimate new discoveries (e.g. Erdős and Jacobian conjectures) and these models are not?
To me, a better signal of capability would be similarly performing novel work at the same or better level, which they presently are not. I say this as someone who very much looks forward to open models being more capable, but to deny the gap is misguided hopeful hype.
The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot.
As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does the economic picture that justified the massive infrastructure building that’s now broadly funded by a complex network of debt.
This is what makes open weight models so threatening to them. The political and “it’s China” angle is mostly just a cover for the real reasons why they’re freaked out.
The fact that models are now a pure commodity is bad enough for the big labs. If small open weight models become the norm the big labs are toast.
Fine timing takes time and data, though. If smart enough models get cheap enough, then most people have lots of use cases that are cost insensitive enough that it's not worth the effort.
Of course, "smart enough" is a low enough threshold for most uses that this is still a problem for the frontier labs.
But at the same time, a truly smart enough closed model could also potentially command almost whatever they'd care to charge for it.
Whether they can actually get to that level remains to be seen, but I can definitely see a situation where most people are perfectly happy with cheap middle of the tree models while large corporations pay magnitudes more than current API pricing for access to models never even marketed as a mass market product and keep the labs afloat.
It's of course be a lot easier for them to find the path towards that of they didn't need to compete with open models in the meantime.
But then can't the case be made that narrower and stricter-defined use cases are better served by more conventional ML? If/Wherever efficiency is a concern, that is.
Yeah and these latest and greatest models universally suck hard for any use case outside of the few 'blessed' ones. Like I'm sure their ability to write prose and generally sound like a human being has regressed quite a bit, but even if not. Opus 5 is barely above GPT4 when it comes to stuff like home improvement advice.
Going by the chart in the article, if your total workload is 1k cataloged items and your quality threshold is 70%, why wouldn't you just pay $19 to gemini API instead of $500 + time to make a custom fine tune?
> The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages.
Most generic use cases do benefit from a model that has been trained broadly. When you don’t know the specific use case ahead of time, you have to have world knowledge ready to go. Even when coding it’s helpful to have all that knowledge on tap so the model can understand product intent and use cases for the product you’re building.
> As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles
The real expensive part of fine tuning is gathering a good data set. That’s been the hard part of training anything for a long time. If you’re lucky enough to have a neatly organized and clean data set then you can attempt it, but you need to be in a position to run evals and measure quality.
I’ve done it, but there are so many use cases where the engineering, data labeling, and ongoing quality review hours cost so much that it would be cheaper to continue using a frontier lab model that just works from the start.
Right now, we keep creating new models with larger and larger weights and throwing hardware at it. His argument is that this is a dead end and ultimately we'll eventually go back to purpose fit algorithms like we've always done in the history of AI development.
I actually prefer using less powerful models most of the time, I might use a stronger model initially and then switch to weaker models as I fine tune the output.
Offloading too much of your task to models eliminates the human ownership, without ownership you can't move forward. Context sizes can't keep up with large codebases and markdown files with instructions and guidelines only take the LLM so far in the ownership aspect.
My gut feeling is that to start offloading ownership to the LLM we would need to see at least two order of magnitude increases on context size.
I’m not sure it will go this direction. Using a generic (near) frontier model is often cheap enough that you need to talk really big volumes before it pays off to fine tune.
My example: we were doing single digit millions of automated call summaries a few years back at a major bank with GPT-4o. Smaller model gave more rejected summaries (compliance not happy), so we briefly looked at fine tuning a smaller model and basically concluded that even at that scale the effort of data collection, management, fine tuning, hosting the model, etc didn’t have a sufficient business case vs picking up other projects.
I mostly agree, but at the same time I wouldn't blindly ignore the "it's China" angle at all. These are models you cannot open up and do a thorough audit of. I do think local compute will eventually be good enough that most people will just use local models. Microsoft might be the ultimate winner of that if they can stop ruining Windows and focus on making an OS similar to how straightforward and simple macOS is, no ads literally everywhere for Microsoft Office. Just make a good OS Microsoft, is that so much to ask? You HAD a really solid OS and you ruined it.
>As does the economic picture that justified the massive infrastructure building that’s now broadly funded by a complex network of debt.
I have a genuine question and I'd like to hear people's good faith thoughts on this.
There's a fair case that open models are threatening to institutions who spent a lot on training proprietary SOTA models.
But to my understanding, the massive investment spend (much of it debt backed as you note) is on data centers, chips, physical infra.
Yes, this infra is needed to train the models, but it is also needed to serve inference.
Perhaps the costs associated with training SOTA models is ultimately a "bust" given open models eroding the SOTA closed model performance advantage.
But demand for inference is skyrocketing and there seems to be no end in sight.
The physical hardware underpinning inference is in fact a scarce good (currently, and this seems sustainable at least over mid-term). And inference is a scarce service as such.
I know that cost of inference constantly goes down as models, technical infrastructure, and applied AI techniques become more efficient (specialized SLMs etc). So this puts downward pressure on prices.
But still... demand for inference is just growing like crazy regardless. Putting upwards pressure on prices.
Doesn't this mean that all the spending on AI infra is much better positioned to get positive ROI regardless of the type of model being served?
Put another way, models seem to be commoditizing, but physical hardware is not (currently).
The vast majority of the AI boom spend is on hardware to my understanding (even training capex can be repurposed for inference).
Doesn't this suggest that the economics for the "railroads level of build-out spending" are healthier than they might seem at first glance?
> The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages.
I think that's true across the board. Along those lines, most companies all of a sudden started looking for AI researchers, instead of software engineers, which is what they still need :-).
Just like a decade-ish ago, when the FAANG interview style got openly known, then everybody started to conduct interviews in that pattern :-)
Im pretty sure they know quite well but its not their current focus on creating or allowing others to create small finetuned and optimized models.
In the race they are in, its still highly beneficial to be the frontier model.
I use the frontier model every single day through my company and my company happily spends these tokens.
It makes a huge difference if different people can use one interface to do everything.
Big models give you fundamental things: A lot of facts/context, usability (you might be a native english speaker, don't underestimate how hard it is for A LOT of people to formulate what they want/need in english only AND a low complexity.
No one needs to build a router and x sub models and a router architecture. You literaly just have an API, you might choose the model and the effort but thats it.
I find this current state of the art a LOT more telling on the current progress we are in than anything else. I'm confident that small optimized models will become a lot more relevant like lets say java + english + a second language + spring boot + postgresql. It might also be beneficial in the long term to finetune with your project details.
For a lot of very technical non human interfacing things, finetuning is happening left and right.
But what i find very interesting is emerging complexity capabillity. I believe that fables skill to hold more topics and combine more complex solutions together is because of its parameter size.
We will have to figure out if we can extract this complexity out of it while reducing the training data in a way that the training data focus more on thinking. Plenty of smaller thinking models show that this is doable.
Btw. Mixture of Experts is for sure not optimal for this, but it already is a form of optimized sub models. Perhaps we might just have MoE with a million experts in the future. One per lanuage + area of expertise etc.
Also don't forget: IF AGI is coming through a current frontier model, you will let it work for hours, days and weeks on one problem completly independent of any human input and it will be better than a human. If they reach this before a collapse, we are done and they 'won'. For this you need big frontier models.
I want small models to win and I am constantly experimenting with them. I have never tried fine-tuning and do not have that kind of budget. My approach is to remove some of the burden from models and bring into the agent.
Tool calling is an example - in some tasks RAG works really well, including coding agents where code, git log, Epics/Tasks, dependencies sources, etc. are all available in very structured manner. You can save many extra tool calls if you can run separate prompts and retrieve the source data needed for the actual work - rather its prompt.
And I really want to focus on search - this is the key technology if we want to use RAG instead of fine-tuning. If we can present really contextual sources in the prompts using a hybrid search approach - you can see how easily we get better results - either decisions or summaries from even small models.
Every time I see this kind of story, two things bother me.
First, I have watched the free improvement of frontier models surpass the gains from retraining, many times now. Squeezing more out of the models that already exist, or simply doing nothing and waiting, is a real strategy and it often pays better. The fair comparison is not against today's frontier but against whatever ships while you are still maintaining your fine-tune.
Second, the $500 training bill is the cheapest line item in this story. The expensive parts are creating the data and maintaining the model afterwards. How many use cases can actually produce 177k scored episodes? Here they had to generate them synthetically from Amazon Berkeley Objects. To me, that dataset is the strongest evidence in the article of how hard fine-tuning is to apply: if the data existed naturally, nobody would need to manufacture it.
Regarding your first point, if the cost of not doing it while you wait exceeds the cost of training/maintaining the model, then this approach still makes sense. This really depends on how fast you expect cheap (and open-source) models to improve. So it might only be a temporary strategy, but still worthwhile. Also, I expect that specialized models will always be better (cheaper or better outcomes) than a general model, similar to how specialized HW like GPUs is still used even though CPUs have improved a lot too.
Maintaining in the way of ongoing training is relatively cheap, considering that the initial training run is only 500$ (in this example). Creating new examples to account for drift in the training data is more expensive, but these examples can then be reused when training a new model. And to some degree you need them anyways, to evaluate the models and prompt changes you make.
Overall, as you pointed out, this won't always make sense (either due to the cost of creating examples and training, or simply the lack of available data). But they point this out themselves in the diagram towards the bottom of the page: it only makes sense for frequent and verifiable tasks. This might result in this only being a sensible approach for very large companies (they talk about millions of decisions), but it might make sense for them.
Many of the most valuable companies in the world (who, incidentally, are spending the most money on frontier model inference), already have "the expensive part", the labeled dataset.
The cost of maintaining the datasets and the models is getting commoditized by startups like braintrust and huggingface.
Amazing! I've been in a similar autoresearch-y rabbit-hole lately with getting Apple's 3B Foundation Model to match Sonnet 4.6 on a very specific task.
The result was: 90% parity achieved with a weird combination of a fine-tuned adapter + 1 deterministic step.
fine tuning a small LLMs or even real small language models (like bert) is what the recommended way since the introduction of LLM. the benefit is pretty much obvious: faster to run, fully controlling the stack, better fit for the custom domain...
but in practice not many people do the fine tuning, for pretty much a single reason: large language models API cost are still very cheap, fast enough, and get improvement all the times. what the point of spend time (and money) to fine tune a specified model, then just when you release it a newer gen generic model is released and beat it?
but if someday the progress for LLM is slowed, or the price increased to the point calling api is not a viable approach any more, then surely the day of fine tuning and small models will come again.
I have great interest in fine-tuning open models, and I'm looking for resources that HN folks can personally recommend. This article looks good and I've bookmarked it to read more thoroughly over time.
I've gotten as far as running Nemotron-3-Nano 30b locally, and plan to target models around 30b - 120b parameters. Based on brief examination of the results I can get, I think these vanilla models are capable enough to add real value, but training could push them over the finish line for specialized tasks.
What I really appreciate is that the author is thinking about the whole process, which is also my goal. Confirmation that others are identifying the same use case, and the same strategy for adding value using this technology.
This is a long term project, so I plan to buy hardware to conduct the fine-tune. ..
WTF is happening here? It's a clear marketing piece with clear bunch of bot commenting like the article is a real deal. Is this a regular the modern day HN experience?
60 comments
[ 0.24 ms ] story [ 105 ms ] threadFrom the article it’s not clear how the scorer grades every episode - was it a frontier model that assigned the grade? How does that continue to work as the model that is being fine-tuned becomes better at the task than the frontier model?
We can see this already with GPT how luna can do 90% of what sol is used for. The only reason why china still bothers 'distilling' models is accurate training data generation, something that oai and anthropic had to spend years collecting while trying to dodge legal challenges.
The more intelligent models get, the more people will offramp to cheaper solutions that get the job done. There's no real benefit to using a sota model when the accuracy is already 99% and I think that is the biggest danger to US labs.
But other tech is like that - I don't even remember when I bought my LCD TV - 2018 I think? I have no inclination of buying a new one.
Technology has a tendency to replace new technology, or intrude into vacant areas, but its very rare for technology to replace non-technology (like human interaction).
Most of the recreation humans do in front of screens is tending to (para)social relationships.
Share of the maximum achievable score our GRPO-trained 9B open-source model reached on catalog review, vs 76.9% for the best frontier configuration: a 13.5% relative improvement over the frontier, and 36% over its own untrained base (64.2%). The five frontier models, even with optimized prompts, plateaued within a tenth of a point of each other; the trained specialist cleared that ceiling."
_______
This is hard for me to believe. I have a lot of skepticism that frontier models like GPT 5.5 that are likely 2T+ parameters in size only got about 12% more accurate than an untrained 9b parameter LLM.
If we were to take these at face value, why is it that the frontier labs' models are making legitimate new discoveries (e.g. Erdős and Jacobian conjectures) and these models are not?
To me, a better signal of capability would be similarly performing novel work at the same or better level, which they presently are not. I say this as someone who very much looks forward to open models being more capable, but to deny the gap is misguided hopeful hype.
As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does the economic picture that justified the massive infrastructure building that’s now broadly funded by a complex network of debt.
This is what makes open weight models so threatening to them. The political and “it’s China” angle is mostly just a cover for the real reasons why they’re freaked out.
The fact that models are now a pure commodity is bad enough for the big labs. If small open weight models become the norm the big labs are toast.
Of course, "smart enough" is a low enough threshold for most uses that this is still a problem for the frontier labs.
But at the same time, a truly smart enough closed model could also potentially command almost whatever they'd care to charge for it.
Whether they can actually get to that level remains to be seen, but I can definitely see a situation where most people are perfectly happy with cheap middle of the tree models while large corporations pay magnitudes more than current API pricing for access to models never even marketed as a mass market product and keep the labs afloat.
It's of course be a lot easier for them to find the path towards that of they didn't need to compete with open models in the meantime.
P.S. Building something that proves that you don't need those many params even.
https://github.com/guilt/tinytot
Most generic use cases do benefit from a model that has been trained broadly. When you don’t know the specific use case ahead of time, you have to have world knowledge ready to go. Even when coding it’s helpful to have all that knowledge on tap so the model can understand product intent and use cases for the product you’re building.
> As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles
The real expensive part of fine tuning is gathering a good data set. That’s been the hard part of training anything for a long time. If you’re lucky enough to have a neatly organized and clean data set then you can attempt it, but you need to be in a position to run evals and measure quality.
I’ve done it, but there are so many use cases where the engineering, data labeling, and ongoing quality review hours cost so much that it would be cheaper to continue using a frontier lab model that just works from the start.
https://www.youtube.com/watch?v=21EYKqUsPfg
Right now, we keep creating new models with larger and larger weights and throwing hardware at it. His argument is that this is a dead end and ultimately we'll eventually go back to purpose fit algorithms like we've always done in the history of AI development.
Offloading too much of your task to models eliminates the human ownership, without ownership you can't move forward. Context sizes can't keep up with large codebases and markdown files with instructions and guidelines only take the LLM so far in the ownership aspect.
My gut feeling is that to start offloading ownership to the LLM we would need to see at least two order of magnitude increases on context size.
My example: we were doing single digit millions of automated call summaries a few years back at a major bank with GPT-4o. Smaller model gave more rejected summaries (compliance not happy), so we briefly looked at fine tuning a smaller model and basically concluded that even at that scale the effort of data collection, management, fine tuning, hosting the model, etc didn’t have a sufficient business case vs picking up other projects.
I have a genuine question and I'd like to hear people's good faith thoughts on this.
There's a fair case that open models are threatening to institutions who spent a lot on training proprietary SOTA models.
But to my understanding, the massive investment spend (much of it debt backed as you note) is on data centers, chips, physical infra.
Yes, this infra is needed to train the models, but it is also needed to serve inference.
Perhaps the costs associated with training SOTA models is ultimately a "bust" given open models eroding the SOTA closed model performance advantage.
But demand for inference is skyrocketing and there seems to be no end in sight.
The physical hardware underpinning inference is in fact a scarce good (currently, and this seems sustainable at least over mid-term). And inference is a scarce service as such.
I know that cost of inference constantly goes down as models, technical infrastructure, and applied AI techniques become more efficient (specialized SLMs etc). So this puts downward pressure on prices.
But still... demand for inference is just growing like crazy regardless. Putting upwards pressure on prices.
Doesn't this mean that all the spending on AI infra is much better positioned to get positive ROI regardless of the type of model being served?
Put another way, models seem to be commoditizing, but physical hardware is not (currently).
The vast majority of the AI boom spend is on hardware to my understanding (even training capex can be repurposed for inference).
Doesn't this suggest that the economics for the "railroads level of build-out spending" are healthier than they might seem at first glance?
I think that's true across the board. Along those lines, most companies all of a sudden started looking for AI researchers, instead of software engineers, which is what they still need :-).
Just like a decade-ish ago, when the FAANG interview style got openly known, then everybody started to conduct interviews in that pattern :-)
In the race they are in, its still highly beneficial to be the frontier model.
I use the frontier model every single day through my company and my company happily spends these tokens.
It makes a huge difference if different people can use one interface to do everything.
Big models give you fundamental things: A lot of facts/context, usability (you might be a native english speaker, don't underestimate how hard it is for A LOT of people to formulate what they want/need in english only AND a low complexity.
No one needs to build a router and x sub models and a router architecture. You literaly just have an API, you might choose the model and the effort but thats it.
I find this current state of the art a LOT more telling on the current progress we are in than anything else. I'm confident that small optimized models will become a lot more relevant like lets say java + english + a second language + spring boot + postgresql. It might also be beneficial in the long term to finetune with your project details.
For a lot of very technical non human interfacing things, finetuning is happening left and right.
But what i find very interesting is emerging complexity capabillity. I believe that fables skill to hold more topics and combine more complex solutions together is because of its parameter size.
We will have to figure out if we can extract this complexity out of it while reducing the training data in a way that the training data focus more on thinking. Plenty of smaller thinking models show that this is doable.
Btw. Mixture of Experts is for sure not optimal for this, but it already is a form of optimized sub models. Perhaps we might just have MoE with a million experts in the future. One per lanuage + area of expertise etc.
Also don't forget: IF AGI is coming through a current frontier model, you will let it work for hours, days and weeks on one problem completly independent of any human input and it will be better than a human. If they reach this before a collapse, we are done and they 'won'. For this you need big frontier models.
Tool calling is an example - in some tasks RAG works really well, including coding agents where code, git log, Epics/Tasks, dependencies sources, etc. are all available in very structured manner. You can save many extra tool calls if you can run separate prompts and retrieve the source data needed for the actual work - rather its prompt.
And I really want to focus on search - this is the key technology if we want to use RAG instead of fine-tuning. If we can present really contextual sources in the prompts using a hybrid search approach - you can see how easily we get better results - either decisions or summaries from even small models.
First, I have watched the free improvement of frontier models surpass the gains from retraining, many times now. Squeezing more out of the models that already exist, or simply doing nothing and waiting, is a real strategy and it often pays better. The fair comparison is not against today's frontier but against whatever ships while you are still maintaining your fine-tune.
Second, the $500 training bill is the cheapest line item in this story. The expensive parts are creating the data and maintaining the model afterwards. How many use cases can actually produce 177k scored episodes? Here they had to generate them synthetically from Amazon Berkeley Objects. To me, that dataset is the strongest evidence in the article of how hard fine-tuning is to apply: if the data existed naturally, nobody would need to manufacture it.
Maintaining in the way of ongoing training is relatively cheap, considering that the initial training run is only 500$ (in this example). Creating new examples to account for drift in the training data is more expensive, but these examples can then be reused when training a new model. And to some degree you need them anyways, to evaluate the models and prompt changes you make.
Overall, as you pointed out, this won't always make sense (either due to the cost of creating examples and training, or simply the lack of available data). But they point this out themselves in the diagram towards the bottom of the page: it only makes sense for frequent and verifiable tasks. This might result in this only being a sensible approach for very large companies (they talk about millions of decisions), but it might make sense for them.
Many of the most valuable companies in the world (who, incidentally, are spending the most money on frontier model inference), already have "the expensive part", the labeled dataset.
The cost of maintaining the datasets and the models is getting commoditized by startups like braintrust and huggingface.
LLMs are great because they can handle open domain problems in part because they are generative.
The result was: 90% parity achieved with a weird combination of a fine-tuned adapter + 1 deterministic step.
I wrote the whole thing down here: https://alexisrondeau.me/tada/research/FMDiscovery/dashboard... which includes the question, the answer, the 96 experiments and their lineage etc. etc.
but in practice not many people do the fine tuning, for pretty much a single reason: large language models API cost are still very cheap, fast enough, and get improvement all the times. what the point of spend time (and money) to fine tune a specified model, then just when you release it a newer gen generic model is released and beat it?
but if someday the progress for LLM is slowed, or the price increased to the point calling api is not a viable approach any more, then surely the day of fine tuning and small models will come again.
I've gotten as far as running Nemotron-3-Nano 30b locally, and plan to target models around 30b - 120b parameters. Based on brief examination of the results I can get, I think these vanilla models are capable enough to add real value, but training could push them over the finish line for specialized tasks.
What I really appreciate is that the author is thinking about the whole process, which is also my goal. Confirmation that others are identifying the same use case, and the same strategy for adding value using this technology.
This is a long term project, so I plan to buy hardware to conduct the fine-tune. ..