I would really love if we brought back some colloquialisms in this field. Not that long ago most folks in tech would have had pretty blank looks on their faces when someone started talking about the "Pareto frontier"
The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).
Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
> Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.
So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.
I think people make mistake here, google’s approach is not to spend $2.3 on every $1.0 earned, they’re riding on serving to masses “luna”, they absolutely have way more powerful models internally but they don’t clutter their infrastructure with fragile and costly intelligence/size frontier. I think “underdog” perception is illusory/temporary.
I tend to agree, but I also should highlight how expensive this shit really is.
in one month, Google actually went cash-negative. [0] even still, they are subsidizing their stuff a lot less, have the most opaque and variable limits, and increase adoption through bundling and shuffling features. I can't even share my Google One storage without subscribing to a Google AI plan anymore, but previously any plan except Google One Lite was shareable.
if you tell me that's not enough to go after frontier, then how much money are Anthropic and OpenAI burning?
Where do you get this "not to spend $2.3 on ever $1.0 earned" from?
Google might not have compelling frontier offerings, their chat harness is complete garbage compared to any other lab (in large part due to a bizarrely badly designed harness where something like code execution requires the prompt to undergo some sort of classification step, no idea what they are doing).
But they absolutely kill in terms of usage offerings. Google lets one subscription be used by *SIX* different google accounts on a family plan.
Plus I currently literally get *$40/month* of Gemini API credits on developer.google.com because they gave me a $10/month grant 4 times.
They give you 200 cloud compute units on google collab, this literally lets you spin up an H100 for around 40 hrs or something if you want to try spinning up local models.
You get Jules (huge allotment btw), Image gen, Video gen, Music Gen, antigravity usage, 5 TB of cloud storage, Notebook LLM...
Okay 5tb cloud storage is their most expensive plan. But what do you actually use video or image gen for? Or music gen? Antigravity is garbage, I guess you can do some stuff with Gemini models over API.
I was paying for Google ai and then realized that between obscure limits and gimmick features I don't really need it. Canceled my subscription and didn't even notice a difference
It's an example number, which doesn't matter much but it comes from widely reported $2.30 spent on ops for every $1 of revenue in spring of 2025 for Anthropic.
The difference between contributing to OS and AI, is that the first is a hobby alternative to woodworking or hiking, while the other can easily bootstrap you a company you can get millions in investment, at least for time being.
"Oh darn, you know that thing I made and released with explicit, precise language defining who can use it and what, if any, restrictions apply? Well now someone is using it in complete accordance with those conditions I set out, and that's somehow making me upset"
On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.
I suspect it might generalize to other large models, but I don't think Qwen3.8 27B is one of them. Kimi K3 is a 2.8 trillion parameter model, and I suspect that is playing a big role in being able to reduce the length of CoT without taking a hit in quality.
Technically kimi k-3 weights license is not open weight (it has a lot of restrictions). I would classify it as ‘weight open’ similar to the bsl and fsl ’source open’ licenses.
There is little to no point reading the article as well. It's stripped of all alpha.
> task and environment feedback
> on-policy planning and learning
> feedback connects decisions to their consequences
These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.
Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.
Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs.
Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds)
Sol is at 2/10 vs kimi’s 3/15
Agreed, I think the only place where it’s still interesting is ui design. Visually kimi and muse feel much nicer than frontier models to me, but maybe it’s an artifact of everything terrible being Claude Design
I was surprised by that. I run my benchmark [1] every couple of days and was sure this model will be ath the pareto frontier, if not THE pareto frontier. But no:
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost.
6-Sol has been (in my experience) terrible. I can't explain why, but it just feels like a weird combination of too dumb and too change-happy (like the old 2024/2025 models).
5.6-Sol and 6-Astra have been worlds better for the tasks I've had them do.
I am glad I am not the only one to notice. I feel like I've gone back to Sonnet 4 levels of incompetence!
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/ammend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
This is astonishingly bad and it is nowhere close to Sol 5.6, or even DeepSeek 4.1! I swear even Gemini 3.8 is slightly better.
I keep saying that with self-hosting, you at least know what to expect and don't have to trust they nerf their models as they go. I was a skeptic and considered nerfing a conspiracy theory, but at this point with enough experience, I have experienced enough to fully see this being a thing.
Probably only a matter of time before some class action happens.
100%, I wish for a legislation which would require the providers to give you at least a unique hash identifying the model (and infra running it, if it affects output) - such that the same hash must give the same output given the same seed. Right now it's all just vibes
This is really interesting. I think the Fireworks Serverless Training infrastructure they used to develop it is also unique and needed. Except if someone works at one of a handful of the largest labs, it is very difficult to set up or try any sort of training pipeline. The managed training infrastructure makes it available to more people.
This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking
It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications
LLMs can be too creative and often too verbose. Sometimes there is a right answer and a way to get there with the understanding of language, but despite using structured outputs, the model insists on inventing variations not in the schema or coming up with something completely different. A model like Jev that can not do those things, and can give the same output every time with given the same input, and be able to measure probabilities has many use cases.
LLM inference has two very different regimes of work: prefill & decode. You can think of the former roughly as processing a pre-specified prompt, and the latter as sequential processing (auto-regressive token generation) eg. "chain of thought". The latter is very important for LLMs and cannot be ignored; it deeply influences infra design, even necessitates copious amounts of high-bandwidth memory. Jev-like models can ignore the latter and therefore optimize much better for the former, consequently operating at both better cost and latency.
The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.
Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.
This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
That's a really impressive result. There are all kinds of small tasks like this I use an LLM for, but theoretically if you broke all the sub-use cases into local-only models, and had something lightweight that routed to the right model, you could have faster and cheaper workflows. E.g. something trained on the linux man pages for common commands, since it's usually quicker to ask an LLM for a specific command with flags than to consult the man pages.
If it’s one of thing that you want just for English to bash shell commands, I will create AST, it is deterministic, exceptionally fast, no tokens so no need to fine tune existing model, please let me know your thoughts.
You can go quite far using a human language to Bash grammar based setup but at some point the input prompts are harder to translate. The OP has existing projects that work with AST quite deeply so I assume they know about that already.
I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.
Golden age before the age that ends humanity. Not talking about any "rogue AI", just the known statistical models of what is coming due to climate change.
I'm literally working on context/harness engineering right now (a set of opencode plugins)
Aside on the aside, I welcome this new era of really personal software. Not Ai's being sycophants, rather being able to easily and quickly change, adapt, or extend software I am not familiar with.
Seriously, I'm using a Qwen 3.8 27B on the homelab, distilled from supposed Fable traces. Regardless, the difference is notable, less thinking, better output. Distilled / heavy quant is better than the original (imv)
I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.
A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.
I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.
I think a sub-set of people see this coming, I also think it isn’t just fine tuning open weight LLM models. A few people I know who are thinking along the same lines with architectures like BERT etc.
That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.
For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!
The issue I find with this is that the frontier models still outperform the small finetuned model on its specific task. So much so that the ROI on doing fine tunes is likely negative. I would love to hear some specific example where it did provide value though, if any has any. That would be helpful to start being able to find similar cases.
There are lots of distilled models on huggingface that are much smaller than, say, Opus and show clear improvements. I do not have the budget to fine-tune a 30b or more parameter model but from my tiny model experiments, the results are quite clear. Again, I have only a couple small tests.
Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?
I think LLMs just made us think that is the only way. It is not.
Maybe we're getting into a world where our general models can make their own separate fine-tuned models as utilities just like they would a bash/python script. If making a fine-tune is cheap and fast then it can also be throw-away and constantly improved.
I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
I have no experience doing this, and I dont mean this to be snarky: what is the power (and thermal) usage of running your CPU only LLM?
When I do something "heavy", my 9800x3D will kick on the fans and start making lot of heat. This is fine when I'm intentionally doing say a transcode, but if I'm just querying syntax it could get pretty annoying. Do you "feel" it? I know that will be pretty machine dependent.
I vaguely recall a project from a while back that did something similar without LLMs.
I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.
I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a directory and strip the prefix” is something I google like once a month).
Over on /r/LocalLLaMA there's a group that's been getting popular doing the same thing for the Qwen 27B (and other) models. - https://huggingface.co/ukisai
Does anyone have a real-world usage feedback on this model? Because the benchmarks look, I would say, too good: it scores even higher than Qwen 3.8 27B in some benchmarks - I am not even sure how is it possible. So the question is if this is benchmaxxing or is it actually good for practical usage like coding in complex projects, research, etc.
Aside. I find the "cost per task" charts both useful and uncanny. Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars? Or a different model that too scores 90% in 1 dollar? How much will it cost me the last 10% or 5%? At the end of the day, cost to 100% is what matters and the half (90%) backed solution may require more to reach 100% (or not, who knows?)
> Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars?
It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...
The 90% and 95% are against some blend of tasks meant to be broadly representative. A pricey model seldom fails a problem that cheap models do well, so there's stratification of tasks by difficulty. Someone doing novel research may be in the "hard" 15% of the blend, where P(solution) goes from one third to two thirds.
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
I see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there
What am I missing here? I think of fireworks as an inference provider serving open weights model. The value that they primarily provide to customers is that (i) they improve reliability by balancing across a bunch of clouds/neoclouds, (ii) they get better pricing by buying capacity in bulk, and (iii) they reduce operational costs. So far so good.
I can also see the argument for providing a post-training service from a customer acquisition perspective: "hey, we can fine-tune this open weights model, so it both gives better/more predictable results than OpenAI/Anthropic and also is cheaper. And btw, once we've won your business, please run this model on our infra."
But what I'm struggling to understand is fireworks spending a bunch of money (on salaries and compute) releasing a frontier model that is going to rapidly fall behind the frontier. Is this "just" advertising for them, both for customers and also for hiring? Or are they actually trying to stay on the frontier? If so, to what end?
Why does Cursor or Devin make their own models? If you use their models, then you can't fallback to other people's models on openrouter or anyplace else. You just stick with them.
Been thinking about the feasibility of training a model using synthetic thinking traces that were reduced to caveman-speak prior to being used for training. Seems like it would be fairly easy to generate plenty of suitably lobotomized synthetic traces with a pair of cheap-ish models. Or even just using good old fashioned NLP to aggressively remove stop words and reduce trace words to lemmas.
It’s the first time I know fireworks has a team doing model research. I do have a complex mood in that. On one hand, I’m always happy to see improvement of OSS models, whether that’s on intelligence or cost-efficiency. On the other hand, I would be a little worried about using fireworks as my API provider. Till the moment I saw this news, I had been using fireworks as my provider of deepseek v4 flash, because I thought fireworks acting as a role deploying OSS models and selling calculation resources, should be safe to use without worry of data being used for training since there’s no “conflict of interests”. But I would think twice now.
this work may explain why recent models like qwen-3.8-flash and MiMo-2.6-* have not made it into their offering, which has given me reason to pause my excitement for Fireworks
Seems irrelevant? Of course we don’t use data for training.
…trust me bro.
It’s obviously easier to believe when they’re not training models.
Eh, anyway this whole thing is just an ad:
> Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload.
Probably, I guess, fancy serverless infrastructure actually makes virtually no difference to hosting really large models that people want to use, and “just” being an inference provider for open weight models turns out to have no moat.
So this is a bit of a pivot to “use our training infrastructure too…!” imo.
Pivot? Sure. Go them. Not what I signed up for though. /shrug
I think legal jurisdiction matters in discerning these things. The EU or US both have courts that, despite anyones opinion, are regarded as having robust contract enforcement. If Fireworks is domiciled in either then claiming ZDR and instead training on the data would be an enormous financial footgun. There's reasons why companies on both sides of contracts, even those with little to no US or EU activity, agree to use US or EU courts for enforcement.
Hong Kong I think used to be a popular option as well, before the handover from the UK.
Being terminally 'online' can make people cynical about everything, but have to temper things with reality a bit too.
Another factor in the cynicism is that robust enforcement is very 'pay to play'. I think your point is accurate for a company like Fireworks, largely doing business with peers that could afford litigation if necessary.
Good luck to an individual or small business trying to hold OpenAI or Meta to account if they breach contract over training, or do something crazy like copyright infringement on an industrial scale.
Last I checked they still offer no training ZDR US based hosting. It is one of my three pinned providers for deepseek v4 flash along with Parasail and Deepinfra.
181 comments
[ 0.16 ms ] story [ 21.1 ms ] threadAnalysis paralysis stifles not just human intelligence, but other intelligences too.
The more options you have, the harder it becomes to be satisfied with the one you picked.
Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?
Kimi K3 with less reasoning tokens isn't exactly exciting either, and particularly so if the license is less open than original Kimi K3.
The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).
I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.
So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.
in one month, Google actually went cash-negative. [0] even still, they are subsidizing their stuff a lot less, have the most opaque and variable limits, and increase adoption through bundling and shuffling features. I can't even share my Google One storage without subscribing to a Google AI plan anymore, but previously any plan except Google One Lite was shareable.
if you tell me that's not enough to go after frontier, then how much money are Anthropic and OpenAI burning?
[0]: https://www.techspot.com/news/113214-google-records-first-ne...
Google might not have compelling frontier offerings, their chat harness is complete garbage compared to any other lab (in large part due to a bizarrely badly designed harness where something like code execution requires the prompt to undergo some sort of classification step, no idea what they are doing).
But they absolutely kill in terms of usage offerings. Google lets one subscription be used by *SIX* different google accounts on a family plan.
Plus I currently literally get *$40/month* of Gemini API credits on developer.google.com because they gave me a $10/month grant 4 times.
They give you 200 cloud compute units on google collab, this literally lets you spin up an H100 for around 40 hrs or something if you want to try spinning up local models.
You get Jules (huge allotment btw), Image gen, Video gen, Music Gen, antigravity usage, 5 TB of cloud storage, Notebook LLM...
The lock-in is less pronounced as it is with AWS or MS.
That's just vibes, though.
> creative
Choose one.
The words of a license are what the license is.
Not suggesting this is right or wrong, but is sort of the nature of the technology.
> task and environment feedback
> on-policy planning and learning
> feedback connects decisions to their consequences
These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.
Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
[1] https://philippdubach.com/posts/jev-model-router-for-pi/
6 or 5.6? Because 6 is hot garbage
5.6-Sol and 6-Astra have been worlds better for the tasks I've had them do.
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/ammend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
This is astonishingly bad and it is nowhere close to Sol 5.6, or even DeepSeek 4.1! I swear even Gemini 3.8 is slightly better.
I keep saying that with self-hosting, you at least know what to expect and don't have to trust they nerf their models as they go. I was a skeptic and considered nerfing a conspiracy theory, but at this point with enough experience, I have experienced enough to fully see this being a thing.
Probably only a matter of time before some class action happens.
This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking
It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications
For example, a typical/stock LLM can’t really play Doom in real time, but a Jev-like model can. Just because of latency
Of course, if you want the best Doom player, there are way better and faster adhoc models
"Pareto frontier": 8 hits
"Opus 5.5": zero hits
I ask cause would this be a kind of model distillation?
I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.
we dont know what the result is and how its impressive.
Or are the subagents generating your training data using a closed/paid model?
I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.
https://github.com/brainless/baho
It's good to hear you're enjoying yourself, but I suggest retiring that expression. It's really beginning to grate.
Aside on the aside, I welcome this new era of really personal software. Not Ai's being sycophants, rather being able to easily and quickly change, adapt, or extend software I am not familiar with.
https://huggingface.co/vwdubb/Qwen3.8-27B-Fable-Distill-NVFP...
side quest, are fable distillations only wrong when it's another country?
Not my project
A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.
I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.
That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.
For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!
Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?
I think LLMs just made us think that is the only way. It is not.
I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
When I do something "heavy", my 9800x3D will kick on the fans and start making lot of heat. This is fine when I'm intentionally doing say a transcode, but if I'm just querying syntax it could get pretty annoying. Do you "feel" it? I know that will be pretty machine dependent.
I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.
I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a directory and strip the prefix” is something I google like once a month).
Can you (or your agent) please write a tutorial or share some good links.
(Namely the fine tuning part.)
I'd like to learn how to do this as well.
It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
I see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there
I can also see the argument for providing a post-training service from a customer acquisition perspective: "hey, we can fine-tune this open weights model, so it both gives better/more predictable results than OpenAI/Anthropic and also is cheaper. And btw, once we've won your business, please run this model on our infra."
But what I'm struggling to understand is fireworks spending a bunch of money (on salaries and compute) releasing a frontier model that is going to rapidly fall behind the frontier. Is this "just" advertising for them, both for customers and also for hiring? Or are they actually trying to stay on the frontier? If so, to what end?
this is our preferred open weight token vendor
this work may explain why recent models like qwen-3.8-flash and MiMo-2.6-* have not made it into their offering, which has given me reason to pause my excitement for Fireworks
…trust me bro.
It’s obviously easier to believe when they’re not training models.
Eh, anyway this whole thing is just an ad:
> Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload.
Probably, I guess, fancy serverless infrastructure actually makes virtually no difference to hosting really large models that people want to use, and “just” being an inference provider for open weight models turns out to have no moat.
So this is a bit of a pivot to “use our training infrastructure too…!” imo.
Pivot? Sure. Go them. Not what I signed up for though. /shrug
Hong Kong I think used to be a popular option as well, before the handover from the UK.
Being terminally 'online' can make people cynical about everything, but have to temper things with reality a bit too.
Good luck to an individual or small business trying to hold OpenAI or Meta to account if they breach contract over training, or do something crazy like copyright infringement on an industrial scale.