Ask HN: What default model do you use and why?
I use claude for most of what I do, and Fable is largely overkill for me and frequently burns through my Max plan's session credits in minutes (!) when just doing an initial mobile app planning with 4 agents. After I waited out the timeout period 6 hours later, and I picked up again, the cache had timed out so it burned through 2% of the session in less than a minute. Opus 4.8 is now my goto and I will be avoiding 5 until I see a reason to switch back. 4.8 is 'good enough' for what I need and has been a great value. It mostly gets things right. Most of what I do is web and mobile , largely cloud backend.
102 comments
[ 0.29 ms ] story [ 12.6 ms ] threadFor more long work, I now use Fable to create a PLAN.md. I tell it to make a plan that will be executed by other models, and most of the time it ends up choosing Opus or Sonnet for the actual work.
I didn't start doing this recently. Before that, I would just use the top model for everything. Splitting the work across different models depending on the task has helped a lot. They run faster, and I usually get better results
most engineering tasks don’t require frontier llms
when they get stuck, then i consider moving up to more capable models
purchasing a claude plan seems widely unnecessary to me
the tasks they do better than the average engineer cut both ways: unless you have an existing portfolio of well written and designed work done pre-llms, it looks like you’re producing slop that pretends to be well designed
poor typography choices despite using the mode,
poor layout choices despite using popular CSS frameworks
etc
bad engineers will always be bad engineers
tools don’t make up for it
edit: a follow up to this— everyone is using eyebrows in their layouts and have no fucking clue why it was done to begin with
everyone has a status pill floating above their front page hero display text and its not fucking status related
so gross
- like a fancy auto complete (here are some stub methods, they should do X, fill them in)
- using fairly detailed plans and test harnesses, so blowing up the world is hard
The 3.X Flash family have been fairly capable models, and the selling point for me is just raw speed. Gemini is noticeably faster than the competition, about 3-4x, and I just get work done faster with it.
That said I'm keeping an eye on Open Weights. DS4 Flash was good until price hikes, and finding a provider that serves at high speed and without quantisation at the prior price is tricky.
There's no Gemini pro model currently, so you gotta pair that with a 20 openai plan for access to more advanced stuff if you need it.
A few months ago that would've been bigger model stuff, but now you can do that in a couple hours with Flash and direction.
- Preferred: Claude Code with Opus 5 Medium
I also don't want to use the Claude Code and Codex agent harnesses. The good thing with Codex subscription is that it can be used in other harnesses, unlike Claude. As far as I know, only Anthropic has this restriction.
It's really strange because when Opus 5 released, there were some that pointed this out, but a bunch simply said it was the best and as good as Fable etc etc.
but for me, it caused me to get very demotivated and avoid interacting with the model, at least when using Claude Code.
I've moved to Codex 5.6-Sol. Much saner English, much better at execution, and gets stuff done in a matter-of-factly kind of way (Claude Code is a mess these days -- it gets things wrong and goes around in circles).
But I'm harness agnostic and am not locked in. I just keep my issues in Kata Tracker (https://www.katatracker.com/) and switch harnesses/model when I need to.
Being loyal to a particular model/harness seem unwise to me.
I suspect this degradation is happening because the AI labs are using the LLM's output to feedback into the input, to create a thinking loop, and they're optimizing that.
Now I’ve gone back to Codex because I simply find the ChatGPT models much more comfortable to work with. I spend less time fighting the model, correcting its direction, or re-explaining what I meant.
unbeatable price/intel ratio per M tokens:
$0.10 (input)
$0.20 (output)
$0.002 (cached-input)
Granted, yesterday I threw a few tasks to Astra which the former 2 botches; it produced clean, correct solutions quickly, so pending further eval, this may take over.
IMO unless it's a mechanical tasks, it's worth it to use carefully -crafted queries on the more expensive models, than iterate through messier solutions on the cheaper ones.
For my personal stuff, I'm on a small $20 plan, so I need to use tokens conservatively. I was very rarely exceeding limits until I built a Dark Software Factory. It's not as efficient at token use. So, I use Sonnit over Opus here.
At work I have a $100 plan that I rarely exceed so I use Opus. I have access to Fable too, and I did use it a lot while it was new, but I don't find it improves most of my work by too much. I do mostly bug fixing across several hundred repositories with hundreds of thousands of lines of code, mostly written by humans over the past 20-years.
I also use Sol as a secondary for my personal work. I pay for it because I like to talk to ChatGPT. Since I already have the subscription, I let Sol write plans for me. It does a better job at certain tasks and it saves me some Claude tokens. Maybe I should consider Terra for the task, but I don't run up against my usage limits for the little bit I use it.
https://joeldare.com/creating-a-minimal-dark-factory
So far I’ve created a couple simple web based games and a BASIC32 interpreter that boots on an ESP32 and turns it into a 1980’s style computer.
...and then sprinkle in some other models when i think a second opinion will help
Privacy policy is not great with deepseek api, but you are always just trusting their word with any hosted llm and in theory i could at least self host the models i use from deepseek.
I prefer Cursor at this point just because of Composer. Claude Code is passable but the lack of a good, fast, cheap workhorse sucks. Sonnet and Haiku aren’t it.
I also do like Codex and Sol, Terra, and Luna. They are decent but I don’t find they stand out enough to use over the others.
Additionally, I have not tried Astra and found Fable to really not worth the cost for the tasks I do.
Finally, the latest Grok is actually a beast of a model, but expensive enough to not be a stand out.
A day ago I activated Google AI Pro free via Google's tie-up with a local company (I do pay for this company's product though and it's anything but costly). Now I will use this too.
No other reasons to pick these, or not picking anything else.