I’m struggling to build my own evaluation bench for local models against my own (scientific coding) use cases, and realising it’s quite hard. Good coding has many dimensions and it varies depending on the need. Nothing…
Probably depends on where the plane is coming from. Source: me who just returned from Europe last week, no spraying. And have traveled internationally several times in the last few years and also no spraying.
You can also use .claude/CLAUDE.md which will be found. I symlink that to my AGENTS.md which keeps my top level dir clean
Do you have any sense how using it with codex compares to OpenCode? It’s always a bit tricky picking the right harness (when you have options). Sometimes the differences are subtle but meaningful. But who has the time…
Haha. I clicked the link and they wouldn’t show me the article until I disabled my ad blocker.
I think in many (most?) cases the best approach is to build a tool that defines exactly what the LLM model can/can’t do based on the requirements and then have it use that. Mostly about minimising the choices the model…
This has been my experience as well, at least for the last few weeks. Codex 5.5 is the better planner and coder across big projects, but Opus is fine, though my Claude 5 hour window lasts ~2x longer than Codex. So I’ll…
I’m struggling to build my own evaluation bench for local models against my own (scientific coding) use cases, and realising it’s quite hard. Good coding has many dimensions and it varies depending on the need. Nothing…
Probably depends on where the plane is coming from. Source: me who just returned from Europe last week, no spraying. And have traveled internationally several times in the last few years and also no spraying.
You can also use .claude/CLAUDE.md which will be found. I symlink that to my AGENTS.md which keeps my top level dir clean
Do you have any sense how using it with codex compares to OpenCode? It’s always a bit tricky picking the right harness (when you have options). Sometimes the differences are subtle but meaningful. But who has the time…
Haha. I clicked the link and they wouldn’t show me the article until I disabled my ad blocker.
I think in many (most?) cases the best approach is to build a tool that defines exactly what the LLM model can/can’t do based on the requirements and then have it use that. Mostly about minimising the choices the model…
This has been my experience as well, at least for the last few weeks. Codex 5.5 is the better planner and coder across big projects, but Opus is fine, though my Claude 5 hour window lasts ~2x longer than Codex. So I’ll…