The term that Anthropic is now using is "mannered prose". If the creator of this example simply prompted "Remove all mannered prose" then the entire experiment would suddenly become normal sounding. In the fable 5.1…
This post isn't convincing me. I spent so much time meticulously organizing my techno box. I bought a back of the door shoe holder for tech. Every wire, charger, usb key, web cam, airline earbud, everything. It's been…
>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs Not sure…
The government will pay because it's not his money, it's our money. He loves spending our money...
I'm sorry, but just because you achieve results you consider acceptable with this method doesn't mean everyone does. I don't work where we can ship slop. I don't work where PRs can be merged based on what the agents…
It's not free. You're paying electricity and you're ignoring the cost of the hardware. Even on electricity alone, there are cloud providers who may beat your laptop on price per million tokens. Qwen 3.8 flash is…
Sonnet 5 is the worst model of 2026. Literally just turn effort slider down on Opus, it's smarter, faster and cheaper than whatever Sonnet is. Beyond that, I find this whole plan and build thing to be a pointless waste…
Those prices are just tokens? Since each model uses different amounts of tokens to do the same thing, it's a misleading price that often makes open-weights look more competitive than they are, since most open weights…
>At some point one has to wonder if it's still worth using anthropic's models if we need to babysit 100% of its output with another vendor's model. Why not just use that other vendor's model for everything? If you're…
I absolutely experienced this in college. I signed up as a computer science student, as one does. I took all of the freshmen classes across broad topics, and the first biology class was basically just like the article…
I append this to many of my opus claude code prompts `You may use a Fable subagent to answer questions, solve problems, and provide an adversarial review of your ideas and code` You can use a similar pattern in most any…
It's pretty easy for the US to functionally ban chinese models. They only have to target US firms like inference providers or the biggest users, and pretty much the whole domestic market will fall into line. They don't…
I totally agree - designing a competent AI agent with a fully customized harness to successfully pull off this task is a much more challenging engineering effort than merely creating an ordinary computer program. Had OP…
What a sloppy reply. You've hijacked a thread on mathematics first to complain that your incompetent attempt to use ChatGPT to find a job failed, but it seems now that this was a ruse to instead begin arguments…
[flagged]
A business does need a small number of their most senior engineers doing high altitude work that can, at times, include helping sales estimate new features. But in my experience, it's not rocket science and a good…
I think you're confusing product and engineering. I get that programmers are smart so we just assume we can do every job, but it's a waste of your time and salary to talk extensively to customers and create product…
Scenario one: you use software to connect to their server and download a webpage. You are a user. Scenario two: you use software to connect to their server and download a webpage. You are a "bot". Make it make sense
[dead]
GPT5.6Sol completes the suite in 70M tokens, while Qwen3.8Max needs like 145M tokens. So this is a case where models like Qwen 3.8 and Kimi K3 use a lot more output (reasoning) tokens, go a good bit slower, so they can…
While I agree with the premise that there are Thinkers and Shippers, I reject labeling of tinkerers and entrepreneurial. There's nothing entrepreneurial about working for a big business and shipping cool things. But…
There is already something on HuggingFace at the level of Mythos. It's called Kimi K3 and it's running laps around Opus5, Fable5, and Sol56 at cybersecurity. It's so good that the US government is rushing to ban all…
>No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models. Are you talking about Kimi's training cost or the training cost of the model(s) that Kimi distilled? Because…
I don't think the invention of writing is as awe-inspiring as presented or as difficult/impossible for an LLM to achieve as is commonly believed. The invention of writing was a long series of micro-improvements over…
Zen is nice, but they require US hosting so they don't get new Chinese models right away. There is no Kimi K3. Go is nice for the ten minutes you can use it until your hit your cap.
The term that Anthropic is now using is "mannered prose". If the creator of this example simply prompted "Remove all mannered prose" then the entire experiment would suddenly become normal sounding. In the fable 5.1…
This post isn't convincing me. I spent so much time meticulously organizing my techno box. I bought a back of the door shoe holder for tech. Every wire, charger, usb key, web cam, airline earbud, everything. It's been…
>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs Not sure…
The government will pay because it's not his money, it's our money. He loves spending our money...
I'm sorry, but just because you achieve results you consider acceptable with this method doesn't mean everyone does. I don't work where we can ship slop. I don't work where PRs can be merged based on what the agents…
It's not free. You're paying electricity and you're ignoring the cost of the hardware. Even on electricity alone, there are cloud providers who may beat your laptop on price per million tokens. Qwen 3.8 flash is…
Sonnet 5 is the worst model of 2026. Literally just turn effort slider down on Opus, it's smarter, faster and cheaper than whatever Sonnet is. Beyond that, I find this whole plan and build thing to be a pointless waste…
Those prices are just tokens? Since each model uses different amounts of tokens to do the same thing, it's a misleading price that often makes open-weights look more competitive than they are, since most open weights…
>At some point one has to wonder if it's still worth using anthropic's models if we need to babysit 100% of its output with another vendor's model. Why not just use that other vendor's model for everything? If you're…
I absolutely experienced this in college. I signed up as a computer science student, as one does. I took all of the freshmen classes across broad topics, and the first biology class was basically just like the article…
I append this to many of my opus claude code prompts `You may use a Fable subagent to answer questions, solve problems, and provide an adversarial review of your ideas and code` You can use a similar pattern in most any…
It's pretty easy for the US to functionally ban chinese models. They only have to target US firms like inference providers or the biggest users, and pretty much the whole domestic market will fall into line. They don't…
I totally agree - designing a competent AI agent with a fully customized harness to successfully pull off this task is a much more challenging engineering effort than merely creating an ordinary computer program. Had OP…
What a sloppy reply. You've hijacked a thread on mathematics first to complain that your incompetent attempt to use ChatGPT to find a job failed, but it seems now that this was a ruse to instead begin arguments…
[flagged]
A business does need a small number of their most senior engineers doing high altitude work that can, at times, include helping sales estimate new features. But in my experience, it's not rocket science and a good…
I think you're confusing product and engineering. I get that programmers are smart so we just assume we can do every job, but it's a waste of your time and salary to talk extensively to customers and create product…
Scenario one: you use software to connect to their server and download a webpage. You are a user. Scenario two: you use software to connect to their server and download a webpage. You are a "bot". Make it make sense
[dead]
GPT5.6Sol completes the suite in 70M tokens, while Qwen3.8Max needs like 145M tokens. So this is a case where models like Qwen 3.8 and Kimi K3 use a lot more output (reasoning) tokens, go a good bit slower, so they can…
While I agree with the premise that there are Thinkers and Shippers, I reject labeling of tinkerers and entrepreneurial. There's nothing entrepreneurial about working for a big business and shipping cool things. But…
There is already something on HuggingFace at the level of Mythos. It's called Kimi K3 and it's running laps around Opus5, Fable5, and Sol56 at cybersecurity. It's so good that the US government is rushing to ban all…
>No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models. Are you talking about Kimi's training cost or the training cost of the model(s) that Kimi distilled? Because…
I don't think the invention of writing is as awe-inspiring as presented or as difficult/impossible for an LLM to achieve as is commonly believed. The invention of writing was a long series of micro-improvements over…
Zen is nice, but they require US hosting so they don't get new Chinese models right away. There is no Kimi K3. Go is nice for the ten minutes you can use it until your hit your cap.