Yes, I can churn out a lot more stuff as can most of my peers. Experiments etc are all way faster to run with coding agents. But I think the overall creativity and originality is a lot lower. I think this is what many…
I didn't say I'm immune to those effects, I'm including myself in this as well. (also, I'm not older than my colleagues). Most people definitely can't meditate for 30 minutes, so if you can do this, it's very…
I have some sympathy for these kids. If LLMs were around when I was a student, I would've also used them to "speed up" my homework assignments then proceed to fail all my tests. Now I work mostly with PhDs who were at…
LLM written article. It's also not accurate; the fact that language models have human-interpretable representations and neurons has been known since BERT. Circuits research also does not come from Anthropic. Mech interp…
Huh, according to that model card this is a 137B total parameter model. Performance doesn't seem that good: - MAI-Code-1-Flash (137B-A5B) = 51% on SWE-bench pro - Qwen3.6-35B-A3B = 49.5% on SWE-bench pro…
Is this training data even valuable? Usually AI data annotators get paid to write LLM responses, but here all they'd be getting is a bunch of user queries.
Yes, I can churn out a lot more stuff as can most of my peers. Experiments etc are all way faster to run with coding agents. But I think the overall creativity and originality is a lot lower. I think this is what many…
I didn't say I'm immune to those effects, I'm including myself in this as well. (also, I'm not older than my colleagues). Most people definitely can't meditate for 30 minutes, so if you can do this, it's very…
I have some sympathy for these kids. If LLMs were around when I was a student, I would've also used them to "speed up" my homework assignments then proceed to fail all my tests. Now I work mostly with PhDs who were at…
LLM written article. It's also not accurate; the fact that language models have human-interpretable representations and neurons has been known since BERT. Circuits research also does not come from Anthropic. Mech interp…
Huh, according to that model card this is a 137B total parameter model. Performance doesn't seem that good: - MAI-Code-1-Flash (137B-A5B) = 51% on SWE-bench pro - Qwen3.6-35B-A3B = 49.5% on SWE-bench pro…
Is this training data even valuable? Usually AI data annotators get paid to write LLM responses, but here all they'd be getting is a bunch of user queries.