Part of the problem is that from that perspective there was no independent analysis done that could vindicate them. METR is absolutely part of the EA/LessWrong/rationalist ecosystem, so of course their investigation…
This stacks with the 50% discount in OpenRouter, making it $2/$10. https://openrouter.ai/openai/gpt-5.6-sol
In my experience LLMs do find a lot of embarrassing bugs as Linus says but it can constantly turn into a game of whack-a-mole where most of the bugs it finds were written in previous LLM sessions. It's a huge struggle…
I know the company I consult for (not cybersecurity) is not in these programs and if attacked would need to use open weight models.
The Hugging Face incident is a great example of why open source models with defensive cyber capabilities are needed. Hugging Face did not have access to cyber-capable frontier models and kept hitting safeguards. Only by…
The null is surely different in this case because we know that supplemented creatine goes to the brain and the brain uses it, which is definitely not a claim all supplements can make. The fact that there's a…
Is this methodologically similar to Anthropic's recent J-Space paper?
The task could be verifiable in the environment so limiting its CPU and RAM could be to discourage brute forcing the answer.
Both are true! The Confederacy did secede largely to preserve slavery, but the war was started to bring the Confederacy back into the Union, initially without the goal of also immediately abolishing slavery.
LLMs are good at writing complex regex, from my experience
When I use ChatGPT for work it frequently reads my Slack DMs even if they’re not directly relevant, so I’d question a lot of the premises of the article.
“they’ve been folded into Frankenstein positions that demand constant multitasking, social performance, sensory endurance, and emotional labor on top of technical skill” My sense is that this is due to automation, not…
It would be interesting if most of our confusion with quantum mechanics came from treating probabilities as independent when they are actually highly correlated. I don’t really know any physics, but I’m familiar with…
SubredditSimulator was a markov chain I think, the more advanced version was https://reddit.com/r/SubSimulatorGPT2
It's so easy to ship completely broken AI features because you can't really unit test them and unit tests have been the main standard for whether code is working for a long time now. The most successful AI companies…
Happy smart people are generally very focused and only care about a few things. Unhappy smart people are constantly getting "nerd-sniped" into focusing their intelligence on things that don't make them happy.
> It’s the same reason why most of the people who pass your leetcode tests don’t actually know how to build anything real. They are taught to the test not taught to reality. True, and "Agentic Workflows" are now playing…
[flagged]
An answer to the productivity paradox (https://en.m.wikipedia.org/wiki/Productivity_paradox) could be that increased technology causes increased complexity of systems, offsetting efficiency gains from the technology…
Part of the problem is that from that perspective there was no independent analysis done that could vindicate them. METR is absolutely part of the EA/LessWrong/rationalist ecosystem, so of course their investigation…
This stacks with the 50% discount in OpenRouter, making it $2/$10. https://openrouter.ai/openai/gpt-5.6-sol
In my experience LLMs do find a lot of embarrassing bugs as Linus says but it can constantly turn into a game of whack-a-mole where most of the bugs it finds were written in previous LLM sessions. It's a huge struggle…
I know the company I consult for (not cybersecurity) is not in these programs and if attacked would need to use open weight models.
The Hugging Face incident is a great example of why open source models with defensive cyber capabilities are needed. Hugging Face did not have access to cyber-capable frontier models and kept hitting safeguards. Only by…
The null is surely different in this case because we know that supplemented creatine goes to the brain and the brain uses it, which is definitely not a claim all supplements can make. The fact that there's a…
Is this methodologically similar to Anthropic's recent J-Space paper?
The task could be verifiable in the environment so limiting its CPU and RAM could be to discourage brute forcing the answer.
Both are true! The Confederacy did secede largely to preserve slavery, but the war was started to bring the Confederacy back into the Union, initially without the goal of also immediately abolishing slavery.
LLMs are good at writing complex regex, from my experience
When I use ChatGPT for work it frequently reads my Slack DMs even if they’re not directly relevant, so I’d question a lot of the premises of the article.
“they’ve been folded into Frankenstein positions that demand constant multitasking, social performance, sensory endurance, and emotional labor on top of technical skill” My sense is that this is due to automation, not…
It would be interesting if most of our confusion with quantum mechanics came from treating probabilities as independent when they are actually highly correlated. I don’t really know any physics, but I’m familiar with…
SubredditSimulator was a markov chain I think, the more advanced version was https://reddit.com/r/SubSimulatorGPT2
It's so easy to ship completely broken AI features because you can't really unit test them and unit tests have been the main standard for whether code is working for a long time now. The most successful AI companies…
Happy smart people are generally very focused and only care about a few things. Unhappy smart people are constantly getting "nerd-sniped" into focusing their intelligence on things that don't make them happy.
> It’s the same reason why most of the people who pass your leetcode tests don’t actually know how to build anything real. They are taught to the test not taught to reality. True, and "Agentic Workflows" are now playing…
[flagged]
An answer to the productivity paradox (https://en.m.wikipedia.org/wiki/Productivity_paradox) could be that increased technology causes increased complexity of systems, offsetting efficiency gains from the technology…