That is in fact Anthropic's entire reason for existence, the belief that AGI is too dangerous to be controlled by OpenAI/Sam Altman. It naturally follows that it would also be too dangerous to be in the hands of…
I mean, surely the main motivation for "use an LLM to rewrite a huge project in a new language" was excitement about the shiny new tech that made it possible.
TBF, they did it first with ada/babbage/curie/davinci. "Sol" is a much weaker branding, though.
Does anyone understand why LLMs have gotten so good at this? Their ability to generate accurate SVG shapes seems to greatly outshine what I would expect, given their mediocre spatial understanding in other contexts.
The Claude Plays Pokemon stream with a minimal harness is a far more significant test of model intelligence compared to the Gemini Plays Pokemon stream (which automatically maintains a map of everything that has been…
In my experience with the models (watching Claude play Pokemon), the models are similar in intelligence, but are very different in how they approach problems: Opus 4.5 hyperfocuses on completing its original plan, far…
This writeup on the underground puzzle is worth reading, it's a pretty baffling "puzzle" design. https://pokemow.com/Gen2/ShutterPuzzle/ That said, it's definitely Gem's fault that it struggled so long, considering it…
There were no such writeups, 99% of the discussion about difficulties in Crystal were in twitch and discord chats where Google doesn't scrape. (It hadn't yet gotten the public attention that Claude and Gemini's runs of…
It's hard to say for sure because Gemini 3 was only tested with this prompt. But for Gemini 2.5, which is who the prompt was originally written for, yes this does cut down on bad assumptions (a specific example: the…
Exactly what I was going to post. Optimizations like loop unrolling slow down the N64 because keeping the code size small is the most important factor. I think even compilers of the time got this wrong, not just modern…
The busy beaver function is interesting precisely because you _can't_ come up with any computable function that grows faster.
Claude almost universally reacts to everything with a positive exclamation as its first sentence, regardless of whether it's good or bad. If you don't believe me, just watch https://www.twitch.tv/claudeplayspokemon for…
With 2, the real problem is that approximately 0% of the OpenAI employees actually believed in the mission. Pretty much every single one of them signed the letter to the board demanding that if the company's existence…
I've wondered myself why there's so little overlap between these two closely related interests of mine. Some of it seems to be the "But I don't want to cure cancer. I want to turn people into dinosaurs." effect, where…
How long until we get to the point where models know that LLMs get this wrong, and that it is an LLM, and therefore answers wrong on purpose? Has this already happened? (I doubt it has, but there ARE already cases where…
If, like me, you are suddenly curious what would happen if you added a small fourth body: https://youtu.be/WrahPSY9pf0
One of the diagrams says they're using CLIP-G/14 and CLIP-L/14, which are the names of two OpenCLIP models - meaning they're not using OpenAI's CLIP.
You should never ask an LLM to answer questions about itself. The answer is guaranteed to be hallucinated unless Google specifically finetuned it on an answer of that question. The answer it gave you is meaningless.…
So the one thing this article doesn't explicitly justify is whether the ten-byte zlib string is truly the shortest possible. You could imagine that it might be possible to hand-craft a DEFLATE block which is only three…
It has been done - first by OpenAI (MuseNet, which is no longer available) and later by Stanford (Anticipatory Music Transformer): https://nitter.net/jwthickstun/status/1669726326956371971
That is in fact Anthropic's entire reason for existence, the belief that AGI is too dangerous to be controlled by OpenAI/Sam Altman. It naturally follows that it would also be too dangerous to be in the hands of…
I mean, surely the main motivation for "use an LLM to rewrite a huge project in a new language" was excitement about the shiny new tech that made it possible.
TBF, they did it first with ada/babbage/curie/davinci. "Sol" is a much weaker branding, though.
Does anyone understand why LLMs have gotten so good at this? Their ability to generate accurate SVG shapes seems to greatly outshine what I would expect, given their mediocre spatial understanding in other contexts.
The Claude Plays Pokemon stream with a minimal harness is a far more significant test of model intelligence compared to the Gemini Plays Pokemon stream (which automatically maintains a map of everything that has been…
In my experience with the models (watching Claude play Pokemon), the models are similar in intelligence, but are very different in how they approach problems: Opus 4.5 hyperfocuses on completing its original plan, far…
This writeup on the underground puzzle is worth reading, it's a pretty baffling "puzzle" design. https://pokemow.com/Gen2/ShutterPuzzle/ That said, it's definitely Gem's fault that it struggled so long, considering it…
There were no such writeups, 99% of the discussion about difficulties in Crystal were in twitch and discord chats where Google doesn't scrape. (It hadn't yet gotten the public attention that Claude and Gemini's runs of…
It's hard to say for sure because Gemini 3 was only tested with this prompt. But for Gemini 2.5, which is who the prompt was originally written for, yes this does cut down on bad assumptions (a specific example: the…
Exactly what I was going to post. Optimizations like loop unrolling slow down the N64 because keeping the code size small is the most important factor. I think even compilers of the time got this wrong, not just modern…
The busy beaver function is interesting precisely because you _can't_ come up with any computable function that grows faster.
Claude almost universally reacts to everything with a positive exclamation as its first sentence, regardless of whether it's good or bad. If you don't believe me, just watch https://www.twitch.tv/claudeplayspokemon for…
With 2, the real problem is that approximately 0% of the OpenAI employees actually believed in the mission. Pretty much every single one of them signed the letter to the board demanding that if the company's existence…
I've wondered myself why there's so little overlap between these two closely related interests of mine. Some of it seems to be the "But I don't want to cure cancer. I want to turn people into dinosaurs." effect, where…
How long until we get to the point where models know that LLMs get this wrong, and that it is an LLM, and therefore answers wrong on purpose? Has this already happened? (I doubt it has, but there ARE already cases where…
If, like me, you are suddenly curious what would happen if you added a small fourth body: https://youtu.be/WrahPSY9pf0
One of the diagrams says they're using CLIP-G/14 and CLIP-L/14, which are the names of two OpenCLIP models - meaning they're not using OpenAI's CLIP.
You should never ask an LLM to answer questions about itself. The answer is guaranteed to be hallucinated unless Google specifically finetuned it on an answer of that question. The answer it gave you is meaningless.…
So the one thing this article doesn't explicitly justify is whether the ten-byte zlib string is truly the shortest possible. You could imagine that it might be possible to hand-craft a DEFLATE block which is only three…
It has been done - first by OpenAI (MuseNet, which is no longer available) and later by Stanford (Anticipatory Music Transformer): https://nitter.net/jwthickstun/status/1669726326956371971