Interesting tests being done but I can't help but think it limits testing innovation in some way given that the requested apps are essentially all clones of others
It's really interesting to see the Sol/Terra/Luna apps side-by-side.
I need to add these stats somewhere in the UI, but one interesting take away: Terra took 1/2 as much wall-clock time as Sol, but Luna took more wall-clock time than Sol (by about 23%). It's still much much cheaper, but it seems like Terra is likely a more optimal time/cost balance for most use cases.
The Terra quality is usually nearly as good as Sol, but much faster and cheaper. I do appreciate Sol's design sensibilities (see, for example, the audio sequencer). It's the first model in a while that is clearly distinct on that front. They'd all converged to very similar visuals for a while.
"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict.
Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.
> We generated a big pile of artifacts, we are publishing all of them, and you can form your own opinion.
My opinion is that spamming HN with two gimmicky "one-shot prompting shootout" marketing pieces in two days does not build confidence about either your technical or marketing expertise.
My concern with most of these visual benchmarks, popular as they are, is that they are likely more indicative of knowledge (i.e. how comprehensive the training data is and how well it can be retrieved from the model) than of reasoning ability. I don't see in particular how a model would construct a CoT that mapped somehow to a representation of the cube geometry and its animations in latent space without a large chunk of that being pre-existing information.
It’s interesting how all the model names and versions are like SKUS taking up space on a display shelf. I look forward to whatever Sagittarius A* does!
> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all.
> so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and GLM are the slowpokes
You put in a lot of good work, and kudos for that, but man, reading paragraphs like these just puts me off of the entire piece.
Like…how hard would it have been really to type these two sentences by hand, in your own natural voice?
Is this how I learn that Bezos now has a beard? Interesting that it is a detail that all of the models chose to include (unless that was in the prompt and just not put in the post).
Obviously AI-written, but I'm confused with the results: Muse Spark has the best Rubik's cube by far, the only one properly animating, yet it gets a 2/5
Really nice breakdown, surprised by the results - especially the fact that OSS models were so behind on most task... (lol at the SVG of the moon without any sign of life by GLM-5.2)
Missing the exact prompts - would love to replicate...but also curious how you prompted these: they could be a big reason why some models failed completely at rendering SVGs (ie. GLM 5.2)
A lot of these are visual-heavy tests that often require first person sight to confirm results. Considering GLM isn’t multimodal, that might explain why it did better on the calculator question and not much else.
I think there's approximately zero value in seeing how a model can turn 100 tokens into a 100k. What workflow is that? It's not useful in the real world.
I want to know how well it can follow instructions, manage various potentially competing desires in the context, and so on. It's much more interesting how it can turn 100k tokens (e.g. a codebase and lots of tool calls) into 100 tokens.
29 comments
[ 0.22 ms ] story [ 43.9 ms ] threadhttps://arena.logic.inc/
It's really interesting to see the Sol/Terra/Luna apps side-by-side.
I need to add these stats somewhere in the UI, but one interesting take away: Terra took 1/2 as much wall-clock time as Sol, but Luna took more wall-clock time than Sol (by about 23%). It's still much much cheaper, but it seems like Terra is likely a more optimal time/cost balance for most use cases.
The Terra quality is usually nearly as good as Sol, but much faster and cheaper. I do appreciate Sol's design sensibilities (see, for example, the audio sequencer). It's the first model in a while that is clearly distinct on that front. They'd all converged to very similar visuals for a while.
Nothing hard. Everybody wins.
My opinion is that spamming HN with two gimmicky "one-shot prompting shootout" marketing pieces in two days does not build confidence about either your technical or marketing expertise.
Agent: https://arena.ai/leaderboard/agent
Web dev: https://arena.ai/leaderboard/code/webdev
Currently Fable and 5.6 are neck and neck on web dev which is basically the same finding as this.
> so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and GLM are the slowpokes
You put in a lot of good work, and kudos for that, but man, reading paragraphs like these just puts me off of the entire piece.
Like…how hard would it have been really to type these two sentences by hand, in your own natural voice?
(edit: seems to be an issue with inline videos)
surprised it isn't a bigger thing, eg artificial analysis doesn't report anything like that
still doesn't measure the human-agent interaction part, but that's pure vibes atp
> Draw a horse riding an astronaut in svg
https://www.svgviewer.dev/s/if4gi3e7
I want to know how well it can follow instructions, manage various potentially competing desires in the context, and so on. It's much more interesting how it can turn 100k tokens (e.g. a codebase and lots of tool calls) into 100 tokens.
We made Grok 4.5, GPT-5.5, and Claude build the same apps - https://news.ycombinator.com/item?id=48838772 - July 2026 (92 comments)
Real world is messy, other benchmarks are clearly gameable by the Chinese open models.
Great job! And I don’t care about the tone of the article, it’s readable just fine.