2 comments

[ 47.4 ms ] story [ 677 ms ] thread
Author here, Life lately: build a hard eval → new model mogs it → build a harder eval → mogged again → repeat. Cafe Bench is the latest round this, it is a small eval ( world sim ) we built at Dot to test and track progress of the product.

But the data points transfer well into generalised model benchmakrs.

surprises: Opus 5.5 made +$227k for $8.57 in 22 minutes; GPT-6 Astra made +$157k but cost $56; and GPT-6 Luna lost $60k on average.

really interesting results, it seems we really crossed another threshold with yesterdays releases where sol 5.6 still lost money, but sol 6 is in the green.