METR can barely measure Claude Mythos – 50% task horizon now exceeds 16 hours (hugonomy.com) 1 points by GlyphWeaver_a 4mo ago ↗ HN
[–] overthinker_jp 4mo ago ↗ Capability benchmarks may become less meaningful once agents operate across long execution horizons with external tools and permissions. The governance problem starts shifting toward execution boundaries and observability.
2 comments
[ 3.3 ms ] story [ 20.8 ms ] thread