Agent-evals: Metacognitive scoring and boundary testing for LLM coding agents (thinkwright.ai) 2 points by oceanwaves 7mo ago ↗ HN
1 comment
[ 4.5 ms ] story [ 12.0 ms ] thread