Interesting idea. Maybe can put it to guide the code reviewer's attention, i.e. which part of the code needs most of their time.
Jeeves pretended to understand. Is there data showing Jev's actually calibrated?
I see this as a problem of people losing grip on the trajectories as models get more capable. Two reasons: you either become more trusting of your agent, or you don't know what it's doing because the CLI is no good at…
Interesting idea. Maybe can put it to guide the code reviewer's attention, i.e. which part of the code needs most of their time.
Jeeves pretended to understand. Is there data showing Jev's actually calibrated?
I see this as a problem of people losing grip on the trajectories as models get more capable. Two reasons: you either become more trusting of your agent, or you don't know what it's doing because the CLI is no good at…