Ask HN: Model Sycophancy Benchmarks? 1 points by gradus_ad 7d ago ↗ HN Model sycophancy is going to become a major issue for individuals and society... Do any benchmarks exist for this? Really needs to become a standard aspect of model quality measures
[–] muddi900 7d ago ↗ I think we should design a "Bullshit Benchmark" that tests for sycophancy and fluff like Opus 5 "two X, but only Y matters" color commentary
[–] tomveber 6d ago ↗ Closest I know of are SycEval and the sycophancy evals in Anthropic's 2023 paper, both built on a user pushing back at a correct answer.
2 comments
[ 0.23 ms ] story [ 12.0 ms ] thread