1 comment

[ 3.9 ms ] story [ 14.5 ms ] thread
LLM evaluations are very sensitive to the details of the prompt's structure. This post shows how using structured generation reduces the results' variance and the ranking shifts.