What are some ways to avoid common methological pitfalls when generating test cases for "groundedness" benchmarks with automation?
Confirmation bias is one obvious pitfall that comes to mind, but also I wonder how it is possible to achieve reproducibility when the input is stochastic.
1 comment
[ 2.4 ms ] story [ 55.5 ms ] threadConfirmation bias is one obvious pitfall that comes to mind, but also I wonder how it is possible to achieve reproducibility when the input is stochastic.