do you somehow control how non-trivial the queries are? The LLM generates them, right? what if every engine returns garbage, or on the other hand, handles them too well? building a benchmark like this in a genuinely…
do you somehow control how non-trivial the queries are? The LLM generates them, right? what if every engine returns garbage, or on the other hand, handles them too well? building a benchmark like this in a genuinely…