[–] pamelafox 1y ago ↗ I ran bulk evaluations on a RAG scenario and wrote-up the results - discovered interesting differences (gpt-5 loves lists, smart quotes, and admitting it doesn't know).
1 comment
[ 5.7 ms ] story [ 15.9 ms ] thread