Wandr Benchmark: Evaluating Research Agents That Must Search Wide and Deep (research.perplexity.ai) 1 points by tagawa 1mo ago ↗ HN
[–] tagawa 1mo ago ↗ Repo with benchmark tasks, evaluation harness, tech report:https://github.com/perplexityai/wandr
1 comment
[ 0.22 ms ] story [ 12.3 ms ] threadhttps://github.com/perplexityai/wandr