We Benchmarked Frontier LLMs on Defensive Security. The Results Surprised Us (cotool.ai) 6 points by logancarmody 10mo ago ↗ HN
[–] mmpollard 10mo ago ↗ Interesting. I wonder if Gemini 3 reverses that performance trend or if the agent harness lended itself to OpenAI / Anthropic more than Google.Would like to see this on more open-source agent harnesses and tools.
1 comment of 2
[ 7.2 ms ] story [ 78.6 ms ] threadWould like to see this on more open-source agent harnesses and tools.