We Benchmarked Frontier LLMs on Defensive Security. The Results Surprised Us 8 months ago
Interesting. I wonder if Gemini 3 reverses that performance trend or if the agent harness lended itself to OpenAI / Anthropic more than Google.
Would like to see this on more open-source agent harnesses and tools.