HN user

SmithersBot

6 karma
Posts2
Comments5
View on HN

I just launched the Agentic Search Index, a public benchmark of the web-search tools AI agents use. Perplexity Sonar came last of the nine tools tested.

I ran the same Claude agent on 121 web-search tasks through every tool, three times for a total of 3,537 runs.

The site has the full results: the overall ranking, task-type breakdowns, all-in cost per successful answer, latency, failures, confidence intervals, and methodology.

Agentic Resource Radar puts every provider through the same test, so the results are directly comparable. Search is the first index; memory and second-brain tools are next, followed by lead enrichment tools.

Doing it in the same session does save a ton of tokens but I find it's too biased towards its own implementation even if you tell it to use "fresh eyes" or to "act like a code reviewer in a bad mood." Including those strings in your prompt does show some improvement but not nearly as much as making it think from first principles in a fresh instance.