Brief
Benchmark Radar indexes 1,283 benchmark records and ships a CLI for prior-art searches
A paper describes a "living database and search engine" for AI benchmarks, combining daily discovery with a searchable catalog, score histories and downloadable evidence, and says it releases a dashboard, CLI and reproducible analysis. The evidence gives no pricing, licence or hosting details, so the access story is unresolved.
Benchmark Radar is described as a living database and search engine for AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety and domain-specific evaluations. Daily discovery draws on 37 sources — 13 direct connectors and 24 first-party research and engineering feeds — and the catalog holds 1,283 source records drawn from 4 benchmark catalogs, with 12,916 numeric observations on 790 records.
The detail that matters for this desk is the prior-art workflow: the authors walk through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. The paper also audits the full catalog and examines benchmark saturation, adoption trends and the limits of score comparisons, and says records retain source identities and citations.
Our reading
For teams that must choose or design evaluations, the prior-art search path and the retained citations are the useful part — they turn "which benchmark should we use?" into an inspectable query rather than a literature guess. The paper's own examination of benchmark saturation and the limits of score comparisons is a caution worth carrying into any leaderboard reading. The people who should care…
What to do or watch
Next step: run a prior-art search against the catalog or CLI before designing a new evaluation, and inspect the cited evidence for each candidate benchmark. What remains unresolved is how the dashboard and CLI are hosted, licensed or priced, since the evidence states none of that.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- Benchmark Radar is presented as a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations.
- Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds.
- The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records.
- The system retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence.
- A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation.
- The release includes the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a CLI for offline queries, and reproducible analysis.
Sources
- arXivText stored 16 September 2026
How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.
What that means
- 6 of 6 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.