Nowadays, benchmarks are everywhere. They offer a standardized way to measure and compare different language model capabilities. On every model release, you see a list of typical benchmarks in the model card announcement. As models evolve, their capabilities and use expand, and we need new benchmarks to understand where we are and where we are heading as a community and society. If you are looking for newer benchmarks to understand current model gaps, I’m sharing a few interesting ones to check out and keep an eye on as models make progress and harnesses evolve. They help to understand how models and agents behave when finding bugs, rebuilding software, learning from feedback, and judging their own abilities.
(a) SWE-sweep
How many bugs can Language Models find and fix in large codebases? Given a real repository, an agent must discover and repair as many bugs as they can. Agents are not given any hint about the type of bug or its location.
(b) ./ProgramBench
Can language models rebuild programs from scratch? Given only a compiled binary and its documentation, agents must architect and implement a complete codebase that reproduces the original program’s behavior.
(c) ReviewBench
An open benchmark for AI code review. Measuring how well AI agents find issues in real-world pull requests against human-validated ground truth. Run the benchmark yourself with your own reviewer and compare the results.
(d) EdgeBench
Unveiling scaling laws of learning from real-world environments. An ultra-long-horizon benchmark built to measure learning from environments. This benchmark asks how an agent learns from a real-world environment when it is given the time, the feedback, and the room to improve.
(e) Integrity Bench
Frontier AI seems broadly overconfident about its own ability. How well does a model’s confidence match its actual performance? This benchmark helps show by how much.
(f) Artificial Analysis Cyber Index Alliance
The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities.
https://artificialanalysis.ai/evaluations/artificial-analysis-cyber-index
(g) Mercor APEX-Agents
A benchmark evaluating whether AI agents can execute long-horizon professional tasks in investment banking, management consulting, and corporate law within realistic simulated work environments.


