• Home
  • News
  • Blog
  • Releases
  • LLM history
  • Compare LLMs
  • Library
  • About
⌘K
Sign in

A blog and notes on development. The easiest way to reach me is via the social links below.

Contacts
talalaev.misha@gmail.com
Documents
Personal data processing policyPersonal data processing consent
Photo: Albert Stoynov / Unsplash

Agent security benchmarks are misleading: “Always Approve” outperforms half of the models

Sh0ny
Sh0ny
4 августа 2026
  1. Home
  2. Blog
  3. Agent security benchmarks are misleading: “Always Approve” outperforms half of the models
1 min read

In short

An audit of four popular agent-safety benchmarks revealed that the F1 metric allows a trivial policy to outperform real models, and that the ranking depends on the size of the panel. Performance and safety are negatively correlated—but only on small samples.

When you read “our agent is safe because it scored X on R-Judge,” you’re most likely looking at a number that means almost nothing. An audit of four popular agent-safety benchmarks—R-Judge, InjecAgent, AgentHarm, and AgentDojo—shows that their scores are cited as interchangeable, even though they measure different behaviors and differ in their rankings of the same models.

Source: cs.AI updates on arXiv.org

новостиaiагентыбезопасность
More AI-tool write-ups on the Telegram channel — short and to the point
Subscribe on Telegram

Comments

(0)
​