In short
An audit of four popular agent-safety benchmarks revealed that the F1 metric allows a trivial policy to outperform real models, and that the ranking depends on the size of the panel. Performance and safety are negatively correlated—but only on small samples.
When you read “our agent is safe because it scored X on R-Judge,” you’re most likely looking at a number that means almost nothing. An audit of four popular agent-safety benchmarks—R-Judge, InjecAgent, AgentHarm, and AgentDojo—shows that their scores are cited as interchangeable, even though they measure different behaviors and differ in their rankings of the same models.
Source: cs.AI updates on arXiv.org