In short
Governing autonomous agents like ordinary DevSecOps is no longer enough: the most dangerous failures appear at the interaction level, where system behaviour cannot be derived from the behaviour of individual models. CASE proposes splitting governance across four levels and looking for the weakest one.
The problem with corporate AI agents is not that companies have no means of control at all. The problem is different: those means mostly check individual agents, whereas real failures are often born of their collective behaviour.
CASE's authors propose treating agent governance as four different tasks. For a single agent you need control theory: the goal as a set-point, constraints as feedback, evaluation as observation. For a collective you need the theory of complex adaptive systems, because emergent behaviour cannot be reliably assembled from checks on individual participants. For a team of humans and agents you need the cybernetics of oversight, and for a fleet of agents you need operations engineering, including error budgets not just for technical availability but for the quality of decisions.
The most important conclusion here is that controlling one layer can make things worse at another. You can, for example, automate deployment almost completely in technical terms, but then human oversight risks becoming a formality. In CASE's terms that is not an accidental process flaw but a conflict between levels of autonomy.
The authors call the gap between risk and available tooling the Emergence Gap. By their data, 82% of documented production agent failures unfold across several levels at once, none of the 22 tools studied fully covers the collective behaviour level, and all 35 public deployments assessed fell into the lowest maturity category. So a checklist for an individual agent can create a feeling of control exactly where there is least of it.
The approach has limitations too. CASE remains a proposed architecture and evaluation model rather than a universally proven standard: the authors themselves speak of the need to account for cross-level links and the paradox of hands-off control. Besides, the figures given describe the studies the paper is based on, not an independent audit of every corporate deployment. A practical question is open as well: exactly how companies are to maintain the required diversity of human oversight across a rapidly changing fleet of agents.
If autonomous systems are going to make decisions collectively, what does your organisation currently check more carefully: the individual agent, or the consequences of their interaction? Source: cs.AI updates on arXiv.org