In short
Most AI benchmarks test how well a model copes without a human, and by doing so push developers towards replacing the user. The position paper's authors propose evaluating not autonomous AI but the outcome of the human-plus-system pairing.
The problem with today's AI benchmarks may be not that they are too hard but that they test the wrong thing. If a model beats a human on a solo test, that does not yet mean the two of them together solve problems better.
The authors of a position paper argue that the current logic of evaluation implicitly sets a goal: to make AI capable of working instead of a human. Such an approach is convenient to measure — there is a task, an answer and a result. But it poorly reflects scenarios where a human sets the direction, checks the conclusions and makes the decision while the model speeds up the routine part.
The shift proposed is to evaluate the performance of the human–AI team. Then what matters is not the model's own record but whether it improves the human's outcome: does it help find a solution faster, does it reduce the number of errors, does it complement the user's abilities.
That changes the criterion of a "good" system too. A model may not be the best in an autonomous test yet be more valuable in practice if it can explain in time, flag a risk or leave the human in control. For applied AI that is arguably more important than another measure of superiority over people.
But for now this is precisely a position paper, not a set of experimental results. The material presented offers no concrete methodology for evaluating teams, no new metrics and no proof that the approach already yields better results. The main open question is how to measure the contributions of human and AI so that the test does not turn into the user's subjective impression.
If you had to choose between a model that works best alone and a model that noticeably strengthens you on a real task, which should count as the stronger result? Source: cs.AI updates on arXiv.org