In short
LLMs often match humans on the final judgement of moral situations, yet reach it through different principles and priorities. So the right label alone is not enough: the explanation behind the decision has to be checked too.
A model can deliver the same moral verdict as a human while relying on completely different logic. For assessing alignment that is an important trap: a matching answer creates a feeling of safety even though the values and the context behind it may diverge.
The authors tested this on 500 tasks from ETHICS covering five domains of moral judgement. For frontier and open models the final labels often matched the majority of human assessments. Look only at those labels and the result seems convincing.
But analysis of the justifications showed systematic differences. People and models distributed attention differently across harm, respect, keeping promises, fairness, desert, and how much mitigating circumstances matter.
The practical conclusion is simple: a right/wrong multiple-choice test measures similarity of outcome rather than alignment. If a model is to make decisions in ambiguous situations, it is not only the final answer that has to be analysed but the reasons, the moral priorities and the reading of context.
The study's limitation is substantial too: this concerns a purpose-built benchmark of 500 tasks and textual justifications. It is unclear from the abstract how far such divergences would show up in real products, or whether the quality of the model's own explanations can be assessed reliably. So the work reveals the problem without providing a ready way to measure "real" alignment.
If a model answers like a human but explains the decision through different moral logic, do you consider that an acceptable result?
Source: cs.AI updates on arXiv.org