In short
Most tests of moral reasoning in LLMs check whether answers match human values, but hardly test how norms are applied in context. The breakdown shows why such an AI can sound ethical and still make the wrong decisions.
The problem is not that language models necessarily "fail to understand morality". The problem is that we often test only half the skill in them — which values they call correct, but not whether they can apply a norm to a specific situation.
That is an important distinction. It is easy to agree with the proposition "do no harm". It is far harder to determine what counts as harm in a given context, which circumstances change the assessment of an act, and which norm applies here at all.
The paper's authors call this the gap between two problems. The first is the moral value problem: does the model's answer match human moral values. The second is the moral norm problem: can the model recognise a situation and correctly apply a context-dependent rule.
Existing evaluations, in the authors' view, are noticeably skewed towards the first problem. So a high score may mean not developed moral reasoning but the ability to reproduce socially acceptable phrasing.
Proper testing requires at least three things: good reference data on how norms are applied, evaluation of the intermediate reasoning, and a check of which features of a situation the model treats as morally relevant. The authors also propose formalising normative theories and separating the level of values from the level of norms in tests.
The limitations of this approach are obvious too. The paper openly notes the shortage of quality expert-annotated data, insufficient attention to the reasoning process and the absence of settled formal representations of norms. Until those exist, test results are easy to misread: to take a convincing moral answer for the ability to make moral decisions.
The practical conclusion for developers is simple: if a model is used in a sensitive context, what needs testing is not only its declarations about good and evil but its handling of specific cases with shifting circumstances. Otherwise we are measuring a moral vocabulary rather than moral competence.
Where in your tasks does the model err more often: in choosing values, or in applying the right rule to a specific situation?
Source: cs.AI updates on arXiv.org