• Home
  • News
  • Blog
  • Releases
  • LLM history
  • Compare LLMs
  • Library
  • About
⌘K
Sign in

A blog and notes on development. The easiest way to reach me is via the social links below.

Contacts
talalaev.misha@gmail.com
Documents
Personal data processing policyPersonal data processing consent
Photo: CDC / Unsplash

The model knows the theorem — until you rewrite it with another formula

Sh0ny
Sh0ny
11 августа 2026
  1. Home
  2. Blog
  3. The model knows the theorem — until you rewrite it with another formula
2 min read

In short

A new test shows that a language model can lose a familiar theorem when the statement is written in a mathematically equivalent but unfamiliar way. That is an important signal for systems meant to handle formal knowledge reliably rather than merely recognise familiar phrasings.

The nastiest problem with mathematical models may lie not in ignorance of theorems but in fragile access to them. On the TREAT benchmark the best model identified a theorem correctly in only 60.73% of cases when its statement was rewritten in an equivalent form.

That is, the meaning of the statement was preserved but the shell changed: instead of the standard notation came residual equations, witness existence conditions, optimisation identities, set relations or operator forms. For a person familiar with the subject these are different roads to the same object. For a model they are sometimes already a different problem.

TREAT tests precisely the recognition of a theorem's identity. The set holds 737 theorems and 29,480 transformed examples, and for each variant the original assumptions and an inverse mapping to the canonical formulation are preserved. That is more useful than an ordinary paraphrase test: the model cannot get by on similar words and has to grasp that this is the same formal result.

The practical conclusion is simple: if an AI system is to pass knowledge into Mathematica, Lean, Coq or another formal tool, it cannot be tested only on familiar formulations. Equivalent representations are needed, along with explicit checks and the ability to say "I don't know" honestly instead of confidently answering the wrong question.

But the benchmark has boundaries. It is built on theorems collected from pages with mathematical expressions and tests recovery of a theorem's name, not the proof, the correctness of a derivation or the ability to solve a new problem. Moreover the systems erred in different ways: they refused to answer, chose the wrong theorem, or produced a poorly formed result. So 60.73% is not a measure of "mathematical intelligence" in general but an indicator of fragility in a specific, well-controlled scenario.

If you use AI for technical work, which is more dangerous for you: an error in reasoning, or confident recognition of a familiar answer merely because the problem is written in the usual form? Source: cs.AI updates on arXiv.org

новостиaillmразработка
More AI-tool write-ups on the Telegram channel — short and to the point
Subscribe on Telegram

Comments

(0)
​