• Home
  • News
  • Blog
  • Releases
  • LLM history
  • Compare LLMs
  • Library
  • About
⌘K
Sign in

A blog and notes on development. The easiest way to reach me is via the social links below.

Contacts
talalaev.misha@gmail.com
Documents
Personal data processing policyPersonal data processing consent
Photo: Brecht Corbeel / Unsplash

A new benchmark tests whether AI can discover the rules from scratch

Sh0ny
Sh0ny
15 августа 2026
  1. Home
  2. Blog
  3. A new benchmark tests whether AI can discover the rules from scratch
2 min read

In short

DiG-bench removes the thing that usually helps AI most: a goal known in advance. Instead the agent has to run experiments, notice regularities and even guess the winning conditions — much like a miniature piece of scientific research.

The most interesting part of the new DiG-bench benchmark is not the 70 games but the absence of instructions for them. The agent is not told the transformation rules and is not even told what counts as winning. First you have to work out the task itself, and only then solve it.

That differs markedly from the usual tests where the goal is stated in advance: write code, answer a question or pick the right option. Here what is tested is a chain closer to research: form a hypothesis, run an experiment, discard a wrong idea and carry the discovered rule over to the next level.

Each game is defined by a short string and has rules of its own. Seven difficulty levels are provided in all. The lowest level, according to the authors, is regularly solved by several models, while the highest already causes trouble even for strong models working inside agent systems.

And that is DiG-bench's practical value: it separates the ability to follow an instruction from the ability to make sense of an unknown environment. A model can look very convincing on tasks with a clear formulation yet get lost when it first has to determine which task to solve at all.

The limitations are substantial too. Of the 70 games only 21 have been published publicly, the rest kept closed for protected evaluation. So independent verification of the full set is impossible. Besides, the fact that at least one person solved every game on the first attempt does not mean it is easy for people on average — the human bar here is drawn rather crudely.

If AI is given neither goal nor rules, what matters more for progress: a stronger model, or a more skilful cycle of experiments around it? Source: cs.AI updates on arXiv.org

новостиaiагентынаука и техника
More AI-tool write-ups on the Telegram channel — short and to the point
Subscribe on Telegram

Comments

(0)
​