• Home
  • News
  • Blog
  • Releases
  • LLM history
  • Compare LLMs
  • Library
  • About
⌘K
Sign in

A blog and notes on development. The easiest way to reach me is via the social links below.

Contacts
talalaev.misha@gmail.com
Documents
Personal data processing policyPersonal data processing consent
Photo: Cristian Pineda / Unsplash

When two channels of perception make AI worse rather than better

Sh0ny
Sh0ny
18 August 2026
  1. Home
  2. Blog
  3. When two channels of perception make AI worse rather than better
2 min read

In short

The new test checks not recognition of a picture or a sound but the ability to reconstruct an unseen result from hand movement and the sound of the pen. People scored above 80% accuracy, while GPT-4o and Gemini 2.5-Pro did not clear 10% — and combining audio with video often made the answer worse.

Multimodal models confidently recognise what they have been shown. But when you need to understand what happened beyond the visible, their confidence runs out fast.

In The Unwritten Benchmark the word cannot be read with the eyes: the ink has been removed. The model receives only video of hand movements and audio of the pen scratching a surface. It has to reconstruct the written word, in three different writing styles at that.

That is an important difference from an ordinary test of sight or hearing. Here it is not enough to notice individual cues — the speed of movement, the direction of a stroke, the sound of the pen. They have to be assembled into a hidden sequence and their causal link understood: which movement and which sound correspond to a particular letter.

The result was harsh. People showed ordered letter-recognition accuracy above 80%, while leading multimodal models, including GPT-4o and Gemini 2.5-Pro, could not exceed 10%.

The most unexpected part is that adding the second channel did not help. In a number of cases feeding video and audio together made the result worse than using a single modality. It seems the models cannot reliably synthesise complementary signals when the answer has to be inferred from a dynamic process rather than found in a ready image.

The study's limitations matter too: the work tests one specific task — recovering words from micro-movements of the hand and the sound of a pen. So the result cannot be turned directly into a conclusion about the uselessness of multimodal models in general. But it does reveal a weak spot: recognising the observable and reasoning about an invisible cause are different abilities.

If a model can simultaneously see a movement and hear it yet errs more often as a result, should it be considered genuinely multimodal? Source: cs.AI updates on arXiv.org

NewsaiNeural networksllm
More AI-tool write-ups on the Telegram channel — short and to the point
Subscribe

Comments

(0)
​