In short
Today: how to review code created by agents and what’s happening with agentic web search. Plus several useful analyses of benchmarks, instructions, and model failures.
Today: how to verify code created by agents and what’s happening with agentic web search. Plus several useful analyses of benchmarks, instructions, and model failures.
🔥 Hot:
🔹 MAGS proposes formally verifying coding agents’ outputs — The authors consider multi-agent autoformalization as a way to find errors that fuzz tests, static analysis, and LLM-based verifiers may miss. 🔹 Study compares web search in ChatGPT, Claude, Grok, and DeepSeek — The paper examines the entire agentic search cycle: search decisions, strategies, results, and final answers. 🔹 New study examines whether AI agents understand computer architecture — Improving hardware design does not yet prove that an agent understood how the system works rather than simply getting lucky while searching through parameters.
➡️ Useful materials:
🔹 Researchers propose virtualizing foundation models through a self-evolving OS layer — The idea is intended to unify state, memory, budgets, and safety constraints in composite agentic systems. 🔹 Study shows that benchmarks reflect changing expectations of LLMs — The authors are interested not only in model rankings, but also in which capabilities researchers consider successful in the first place. 🔹 Paper analyzes what modern tests of systematic generalization are missing — The authors propose viewing such tasks through the lens of reasoning, rather than merely simplified combinations of actions. 🔹 Analysis explains seven productivity traps when working with AI — The piece explores how the promise of “getting more done” can lead to overload and unhealthy work habits.
➡️ Discussions and cases:
🔹 Practitioner explains why an agent’s “stupidity” often begins with the instructions — The analysis traces the problems to three causes: rules written in prose, rules formulated as prohibitions, or the lack of result verification. 🔹 Study revisits the Trusting Trust attack for self-editing AI coders — The paper raises the issue of infecting systems capable of modifying their own code. 🔹 User showed how Ternary Bonsai 2 27B gets stuck in a loop on a complex task — Instead of creating a game involving planetary exploration, the model repeats the same line of reasoning.
📝 If you’d like to add other news and materials to the list, write in the comments.