• Home
  • News
  • Blog
  • Releases
  • LLM history
  • Compare LLMs
  • Library
  • About
⌘K
Sign in

A blog and notes on development. The easiest way to reach me is via the social links below.

Contacts
talalaev.misha@gmail.com
Documents
Personal data processing policyPersonal data processing consent
Photo: Growtika / Unsplash

Web agents learn on the fly, while LLMs learn to compress without losing their reasoning.

Sh0ny
Sh0ny
10 September 2026
  1. Home
  2. Blog
  3. Web agents learn on the fly, while LLMs learn to compress without losing their reasoning.
1 min read

In short

In today's roundup: new approaches to adapting and compressing AI models, as well as tools for evaluating agents and conducting systematic reviews. At the end, a discussion of whether Astra has reached the level of AGI.

In today’s roundup: new approaches to adapting and compressing AI models, as well as tools for evaluating agents and systematic reviews. At the end, a discussion of whether Astra has reached AGI-level capabilities.

🔥 Hot:

🔹 Researchers propose budget-friendly online adaptation for web agents — The approach is designed for lightweight local models that need to learn after deployment. 🔹 A new LLM compression method protects critical reasoning chains — Reasoning-Aware Compression abandons uniform quantization of all model components to reduce energy consumption without unnecessarily compromising quality.

➡️ Useful materials:

🔹 SCAFFOLD turns a web agent’s accumulated skills into a hierarchical library — The agent does not start every task from scratch but reuses procedural knowledge across changing interfaces. 🔹 SciLitBench tests LLMs at every stage of a systematic literature review — The benchmark covers everything from title and full-text screening to structured data extraction. 🔹 AutoFyn trains long-term agents without changing model weights — The system stores verified results in a persistent state and brings them into new sessions through explicit interfaces. 🔹 CriticGen turns LLM response evaluation into concrete improvement recommendations — The method takes the generation process itself into account and provides more detailed feedback than conventional general evaluations. 🔹 ARC-Bench highlights the problem of action ranking in frozen world models — The study tests the assumption that the proximity of a future state to the goal in latent space actually indicates the best action.

➡️ Discussions and case studies:

🔹 Habr users discussed whether Astra can be considered AGI-level — The author interviewed the model and examines what lies behind claims of achieving artificial general intelligence.

📝 If you would like to add other news and materials to the list, write in the comments.

News
More AI-tool write-ups on the Telegram channel — short and to the point
Subscribe

Comments

(0)
​