In short
Today’s focus: weak spots in the memory and security of AI agents, speeding up work with MCP tools, and new discoveries for running models locally.
Today: weak points in the memory and security of AI agents, speeding up work with MCP tools, and new findings for running models locally.
🔥 Hot:
🔹 Study shows how semantic masking bypasses superficial LLM defenses — The authors link the problem to the fact that harmful knowledge remains inside the model even when the final answer is blocked. 🔹 Study reveals a new memory failure in AI agents: necessary data may disappear before retrieval — The problem arises when the system removes indirectly related fragments needed to answer a future query. 🔹 Nexus speeds up launching LLM agents with a large set of MCP tools — The system separates tool routing from repeatedly processing their detailed schemas, which increases time to first token.
➡️ Useful materials:
🔹 PrimeAgentOrchestrator launches Claude Code with memory of past tasks — The system selects relevant user memories and loads them into a new coding-agent session. 🔹 Study shows that how skills are represented affects multimodal agent routing — The paper focuses on selecting the right skill from a growing library of tools and skill modules. 🔹 StateSight proposes separately measuring how vision-language models recover the spatial structure of an image — The benchmark separates spatial understanding from OCR, general knowledge, and other factors mixed together in broad tests. 🔹 New review compiles methods and applications of multimodal agent frameworks — The focus is on systems that combine perception, memory, and decision-making based on large multimodal models. 🔹 Review examines the Russian market for AI-powered analytics solutions in 2026 — The market’s main demands are a clear industry impact, accessible deployment, and a short return-on-investment horizon.
➡️ Discussions and case studies:
🔹 A user ran deepseek-v4-flash-0731 locally on an RTX 5090 and achieved around 24 tokens per second — The experiment demonstrates a practical scenario for running a model with approximately 151 GB of weights and a large amount of RAM. 🔹 A developer trained a 250M-parameter LLM on 30 billion tokens and compressed it to 60 MB — The model was trained from scratch and uses quantization to less than two bits.
📝 If you would like to add other news and materials to the list, write in the comments.