In short
Today's roundup covers new risks posed by autonomous AI agents, ways to track their errors, and the latest discussions about API privacy.
In today’s roundup: new risks posed by autonomous AI agents, ways to track their errors, and the latest discussions about API privacy.
🔥 Hot:
🔹 OpenAI noticed warning signs before the AI-agent attack on Hugging Face — The story shows why autonomous agents need restrictions and monitoring before, not after, an incident. 🔹 Researchers proposed reconstructing agent behavior from traces of their actions — A model of the overall structure of runs helps predict the next step and possible failures instead of analyzing each trace separately.
➡️ News:
🔹 GLM-5.3 will release its weights publicly — The LocalLLaMA community reported that the release promise has been fulfilled.
➡️ Useful materials:
🔹 Gated Activation Steering reduces sycophancy and hallucinations in medical responses — The paper focuses on steering model activations so that responses stay better grounded in context and are less susceptible to user pressure. 🔹 FLARE proposes evaluating not only AI accuracy in medicine, but also the benefits of deployment — The framework accounts for the financial and operational consequences of using models in real-world clinical workflows. 🔹 The study outlines rules for the responsible delegation of scientific work to LLMs — The authors examine what can be delegated to a model, how to verify the result, and which elements of reasoning should remain under human control.
➡️ Discussions and cases:
🔹 Hacker News users are discussing APIs for open-weight models with guaranteed zero prompt retention — The idea assumes no training on prompts and responses, while retaining only technical data needed for operation and billing. 🔹 Developers are discussing how to protect self-hosted sites from aggressive data collection by LLM scrapers — The discussion considers AI firewalls, traps in HTML, and other ways to limit automated access.
📝 If you would like to add other news and materials to the list, write in the comments.