In short
Today: why agents need real limits on their actions rather than just good prompts, and how reasoning is becoming part of the API contract. Plus practical finds from the community and some unusual applications of AI.
Today: why agents need real limits on their actions rather than just good prompts, and how reasoning is becoming part of the API contract. Plus practical finds from the community and some unusual applications of AI.
🔥 Hot:
🔹 Researchers have proposed Aegis for safely controlling AI agents' actions — The system puts the boundary at the execution level: it checks an action's provenance and, in case of doubt, refuses outright, not letting the agent change files, send messages or launch tasks. 🔹 A study calls reasoning effort part of a model's API contract — When buying an AI service what matters is not only the model's brand but the chosen level of reasoning, the model actually served, the output format and the tariff. 🔹 GxP-Agent turned clinical trial software development into a managed process graph — The authors note that across 11 attempts five frontier models failed to produce a correct subject-level dataset, and propose using a Process-DAG instead of generating code in a single request.
➡️ News:
🔹 PlugClaw has unveiled a private device for AI agents on phone and PC — The Show HN project promises to run a personal agent through a separate piece of hardware.
➡️ Useful reading:
🔹 The Model Hypnosis study explores steering AI through additive hidden effects — The work is about a way to change models' behaviour substantially without the usual textual input. 🔹 Block has described how its team designed AI with character for Berd — This is a practical account of the decisions behind building an AI product with a pronounced personality.
➡️ Discussions and cases:
🔹 Qwen3.8-27B has been pushed to 218 tokens per second on two RTX 3090s — The author used vLLM and DFlash2; the test also reached a context of up to 131K tokens with peak usage of 22.3 GB of memory per card. 🔹 AI is analysing 80 million pages of archives to hunt for sunken treasure — The system studies 500 years of Spanish colonial records and helps find information about sunken ships and lost cargo. 🔹 A Hacker News user asked why AI credits cannot be moved between applications — The occasion was a case where credits were available in an IDE but not in another agent used for deployment.
📝 If you would like to add other news and materials to the list, write in the comments.