In short
Today’s topic is how MCP tools are starting to return not only data but also ready-made interfaces, while agents gain more autonomy in research and engineering.
Today: how MCP tools are starting to return not only data but also ready-made interfaces, while agents gain more autonomy in research and engineering.
🔥 Hot:
🔹 MCP tools have learned to return an interface directly to the agent — An analysis of OpenSearch MCP Apps shows how the result of a tool call can turn into an interactive screen instead of remaining text. 🔹 A reinforcement learning agent independently explores and changes complex systems — The Artificial Experimentalist operates in a closed loop: it chooses which states of cellular automata and other systems to explore and intervenes in the simulation during the experiment. 🔹 PICasso translates a natural-language description of a photonic circuit into a verifiable design — The framework combines YAML and GDS generation, knowledge of the manufacturing process, component routing, and DRC/LVS checks.
➡️ Useful materials:
🔹 A study examines how context window size affects the quality of literature reviews produced by LLMs — The authors compared reviews created by models in short and long contexts and evaluated the role of LLMs in preparing academic papers. 🔹 An LLM pipeline processed data from 536 papers on disease-spread models — The researchers are studying how capable language models are of extracting the necessary information for systematic reviews of agent-based modeling. 🔹 CIFQA combines multiple LLM agents with tools for accurate financial calculations — The framework is designed for questions involving rates, dates, formulas, and rules—areas where language models often produce plausible but incorrect numbers. 🔹 A review examines how large models are used to diagnose battery conditions — The paper covers battery prognostics and health management for electric vehicles, energy storage systems, and consumer electronics.
➡️ Discussions and case studies:
🔹 A LocalLLaMA user praises Ornith 1.5 for its speed and tool calling — In personal testing, the model achieved around 130 tokens per second with MTP and proved convenient for a fast tool-testing cycle. 🔹 A Habr analysis explains the story behind the “Cult of the Swarm” of AI models and the Hugging Face hack — The article recounts the results of an experiment in which a large number of models solved deliberately difficult tasks and generated unusual collective dynamics.
📝 If you would like to add other news and materials to the list, write in the comments.