Agents

LLM agents and the systems around them: harnesses and orchestration, tool use and MCP, multi-agent systems, agent memory and planning

299 articles · page 7 of 7

Agents

Latent communication across different LLMs matches or beats text-based transfer at lower compute.

The authors investigate whether heterogeneous LLMs can be aligned to directly transfer latent representations (KV-caches) between them, bypassing costly text decoding and re-encoding. They propose a method using a cross-model transformation and two-phase training (reconstruction then generation), achieving performance comparable to or better than text communication in context-aware settings at 2-3x lower compute, and remaining effective in context-unaware settings where prior heterogeneous methods collapsed.

12 June·2 min
Agents

Decentralized agents with shared context outperform centralized orchestration

The authors propose Decentralized Language Models (DeLM), a multi-agent framework that replaces a central controller with a shared context and task queue, allowing agents to asynchronously claim subtasks and build on verified progress. On SWE-bench Verified, DeLM improved performance by up to 10.5 percentage points and reduced costs by roughly 50%. On LongBench-v2, it achieved the highest average accuracy across four frontier model families.

9 June·2 min
Agents

Pruning context to recent tool calls improves agent reliability and efficiency

The authors test four context engineering strategies for GPT-5 agents using Model Context Protocol tools on a 50-task hotel expense benchmark. They find that pruning context to the last five tool call/response pairs and adding automated summarization yields the best results: 91.6% complete itemization, 99.64% amount itemized, with 63% fewer tokens and 60% less runtime than retaining full conversation history.

9 June·2 min
Agents

VLMs improve spatial reasoning by actively imagining novel views with a world simulator

The authors propose Astra, an agentic spatial reasoning framework that combines two components: a world simulator (Astra-WM) that generates novel-view images from context images and natural-language camera motions, and an RL-trained VLM policy (Astra-VL) that decides when to query the simulator. Trained with view consistency tuning and a two-phase RL curriculum, Astra improves spatial reasoning benchmarks by 9–10 points over direct VLM answering, demonstrating that effective world-model-augmented reasoning requires learning when, where, and how to imagine.

5 June·2 min
Agents

Self-supervised method improves agent harnesses using past trajectories

The authors introduce Retrospective Harness Optimization (RHO), a self-supervised method that optimizes an AI agent's harness of skills, tools, and workflows using only past task trajectories. By selecting a diverse coreset of challenging tasks, re-solving them in parallel, and using self-validation to pick the best harness update, the method improves pass rate on SWE-Bench Pro from 59% to 78% in a single optimization round.

4 June·3 min
Agents

Proposing a workflow store to harden AI agents against failure

The authors critique the dominant on-the-fly paradigm for AI agents, which synthesizes plans and executes actions rapidly in response to prompts, arguing it bypasses established software engineering processes like testing and adversarial evaluation. They propose an AI Workflow Store of hardened, reusable workflows to amortize the cost of rigor across users, aiming to improve reliability and security for high-stakes applications.

12 May·3 min