Agents

LLM agents and the systems around them: harnesses and orchestration, tool use and MCP, multi-agent systems, agent memory and planning

299 articles · page 3 of 7

Agents

Hierarchical supervision allocation improves long-horizon model distillation

The authors propose LENS-OPD, a framework that distills a large teacher model into a smaller student agent for long-horizon tasks by hierarchically allocating supervision: first deciding how long the student should explore, then which decision to intervene on, and finally which tokens to correct. On agent benchmarks including WebShop and ScienceWorld, LENS-OPD consistently improves task success rates over vanilla on-policy distillation and strong baselines.

30 Sept 2026
Agents

On-device caches convert natural language to actions with classification not generation

Fereidouni et al. introduce on-device semantic operation caches that transform the natural language to action problem from a generative task into a classification-centric formulation. Tested on generating Excel formulas from natural language requests, the system reduces total inference cost by 56 percent compared to cloud-only model routing and improves response latency by five times on cache hits.

30 Sept 2026
Agents

Interface support helps developers supervise more coding agents in parallel

The authors studied how developers coordinate multiple autonomous coding sessions, then distilled five supervisory practices into the PILOT framework: Planning, Isolating, Logging, Observing, and Triaging. In a 16-person experiment, a prototype implementing these practices increased ticket throughput by 63 percent and let participants supervise one additional agent at peak, although it did not significantly improve perceived control.

30 Sept 2026
Agents

Event graph corrections improve LLM agent stock forecasts by modeling historical reliability

The authors present RICE-Alpha, a point-in-time stock-scoring framework that combines a history-aware multi-view Base Alpha with a reliability-weighted residual correction from event graph transitions. Tested on daily Nasdaq-100 and Hang Seng Index panels from 2024 to 2026, it achieves the strongest results among evaluated LLM-based agents and momentum baselines, with an ICIR more than doubling that of the strongest baseline and net Sharpe ratios of 1.656 (US) and 1.725 (Hong Kong).

30 Sept 2026
Agents

Terminal agents struggle when familiar environmental assumptions stop holding

Singh and colleagues introduce AGNI, a pipeline that identifies assumptions behind successful agent trajectories, invalidates them through controlled environmental changes, and verifies that the modified tasks remain solvable. Multiple LLM agents performed substantially worse on these novel variants, often observing evidence of a change without correctly diagnosing it. Post-training on environmental novelty improved performance on both held-out novel tasks and unchanged base tasks.

30 Sept 2026
Agents

Draft-KV ensures the receiver actually uses the sender's latent message

The authors identify a fundamental flaw in prior latent communication methods: replacing the sent message with one from an unrelated question barely changes accuracy, meaning the receiver relies on the interface, not the message content. They propose Draft-KV, which transmits key-value states from the sharer's own answer draft via a gated attention branch with progressive training. A frozen 8B sharer lifts a frozen 0.5B receiver from 37.45% to 78.04% on MMLU-Redux, and message reassignment drops accuracy to 36.40%, confirming the gain stems from genuine latent information exchange.

30 Sept 2026
Agents

Agent distillation improves when supervision follows outcomes, not model disagreement

Zhang et al. test a common assumption in on-policy distillation: that large token-level differences between teacher and student mark the most useful moments for supervision. They show that these local gaps often mispredict downstream benefit, then introduce OG-OPD, which weights teacher guidance using final outcomes from paired student continuations and improves success across three interactive benchmarks.

30 Sept 2026
Agents

Production-scale AutoResearch exposes five failure modes and a three-principle fix

The authors applied AutoResearch, an LLM that iteratively edits training scripts and keeps changes that improve a held-out metric, to two production embedding systems at Amazon. Over 220 experiments spanning 12 weeks, they identified five recurring failure modes and a three-principle multi-agent scaffolding design that maps each to a structural remedy. The framework improved Recall@6 by 1.82x, coherence by 2.1x, and expanded catalog coverage 5.8x.

30 Sept 2026
Agents

A method that links failure localization to repair by identifying which editable system asset to change

The authors propose DAAF, a framework that moves beyond localizing where a failure occurs in an LLM agent to determining which persistent, editable system asset (e.g., a routing rule or knowledge segment) should be changed. By learning the effects of valid attribute replacements from controlled replays and amortizing this evidence at deployment time, DAAF achieves 80.72% attribute Hit@1 on held-out Telecom tasks while recovering 62.65% of failed executions and limiting clean-task regression to 3.23%.

29 Sept 2026