Agents

LLM agents and the systems around them: harnesses and orchestration, tool use and MCP, multi-agent systems, agent memory and planning

299 articles · page 5 of 7

Agents

Trajectory shortcut trees improve agents without outcome labels or annotations

Liu et al. developed DENSE, a method that extracts completed work, recovery evidence, and unresolved requirements from agent trajectories without knowing whether the original attempt ultimately succeeded. On Terminal-Bench 2.1, reruns guided by this distilled evidence improved strict pass rates across four recipient models by 7.12–15.64 percentage points while reducing their observed token use by 19.0–43.6%.

22 Sept·2 min
Agents

Obstacle-aware harness improves safety of coding agents for robot manipulation

Bingxin Xu and colleagues evaluate whether coding agents—large language models that write robot controllers as programs—can respect a physical safety constraint, pairing each manipulation goal with an obstacle the robot must not touch. They find the agents prioritize task completion and collide in most cases, but their SafeHarness harness improves collision avoidance from 60.5% to 87.5% while raising task success to 71.9%.

20 Sept·2 min
Agents

Cross-task failure diagnosis makes LLM agent harness training faster

The authors developed Ecdysis, a framework for improving the runtime code and instructions that govern an LLM agent’s execution. Rather than revising a harness after each failed task, it analyzes batches of failures for recurring patterns and uses multiple diagnostic roles to propose refinements. In the reported experiments, this reduced training time by up to 1.84× while increasing reasoning accuracy by 18.56% over existing harness-evolution methods.

12 Sept·2 min
Agents

Personalizing agents through cross-session interaction data boosts task success

Wang et al. propose TAHI (test-time adaptation through human-agent interaction), which leverages a user's cross-session interaction history to update agent context and weights while maintaining an evolving rubric that captures personal evaluation criteria. Applied to writing and visual creation across 30 individuals, the adapted agents outperformed non-adapted baselines on individual tasks, and part of the personalization gains transferred to other users.

7 Sept·1 min
Agents

Graph-based policy constrains LLM agents for topology-aware incident response

Vallabhaneni, Cagwin, and Wild built a security-operations architecture in which a graph encoder and reinforcement-learning policy analyze enterprise authentication activity instead of placing the full network state into an LLM’s context. On public cyber-security data and an Indiana University computing cluster, the system achieved 0.91 precision and 0.87 recall on labeled red-team events, with a median 6.3-second end-to-end response cycle.

6 Sept·2 min