Agents

LLM agents and the systems around them: harnesses and orchestration, tool use and MCP, multi-agent systems, agent memory and planning

299 articles · page 2 of 7

Agents

Contrasting multiple trajectories sharpens skill learning for LLMs

Li et al. propose SkillCome, a method that improves large language models by iteratively refining textual skills. For each question, the system generates multiple trajectories, contrasts successful and failed ones to pinpoint effective actions, and accumulates evidence across iterations in a dual memory system. On six benchmarks spanning reasoning and agentic tasks, SkillCome consistently outperformed prior skill evolution baselines, with gains up to +5.69 points.

30 Sept 2026
Agents

StepLearn lets agents use new knowledge immediately while requiring validation before persistent reuse.

The authors introduce StepLearn, a nonparametric framework for test-time learning in LLM agents. Instead of learning only from completed episodes, StepLearn distills individual transitions into hypotheses that can guide the next decision, while deferring persistent reuse until those hypotheses are prospectively validated in later episodes. Across five rounds on WebArena-Lite and ALFWorld, StepLearn outperforms the strongest baseline EvoTest by 2.2 to 12.7 percentage points.

30 Sept 2026
Agents

Statistical gates help self-editing agents avoid lasting performance regressions

Wang et al. examine how agents decide whether to retain proposed edits to persistent instructions that encode workflows, tool use and decision rules. Their SAGE gate compares old and edited skills on the same validation items, penalizes regressions, and abstains when evidence of improvement is statistically weak. Across five benchmarks and four language models, it reduced regressions in 19 of 20 settings while producing the highest final score in every setting.

30 Sept 2026
Agents

Specialized executors follow declared agent plans more faithfully

Oota et al. separate agent planning into two problems: choosing an appropriate planning mode and executing that mode faithfully. Across four benchmarks and three language models, generic Plan+ReAct agents frequently abandoned their declared structure, while pattern-specific executors largely closed this execution gap. The remaining weakness was selection, since models did not reliably choose the best mode for each task.

30 Sept 2026
Agents

Counterfactual credit teaches language models which memories stay useful

Tang, Liu and Sarabi address a central problem in persistent language-model memory: a rewrite may help only much later, and its apparent value may actually come from information already stored. Their Memory Gain Policy Optimization method compares updated and prior memory states to isolate the value created by each rewrite, improving document-level information extraction while cutting average memory length by nearly 80% from the initial policy.

30 Sept 2026
Agents

Retrieving existing agent skills improves optimization across different execution harnesses

Chu et al. introduce Retrieval-Augmented Skill Optimization, or RASO, for improving the natural-language instructions that guide agents inside particular tool and execution environments. Across four benchmarks and two language models, the authors report that retrieving and adapting existing skills consistently outperformed optimization baselines that lacked retrieval during initialization and updating.

30 Sept 2026
Agents

Adaptively choosing how to revise LLM skills outperforms fixed revision strategies

The authors propose StraTune, a method that adaptively selects which revision operator to apply when evolving textual skills for LLMs. By maintaining an optimization state of past operator outcomes and letting a frozen optimizer LLM choose the operator each round, StraTune outperforms five baselines on four benchmarks. Gains are attributed to the adaptive selection, as fixed, random, scheduled, and bandit strategies all score lower.

30 Sept 2026
Agents

Shared scheduler enables realistic training and testing of continual-learning agents

Jung et al. built an execution substrate for evaluating and post-training agents across sessions, background jobs, memory consolidation and simulated weeks of activity. Across seven ported benchmarks, they found that external memory did not consistently improve on harness-native memory, but post-training a 4-billion-parameter model to use its harness and memory produced substantial efficiency and accuracy gains.

30 Sept 2026
Agents

LAM Cuts Agent Memory While Bounding Retrieval Score Errors

The authors introduce LAM, a memory framework that deduplicates agent observations without using an LLM summarizer, preserves cached context prefixes, and estimates compaction costs before deployment. It removed 22.47% of observation tokens while retaining 99.984% of measured evidence associated with correct patches, although its mathematical bound covers retrieval-score changes rather than guaranteeing identical retrieval rankings.

30 Sept 2026