Agents

LLM agents and the systems around them: harnesses and orchestration, tool use and MCP, multi-agent systems, agent memory and planning

299 articles · page 4 of 7

Agents

Scientific agents gain auditability through enforced, re-verifiable research workflows

The IdeaHorizon Team describes AfS, a platform for scientific projects spanning dozens of agent runs with only occasional human involvement. Rather than relying on model instructions alone, AfS mechanically enforces commitments, evidence handling, review boundaries, and memory rules. The paper illustrates these mechanisms through failure traces and two completed campaigns, but does not provide a comparative benchmark.

29 Sept 2026
Agents

Preference learning improves collaboration among teams of language-model agents

Liu et al. formulate preference learning for multi-agent language-model systems and introduce MAPL, which iteratively improves agents by comparing their joint solutions with alternatives from comparator models. Across writing, coding, tool-use and travel-planning tasks, MAPL improved collaboration quality and efficiency, with its learned-reward variant generally outperforming direct preference optimization.

29 Sept 2026
Agents

Multi-agent LLM scaling benefits depend on task type and aggregation method

Fortuna and Bertalanič introduce Steiner's taxonomy of group tasks to analyze multi-agent LLM scaling, modeling independently sampled agents as conditionally independent given an item. Testing 13 open-weight models on teams of up to 30 agents across disjunctive and compensatory benchmarks, they find that on disjunctive tasks, plurality voting captures almost none of the potential from larger teams (model predicts within 0.5 points of true accuracy), while multi-round revision yields strong gains that saturate with just one peer. On compensatory Fermi estimation, averaging reduces error by only about 6% because item-level biases shared across samples account for about 87% of squared error.

29 Sept 2026
Agents

Compiling agent skills into state machines improves reliable task execution

Minghao Li introduces a compiler that converts reusable agent instructions and tool interfaces into extended finite state machines, which track progress, intermediate results, and permitted next operations. Across four benchmarks and four executors, HEXIS improved average task success by 16.1 percentage points over a Skill + ReAct baseline while substantially reducing execution tokens for one tested model.

26 Sept 2026
Agents

Editing agent reasoning history boosts long-horizon task performance

The authors propose AEWM, a world model that edits agent reasoning and action histories to remove noisy or outdated state, rather than predicting environment observations. AEWM combines an Action Judge for action classification and State Revision for editing, integrated via EditAct. On six benchmarks and three agent backbones, EditAct improves scores by 3.2–6.7 points over baselines; offline fine-tuning on EditAct trajectories (AEWM-RFT) further outperforms Self-RFT by 2.2–2.6 points without online guidance.

25 Sept 2026
Agents

Reinforcement learning trains video AI agents to use external tools effectively

The authors present VideoGen-Agent, a multimodal agent trained via multitask agentic reinforcement learning to orchestrate external tools for video generation. On the introduced VABench benchmark, the agent improves its base text-to-video generator from 56.5 to 75.6, and upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons.

23 Sept 2026