Agents

LLM agents and the systems around them: harnesses and orchestration, tool use and MCP, multi-agent systems, agent memory and planning

299 articles · page 6 of 7

Agents

Layered agent harnesses improve software over repeated autonomous development cycles

Yan and colleagues developed Harness-of-Harness, a supervisory framework designed to help coding agents improve a software project across repeated development loops without human intervention. Across three benchmarks and three harness–model pairings, the authors report a 52.25% average relative gain over standalone harnesses, alongside a multi-day demonstration lasting more than 70 iterations.

3 Sept·2 min
Agents

Dual-brain memory keeps real-time speech agents accurate and emotionally aware

VoiceMem is a dual-stream memory architecture for duplex speech language models, pairing a factual left brain for personal information with an emotional right brain for affect and persona. The authors show that top-5 retrieval beats Mem0 at top-200 by nearly 30 points, and that retrieval completes in 134 ms, within voice-activity-detection latency. The result is a practical memory foundation for real-time, personalized, emotionally aware spoken interaction.

28 Aug·2 min
Agents

Separating state tracking from actions improves multi-turn tool reliability

Guo et al. introduce a structured tool-use policy that separates reconstructing task state, deciding whether to act, specifying an admissible action, and producing the final output. Across multi-turn, multi-tool, and incomplete-information evaluations, the authors report more reliable task completion than direct function calling and ReAct, especially for smaller models and tasks dependent on earlier context.

27 Aug·2 min
Agents

Selecting complementary LLM skills improves success while reducing context use

Chen et al. formulate the choice of reusable skill documents for LLM agents as a budget-constrained optimization problem rather than an independent relevance-ranking task. Their Best Prefix Selection algorithm achieved 0.73 task success on a contamination-controlled BigCodeBench variant, compared with 0.20–0.52 for existing selection methods, while using fewer tokens than the strongest released router.

23 Aug·2 min
Agents

New framework lets you train AI agents inside the same harness systems they use at inference

The authors present OpenForgeRL, an open-source framework that trains stateful, harness-based AI agents end-to-end using standard reinforcement learning libraries. It wraps the harness's model calls via a proxy for data recording and orchestrates rollouts in remote Kubernetes containers, yielding competitive results on tool-use, browser, and computer-use benchmarks with relatively few training tasks.

24 July·2 min
Agents

Python objects become AI agents in new NVIDIA framework

The authors introduce NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework where any Python object can act as an AI agent: methods define available actions, fields store state, docstrings serve as prompts, and type annotations act as contracts. Methods whose body consists only of '...' are completed at runtime by an LLM-driven loop, while normal methods remain deterministic. The framework is demonstrated on SWE-bench Verified, Terminal-Bench 2.0, and ARC-AGI-3, showing that current models use this interface effectively.

23 July·2 min
Agents

Re-evaluation shows harness evolution for agents may not outperform simple test-time scaling

Wang et al. revisit the evaluation of automatic harness evolution for LLM agents, comparing it with simple test-time scaling and discovery baselines under matched feedback and inference budgets. They find that harness evolution does not consistently outperform these baselines on Terminal-Bench 2.1 and shows limited generalization to held-out tasks, raising concerns about overfitting and the effectiveness of the approach.

14 July·2 min
Agents

Treating memory management as a trainable cognitive skill boosts LLM performance 2x–4x in long-horizon tasks

The authors introduce AutoMem, a framework that treats memory management as a learnable skill for LLMs by promoting file-system operations to first-class actions. By automating the optimization of memory structure (prompts, schemas, action vocabulary) and model proficiency (training on good memory decisions), AutoMem improved base agent performance by approximately 2x to 4x across three long-horizon games without altering task-action behavior.

2 July·2 min
Agents

Three-stage training instills world model planning into LLM agents

Zhang et al. propose a training paradigm that equips LLM agents with an internal world model for prospective reasoning. By first injecting predictive capabilities through a world model mid-training stage, then eliciting structured foresight via supervised fine-tuning, and finally calibrating with reinforcement learning, the authors demonstrate significant improvements on search and mathematical reasoning tasks over standard fine-tuning approaches.

26 June·2 min