Agents

LLM agents and the systems around them: harnesses and orchestration, tool use and MCP, multi-agent systems, agent memory and planning

299 articles · page 1 of 7

Agents

Codetta protocol enables high-capacity, keyless, undetectable steganography for LLM agents

Qi Pang, Virginia Smith, and Wenting Zheng propose Codetta, a steganographic protocol that allows independently deployed LLM agents to communicate covertly without a pre-shared key. The protocol achieves high capacity while maintaining provable undetectability. On three agent workloads, Codetta transmits up to 94 times more bits per visible token than the prior state-of-the-art asymmetric protocol.

3 Oct 2026
Agents

Evidence-driven harness edits improve frozen GUI agents across benchmarks

Yang et al. introduce an automated method for improving the runtime harness around a GUI agent without changing the underlying model. Across six backbones on OSWorld-Verified, the authors report consistent held-out improvements, including a 12.33-point gain for Qwen3-VL-32B-Instruct. A harness optimized on one benchmark also transferred to WindowsAgentArena, where it improved GPT-5 by 13.87 percentage points.

2 Oct 2026
Agents

Visual memory harness helps multimodal agents solve long interactive tasks

Han et al. introduce VISTA, a model-agnostic harness designed to extend multimodal agents beyond their immediate visual context. Rather than compressing each observation once and relying on that representation throughout a task, VISTA stores past images in their original form and lets the model actively retrieve and arrange them as reasoning proceeds. With Claude Opus 5.0, this approach completed all 25 public ARC-AGI-3 games and achieved the benchmark’s maximum Relative Human Action Efficiency score.

2 Oct 2026
Agents

Reusable reasoning choices make multimodal search agents substantially faster

Zhu et al. introduce Selection-based Structured Reasoning (SSR), which asks a multimodal search agent to choose among reusable natural-language reasoning templates rather than generate a new rationale token by token before every action. Across seven benchmarks with 2B and 4B parameter models, the authors report performance competitive with similarly sized search agents, alongside more than 90% lower per-turn reasoning latency and 28% to 54% lower total model inference latency per question.

2 Oct 2026
Agents

Models Learn to Allocate Context and Scale Inference More Effectively

Li and colleagues study contextual reasoning, the ability to allocate multiple context windows and decide what information should pass between them during inference. Their Hermes harness gives models varying degrees of control over those decisions, while Hermes-Learn uses supervised fine-tuning followed by reinforcement learning to teach smaller models how to use that control. The authors report that training produces adaptive strategies and performance gains that persist across benchmarks, models, and inference budgets.

1 Oct 2026
Agents

Reusable harness components let agents adapt workflows to each task

Kuang and colleagues argue that a single agent harness is poorly suited to heterogeneous tasks because mechanisms such as verification, context management and termination logic can help in one setting but distract or slow an agent in another. Their STITCH framework composes task-specific harnesses from reusable primitives, improving reported task success by up to 12 percentage points over fixed harnesses while adding 2.7% test-time overhead.

1 Oct 2026
Agents

Selective self-distillation helps LLM judges generalize on subjective evaluations

Hong et al. train LLM judges using preference rationales and rubrics rather than relying only on rewards for correct final verdicts. Their position-selective self-distillation method withholds supervision at tokens where feedback sharply narrows the model toward one expression, improving generalization and outperforming outcome-supervised reinforcement learning by 2 to 9 percentage points on evaluated subjective categories.

1 Oct 2026
Agents

Cross-checking coding agents improves reliability in automated model research

Chen and colleagues present RankEvolve, a harness that automates repeated proposal, implementation, training and evaluation of ranking-model changes while enforcing a predefined experimental process. Their central result is that combining Claude Code and Codex as cross-checking execution nodes raised fully correct execution from 45.8 percent to 62.5 percent, although consequential defects still passed silently in 10.4 percent of cases.

1 Oct 2026