295 papers this week in Agents12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Agents research

Latest Paper· Agents

Codetta protocol enables high-capacity, keyless, undetectable steganography for LLM agents

Qi Pang, Virginia Smith, and Wenting Zheng propose Codetta, a steganographic protocol that allows independently deployed LLM agents to communicate covertly without a pre-shared key. The protocol achieves high capacity while maintaining provable undetectability. On three agent workloads, Codetta transmits up to 94 times more bits per visible token than the prior state-of-the-art asymmetric protocol.

Qi Pang, Virginia Smith, Wenting Zheng
Reinforcement learningMETA ▲0.3%

Training LLMs as adaptive solvers for industrial-scale optimization

The authors propose Strategy-Diverse Reinforcement Learning (SDRL) to train open-source LLMs as adaptive meta-solvers for industrial-scale optimization. They show that combining solver-integrated reasoning, exact algorithms, and heuristic search with a hierarchical diversity reward allows the model to select the best strategy for each problem, outperforming DeepSeek-V4-Pro and GPT-5.5 on average across benchmarks and on industrial-scale tasks.

today
Evaluation & benchmarks

AI agents reproduce only 41% of machine learning papers with full code and weights provided

The authors introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers with predefined reproduction targets and GPU-hour budgets. They test four AI agents once per paper; the best agent reproduces 41% of Run-tier papers (code, data, weights provided), 27% of Retrain-tier (no weights), and 15% of Reimplement-tier (no code). Failed attempts use only 29% of their budget on average, and the most common error is implementing methods without verifying against the paper's reported numbers.

today
Evaluation & benchmarks

New benchmark measures whether AI systems can explore, not just recall, unfamiliar rules

The authors introduce ExplorationBench, a benchmark that evaluates AI systems' ability to explore unknown environments by forming hypotheses, designing experiments, and iterating on results. Its two executable sandboxes make every answer exactly checkable while ensuring that pretraining recall alone cannot solve the tasks. Evaluating ten AI systems, the authors find that the strongest can acquire and apply these unfamiliar rules, though performance varies across trajectories and continued exploration can stall or reverse earlier gains.

today
Reinforcement learning

Historical trajectories provide step-level credit for training LLM agents via regressor

Zhang et al. introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that extracts step-level supervision from historical successful trajectories to guide reinforcement learning of LLM agents. By pairing each state in a successful trajectory with its remaining number of steps, they train a frozen TabPFN regressor to predict progress in new rollouts; the change in predicted distance between consecutive states provides auxiliary credit for each action. Experiments on ALFWorld and WebShop with 1.5B and 7B models show that T2SPO consistently improves task success over GRPO, a baseline that only uses outcome rewards.

yesterday

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses