Durable lines of agentic-AI research, tracked paper by paper. Each thread carries a living synthesis maintained by the research desk.
17 papers · 5 this fortnight
How to ensure language-model agents follow executable rules and avoid prohibited actions in regulated domains.
26 papers · 7 this fortnight
Methods for defending against and testing vulnerabilities in LLM agent systems, including prompt injection, memory poisoning, and gradual attacks.
19 papers · 7 this fortnight
Methods for training LLM agents to use internal world models for planning and reasoning over long horizons.
24 papers · 9 this fortnight
Benchmarks and methods for enabling LLM agents to coordinate effectively in multi-agent settings.
61 papers · 9 this fortnight
Techniques for optimizing the code, instructions, and training frameworks surrounding LLM agents to improve performance and reduce cost without changing model weights.
23 papers · 7 this fortnight
Techniques for constructing, pruning, and augmenting context windows to improve LLM agent performance and reliability.
4 papers · 3 this fortnight
Methods and benchmarks for LLM agents that modify software behavior or code at runtime in response to observed failures or user needs.
37 papers · 8 this fortnight
Methods for managing memory in LLM agents, including consolidation, compression, and learnable memory skills under budget constraints.
14 papers · 4 this fortnight
Optimizing the selection of reusable skill documents for LLM agents under budget constraints to improve success and reduce context use.
2 papers · 2 this fortnight
Methods that use reinforcement learning to train language-model and multimodal agents to effectively call external tools and APIs.
10 papers · 6 this fortnight
Benchmarks and methods for evaluating and enabling language-model agents to learn from their own past experiences and capability goals without external supervision.
1 paper · 1 this fortnight
Benchmarks that decompose agent tasks into subgoal processes for fine-grained evaluation of agent performance and error analysis.
2 papers · 2 this fortnight
Exploring how diffusion-based generation (iterative masking/filling) can be adapted from pretrained autoregressive models, and how this affects efficiency and generation capabilities across parameter scales.
3 papers · 2 this fortnight
Benchmarks and evaluations for LLM agents operating across multiple devices or operating systems, focusing on workflow continuity and cross-platform coordination.
17 papers · 4 this fortnight
Methods for distributing credit over long trajectories in language-model agent reinforcement learning when only sparse terminal rewards are available.
2 papers · 2 this fortnight
Investigating how safety training can inadvertently transform harmful behaviors (e.g., discrimination) into subtler forms rather than eliminating them, and how to detect or mitigate these side effects.