Research
Concepts
The ideas running through the desk, and what the papers are doing with each. Drawn from the papers themselves rather than from a taxonomy we invented.
agent memory
46 papersA repository where an agent stores past interactions and outcomes to inform future decisions.
verification
35 papersChecking outputs for correctness before accepting them, often to catch errors or reward hacking.
agent skills
31 papersLearned capabilities that enable AI agents to perform specialized tasks with improved effectiveness.
context compaction
31 papersCompressing or summarizing an LLM's past context into a smaller representation to extend memory and improve reasoning within limited token budgets.
coding agent
28 papersAn AI-powered software system that autonomously writes, modifies, or extends source code based on high-level instructions.
multi-agent architecture
28 papersA system where multiple AI agents collaborate, each handling specialized roles, communicating and coordinating to complete complex tasks.
agent harness
26 papersA system that structures, verifies, monitors, or controls how an AI agent interacts with external tools, environments, or execution runs.
agentic reinforcement learning
26 papersTraining approach that lets LLM-based agents learn multi-step tool use and planning via reinforcement learning, often with dense or per-tool rewards.
agentic evaluation
25 papersMeasuring AI agents' performance in complex tasks, highlighting failures in planning, monitoring, and real
vision-language model
25 papersA model that handles tasks combining visual and textual information, such as editing images from conversation or reasoning across images and text.
retrieval augmented generation
23 papersA method that retrieves relevant external information to ground and improve the output of a generative model.
agent safety
22 papersRisks from AI agents including timing errors, shutdown avoidance, collusion, overclaiming, and unsafe execution, mitigated by constrained decoding, safety harnesses, and swarm governance.
failure attribution
22 papersIdentifying which component or step in a system caused an observed failure, often using structured trace representations.
multi-agent coordination
19 papersThe collaborative operation of multiple AI agents to solve tasks more effectively or with greater capability.
harness evolution
18 papersIteratively adapting an agent's environment, prompts, or evaluation setup to improve task performance and generalization.
credit assignment
16 papersDetermining which actions or steps in a sequence contributed to a final outcome, often for reinforcement learning.
benchmark
14 papersA standard set of tasks or data used to evaluate and compare the performance of AI systems.
multi-turn tool use
13 papersThe capability of an agent to invoke external tools sequentially over multiple conversational turns to complete complex tasks.
typed state representation
11 papersTracking state with explicit data types, separated from
memory consolidation
11 papersSummarizing or structuring past information to reduce token usage while preserving task accuracy.
evidence blindness
10 papersHaving access to relevant evidence does not prevent an LLM agent from making incorrect or unsafe decisions.
agentic inference-time scaling
10 papersSpending more computation during inference, often by running multiple agent iterations or expanding context, to improve model output quality.
process-level evaluation
10 papersA granular assessment of a system's intermediate steps or subgoals, rather than just its final output, to diagnose specific failures.
agent-driven simulation
9 papersAutonomous agents automatically generate, control, or extend simulations to improve accuracy, interactivity, or performance.
recursive self-improvement
9 papersAn AI system repeatedly uses its own outputs to enhance its capabilities, creating a feedback loop for autonomous improvement.
model routing
9 papersDirecting each query or subtask to the most suitable model to optimize cost, accuracy, or quality.
deep research
8 papersA multi-step AI system that autonomously gathers, evaluates, and synthesizes evidence from diverse sources to produce comprehensive reports.
co-evolving feedback
8 papersFeedback that evolves alongside the policy or components it guides to improve performance.
group relative policy optimization
8 papersA reinforcement learning method that optimizes policy by computing advantages based on group-level reward statistics.
latent world model
8 papersA learned internal, compressed representation of environment dynamics, enabling future-aware planning, generation, and reasoning directly in latent space.
self-evolving agent
8 papersAn agent that improves by reusing its own experience, memory, or self-generated feedback, with success depending on verification and design.
ai control
8 paperscode as action
8 papersA paradigm where actions are expressed as executable code, such as method calls or program synthesis, rather than direct manipulation.
reinforcement learning with verifiable rewards
8 papersReinforcement learning that uses rewards from objectively checkable signals like correct answers or verification feedback.
vision-language-action model
7 papersA robot policy model that maps visual observations and language instructions to actions, trained on demonstrations and often used for manipulation tasks.
ui grounding
7 papersMapping natural language commands or descriptions to specific visual elements of a user interface.
world model for agents
7 papersA predictive environment model that agents use to simulate outcomes, enabling planning, reasoning, and long-horizon task execution.
embodied artificial intelligence
6 papersAI systems that interact with the physical world through sensors and actuators in a body or robot.
multi-agent collaboration
6 papersMultiple autonomous agents coordinate and combine their specialized capabilities to solve complex tasks beyond individual ability.
mixture of experts
6 papersA neural architecture with many specialized sub-networks, only a few activated per input via a router, balancing capacity and efficiency.
chain of thought
6 papersA step-by-step reasoning process that produces auditable trajectories for verification and accuracy.
llm as a judge
6 papersrole decoupling
6 papersstructured layout reasoning
6 papersA multi-step process that generates or verifies spatial arrangements of objects against explicit constraints and geometry.
preference optimization
6 papersA post-training method that adjusts model behavior using comparisons or self-generated preferences to improve alignment.
memory retention
6 papersThe ability of a system to preserve and recall information over time or across contexts, often balancing compression against fidelity.
agentic video generation
5 papersMulti-agent systems collaborate autonomously to produce long-form videos maintaining narrative and visual coherence.
graph-structured skills
5 papersA directed knowledge structure where skills are linked, enabling targeted updates and efficient retrieval for AI agents.
skill evolution
5 papersThe process of developing, refining, or optimizing reusable capabilities in LLM agents through learning, adaptation, or experience accumulation.
long-horizon robot manipulation
5 papersRobot manipulation tasks that require completing a sequence of many interdependent sub-tasks over an extended time horizon.
knowledge graph
4 papersA structured web of entities and relations, used to store domain facts and support reliable reasoning, retrieval, and model alignment.
amortized planning
4 papersA learned model generates plans directly for new instances, replacing expensive per-instance search with a fast, generalizable forward pass.
non-determinism
4 papersSame inputs give different outputs across runs, confounding testing and evaluation.
agentic scaffolding
4 papersExternal supports like training curricula, tool selection, and success verifiers enable agents to handle long-horizon or out-of-distribution tasks.
advantage estimation
4 papersIn policy optimization, a method that calculates the relative value of actions using group statistics or baselines to guide learning.
fine-grained spatial reasoning
3 papersReasoning precisely about object positions, orientations, and spatial relations across planar, depth, and temporal dimensions for robotics and navigation.
persistent corpus navigation
3 papersOrganizing a document corpus into pre-built, reusable navigation views that reduce online retrieval cost and improve accuracy.
unified multimodal understanding and generation
3 papersA single end-to-end model processes and generates multiple data types (e.g., images, text) with equal competence.
consequence-aware evaluation
3 papersEvaluation that assesses model performance by considering the real-world impact or risk of errors, not just standard accuracy.
agent-integrated software
3 papersSoftware whose behavior is driven by an embedded AI agent, enabling dynamic adaptation, autonomous task execution, and contract-based integration.
composable simulator infrastructure
3 papersA system for building simulators by combining and extending reusable components, often via automated agents.
kv cache compression
3 papers
An idea earns a page at 3 papers, and only when it is central to them. Categories stay categories.