agentic inference-time scaling
Spending more computation during inference, often by running multiple agent iterations or expanding context, to improve model output quality.
- Papers
- 10
- Released code
- 4
- First seen
- June 2026
- Latest
- Sept 2026
3 papers in the last two months, against 7 in the two before.
The papers
Most central to this idea first, not most recent.
- Independentcs.CVcode
Open-source image model rivals closed systems on minimal budget
July 2026
- Big Techcs.CV
Agent framework auto-tunes video diffusion for 2x speedup
NVIDIA Research · June 2026
- Top Universitycs.MA
Decentralized agents with shared context outperform centralized orchestration
Stanford University · June 2026
- Big Techcs.CL
Agentic document QA pays off mainly with stronger vision-language models
Amazon.com · Sept 2026
- Top Universitycs.DC
DynBranch starts and reuses agent work before branches resolve
National University of Singapore · Sept 2026
- Top Universitycs.AIcode
Dropping Rather Than Rewriting Context Cuts Long-Horizon Agent Costs
Carnegie Mellon University, Bosch Center for AI · Sept 2026
- Top Universitycs.AIcode
Re-evaluation shows harness evolution for agents may not outperform simple test-time scaling
Allen Institute for AI, University of Washington · July 2026
- AI Startupcs.LG
Orchestrator models dynamically combine specialized LLMs into collective intelligence
Sakana AI · June 2026
- Big Techcs.CLcode
Letting language models decide when to compact their context windows
Johns Hopkins University, Apple · June 2026
- Big Techcs.AI
Memory as state management instead of semantic retrieval improves long-horizon agents
University of Science and Technology of China, Microsoft · June 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.