verification
Checking outputs for correctness before accepting them, often to catch errors or reward hacking.
- Papers
- 35
- Released code
- 4
- First seen
- June 2026
- Latest
- Sept 2026
33 papers in the last two months, against 2 in the two before.
Who is working on it
The papers
Most central to this idea first, not most recent.
- Big Techcs.LG
Exactly-once behavior shifts from models to contracts with fault type
Microsoft · Sept 2026
- Chinese Techcs.CL
Verified workflow training gives language models reusable procedural skills
East China Normal University, Shanghai AI Laboratory · Sept 2026
- Chinese Techcs.CL
Hybrid agents learn to recreate software across five computing platforms
Alibaba Token Hub, Alibaba Group · Sept 2026
- Industrycs.AI
Formal logic solvers can audit and improve language-model reasoning chains
University of Bristol, King Abdullah University of Science and Technology · Sept 2026
- Independentcs.SE
Engineering agents need evidence-bound authorization before their outputs trigger action
Sept 2026
- Big Techcs.AI
Probabilistic scoring unlocks verification as a new scaling axis for LLMs
Stanford University, UC Berkeley · July 2026
- Top Universitycs.AI
Task-specific plans and terminal checks create different value for LLM agents
The Chinese University of Hong Kong, University of Edinburgh · Sept 2026
- Industrycs.DCcode
Dispatch-aware agents safely optimize compiler-generated GPU kernels within full models
Red Hat · Sept 2026
- Industrycs.AI
Safe agent self-modification requires verifiable, expressive recovery mechanisms
Independent Researcher · Sept 2026
- Research Labcs.CR
Execution Logs Can Mislead Visual Judges in Video-Generation Agents
RIKEN · Sept 2026
- Top Universitycs.AI
LLM agents collude to bypass verification in long-horizon tasks
Stanford University, Georgia Tech · Sept 2026
- Industrycs.CRcode
Cyber agents hide critical weaknesses behind successful attack workflows
National Research Council, University of Windsor · Sept 2026
- Industrycs.AI
Co-trained agent roles improve tool-based reasoning and verification
Waseda University, Adelaide University · Sept 2026
- Independentcs.AI
Deterministic guardrails expose gains that fool LLM-based judges
Sept 2026
- Top Universitycs.LG
Sparse verifier feedback makes broad credit assignment outperform turn targeting
Institute of Science Tokyo, Zhejiang University · Sept 2026
- Industrycs.AIcode
Memory lets traffic-simulation agents build reusable skills without retraining
Jilin University · Sept 2026
- Independentcs.AI
Persistent evidence tracking improves language-agent performance on bioinformatics benchmarks
Sept 2026
- Industrycs.CL
Closed-loop synthesis efficiently trains language models to use external tools
vivo AI Lab · Sept 2026
- Industrycs.CL
Self-evolving loop synthesizes high-quality multimodal training data
vivo AI Lab · Aug 2026
- Independentcs.AI
Benchmark tests whether coding agents improve governance without breaking repositories
Sept 2026
- AI Startupcs.AIcode
Goal pressure exposes wide gaps in security agents’ scope adherence
dreadnode · Sept 2026
- Industrycs.AI
Research agent improves itself through seven successive code rewrites
Weco AI · Sept 2026
- Big Techcs.AI
AI coding agents still miss production inference failures
NVIDIA, University of California, Berkeley · Sept 2026
- Big Techcs.RO
One demonstration can bootstrap reusable robot manipulation programs
Massachusetts Institute of Technology, National University of Singapore · Sept 2026
- Top Universitycs.AI
Attribution-guided skill graphs improve targeted repairs for frozen language models
National Key Laboratory for Novel Software Technology, Nanjing University · Sept 2026
- Big Techcs.CL
Process-based evaluation reveals where computer-use agents go wrong
NVIDIA · Sept 2026
- Top Universitycs.CV
Reinforcement learning trains video AI agents to use external tools effectively
Princeton University, Stanford University · Sept 2026
- Independentcs.SE
Multi-image evidence can help models repair software, but unreliably
Sept 2026
- Industrycs.CL
DolphinBench measures agent memory through tasks, cost, and latency
Mem0 · Sept 2026
- Independentcs.CR
Hierarchical registries enable fast, trust-aware discovery across agent ecosystems
Sept 2026
- Industrycs.AI
Distilled repository skills improve agents conducting machine-learning research
Beijing Academy of Artificial Intelligence, University of Science and Technology of China · Sept 2026
- Industrycs.AI
Layered agent harnesses improve software over repeated autonomous development cycles
Shanghai Artificial Intelligence Laboratory · Sept 2026
- Independentcs.AI
Query-aware evidence forests improve long-term memory for multimodal agents
Aug 2026
- Independentcs.AI
Five-level ladder maps reasoning models beyond direct human oversight
Sept 2026
- Top Universitycs.MA
Decentralized agents with shared context outperform centralized orchestration
Stanford University · June 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.