The authors converted MMLU, MMLU-Pro, MedMCQA, MATH, GPQA and HumanEval into 31,119 dynamic episodes.
Revision-aware agent graphs reuse valid work while avoiding stale answers
RIAG combines deterministic version selection, version-keyed caching and bounded independent reasoning to handle tasks revised over time.
Top University
Yan Luo · Selim-Antoine Lali · Jeremy Moebel · Iliass Khoutaibi · Ahmadou Aidara · Mengyu Wang
Harvard AI and Robotics Lab · Harvard University
Research Digest··2 min read
Luo and colleagues study dynamic task routing, where an agent must identify which version of a document applied at a query time and then answer the associated problem.
Why this paper
From Harvard AI and Robotics Lab and Harvard University
In one line
RIAG achieves 54.24% accuracy at 0.62 calls/query, outperforming baselines that use 18 calls/query.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§