The authors first demonstrate that retrieval and LLM policy learning are order-sensitive: adapting the retriever before the policy yields larger reward gains than the reverse order.
Bilevel optimization bridges the retriever-policy credit gap in agentic RL
BRIDGE jointly adapts retrieval and LLM policy in a hierarchical order, improving multi-hop QA accuracy by up to 9.6 EM points.
Academic
Quan Xiao · Mingda Liu · Gaowen Liu · Katsuki Fujisawa · Tianyi Chen
Cornell University · Institute of Science Tokyo · Cisco Research
Research Digest··3 min read
Thread:RL for Tool Agents
The authors identify an information-credit gap in agentic reinforcement learning (ARL) for LLMs: when failures arise from poor retrieval, the fixed retriever escapes blame while the LLM policy is penalized.
Why this paper
From Cornell University and 2 others · Part of RL for Tool Agents, now 14 papers
In one line
Training retrieval before policy optimization, then co-adapting both with bilevel optimization, improves retrieval-augmented LLM agents’ question-answering accuracy while reducing retriever-update memory.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 3B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§