Bilevel optimization bridges the retriever-policy credit gap in agentic RL

BRIDGE jointly adapts retrieval and LLM policy in a hierarchical order, improving multi-hop QA accuracy by up to 9.6 EM points.

Academic
Quan Xiao · Mingda Liu · Gaowen Liu · Katsuki Fujisawa · Tianyi Chen

Cornell University · Institute of Science Tokyo · Cisco Research

Research Digest··3 min read
The authors identify an information-credit gap in agentic reinforcement learning (ARL) for LLMs: when failures arise from poor retrieval, the fixed retriever escapes blame while the LLM policy is penalized.

The authors first demonstrate that retrieval and LLM policy learning are order-sensitive: adapting the retriever before the policy yields larger reward gains than the reverse order.

Why this paper

From Cornell University and 2 others · Part of RL for Tool Agents, now 14 papers

In one line

Training retrieval before policy optimization, then co-adapting both with bilevel optimization, improves retrieval-augmented LLM agents’ question-answering accuracy while reducing retriever-update memory.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 3B)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.