The authors retained the context-conditioned Transformer architecture used for in-context reinforcement learning but replaced behavior cloning with a Bellman-style objective.
Q-target training improves in-context reinforcement learning from suboptimal data
Replacing behavior imitation with Bellman-style value targets made context-conditioned Transformers more robust to weak offline trajectories.
Chinese Tech
Shanghai Jiao Tong University · Alibaba Group · University at Buffalo
Research Digest··2 min read
Lin et al.
Why this paper
From Alibaba Group and 2 others
In one line
QTPT enables robust in-context RL from weak offline data by training with Q-targets instead of behavior cloning.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§