The authors developed SAGE, a framework that organizes LLM decision-making into three steps.
Structured reasoning framework improves LLM game-playing efficiency and payoff
SAGE anchors strategic reasoning to an equilibrium prior, adapts to opponent tendencies, and recalibrates using counterfactual history, achieving up to 127.6% payoff gains with fewer tokens.
Big Tech
Zhiwei Chen · Tianchun Wang · Zhongtao Rao · Haiming Zhu · Ding Cao · Tianxiang Zhao
The Hong Kong University of Science and Technology (Guangzhou) · Johns Hopkins University · Microsoft · Fudan University · University of Science and Technology of China
Research Digest··3 min read
The authors propose SAGE, a training-free inference-time framework that structures LLM strategic reasoning around three coordinated operations: anchor, adapt, and recalibrate.
Why this paper
From Microsoft and 4 others · Released code
In one line
Structuring LLM strategic reasoning as anchor, adapt, recalibrate improves payoff and cuts token use in imperfect-information games.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§