\n\nOkay" or "Alright,").
Starting token cues can elicit reasoning from base language models
Fixing just the first two tokens of a response dramatically improves performance on math and coding, rivaling reinforcement learning.
Independent
Sophie L. Wang · Amil Dravid · Rulin Shao · Kevin Farhat · Sewon Min · Alexei A. Efros
Research Digest··3 min read
, "Okay" or "Alright,").
Why this paper
Independent
In one line
Base models can reason well when response-opening token cues activate behaviors learned from training data, and reinforcement learning partly works by making those cues likelier.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§