The authors started from native dual-mode Qwen3 checkpoints and jointly trained routing and mode-conditioned responses, without a method-specific supervised fine-tuning warm-up.
Models learn when to reason, cutting tokens while improving accuracy
CounterRoute trains a dual-mode language model to choose between deliberate reasoning and direct answers while separately improving its routing and response generation.
Big Tech
Ruochen Jiao · Besnik Fetahu · Zhenyu Shi · Priyanka Nigam
Amazon.com
Research Digest··2 min read
Jiao et al.
Why this paper
From Amazon.com
In one line
CounterRoute jointly learns routing and response modes in dual-model LMs, cutting generated tokens by 41-51% while improving accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§