, scope or type check) to determine the allowed next tokens, renormalizes the softmax probabilities over that set, and computes cross-entropy against the correct token.
Training language models to ignore constraints that will be applied at inference improves efficiency
The authors show that a constraint-aware loss, which only penalizes tokens allowed by the prefix analysis, reduces prediction loss at matched model size and data.
Big Tech
Jinwoo Kim
Microsoft Research · University of California-San Diego
Research Digest··3 min read
Thread:Constraint-Aware Training
Jinwoo Kim proposes a constraint-aware (CA) training objective that externalizes prefix-definable analyses from the language model, allowing the model to ignore tokens that will later be filtered by a constrained decoder.
Why this paper
From Microsoft Research and University of California-San Diego
In one line
Constraint-aware training reduces model size and data needs by externalizing decoding-time analyses from the training objective.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§