The authors analyze what they call “stale delta,” a mismatch created when quantized values used in the attention forward pass are inconsistent with those used to calculate gradients.
Delta-Matching stabilizes fully FP8 attention training across model scales
The method corrects a forward-backward mismatch in quantized softmax attention, matching mixed-precision training across tested architectures and stages.
Big Tech
Haozhan Tang · Hao Kang · Han Cai · Song Han · Chenyan Xiong
Carnegie Mellon University · NVIDIA · Massachusetts Institute of Technology
Research Digest··3 min read
Tang and colleagues identify a numerical inconsistency that can make FP8 attention gradients accumulate optimization error, particularly as model size and training duration increase.
Why this paper
From NVIDIA and 2 others
In one line
Delta-Matching restores softmax gradient's zero-row-sum invariant, enabling 8-bit FP8 attention training that matches BF16/FP32 performance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§