The authors analyze affine normalization methods such as Reversible Instance Normalization, which scales each input series and can return predictions to their original units before loss calculation.
Computing loss on normalized targets prevents scale from steering training
Across four forecasting architectures, scale-invariant training consistently reduced error by removing unintended gradient weighting from high-magnitude series.
Big Tech
Ignacy Stepka · Willa Potosnak · Kin G. Olivares · Artur Dubrawski
Carnegie Mellon University · Amazon · Nixtla
Research Digest··2 min read
Stepka et al.
Why this paper
From Amazon and 2 others
In one line
Computing homogeneous residual loss on scaled targets prevents series magnitude from weighting gradients and improves accuracy in foundation-model and supervised forecasting.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§