Standard BCE materializes a dense [B, N, V] logits tensor causing OOM for large V.
CutBCE eliminates memory barriers for large-vocabulary recommendation training
On an 876k-item dataset, the authors report 65.7% lower memory use and 225.9% faster training with no accuracy loss.
Big Tech
Yaoyiran Li · Haowen Ning · Mohamed Hammad
Google Cloud
Research Digest··2 min read
The authors present CutBCE, an exact hardware-accelerated Binary Cross-Entropy loss operator for JAX/TPU that avoids materializing full logits in high bandwidth memory.
Why this paper
From Google Cloud
In one line
CutBCE eliminates the memory bottleneck of binary cross-entropy loss for large-vocabulary recommendation by never materializing the full logits tensor.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (hardware 8-chip TPU slice)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§