LoGRA forms low-rank sketches directly during backpropagation, avoiding construction and storage of full gradients for selected weight matrices.
Low-rank gradient sketches cut memory costs in language-model reinforcement learning
LoGRA compresses gradients and optimizer state while controlling policy-update size, enabling reinforcement learning of a 27-billion-parameter model on eight GPUs.
Big Tech
Shaokun Zhang · Yifan Zhang · Jian Hu · Yueying Li · Hao Zhang · Binfeng Xu · +2 more
NVIDIA
Research Digest··3 min read
Zhang et al.
Why this paper
From NVIDIA
In one line
LoGRA reduces LLM reinforcement learning memory by up to 45.7% using low-rank gradient sketches and KL step control.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§