Low-rank gradient sketches cut memory costs in language-model reinforcement learning

LoGRA compresses gradients and optimizer state while controlling policy-update size, enabling reinforcement learning of a 27-billion-parameter model on eight GPUs.

Big Tech
Shaokun Zhang · Yifan Zhang · Jian Hu · Yueying Li · Hao Zhang · Binfeng Xu · +2 more

NVIDIA

Research Digest··3 min read
Zhang et al.

LoGRA forms low-rank sketches directly during backpropagation, avoiding construction and storage of full gradients for selected weight matrices.

Why this paper

From NVIDIA

In one line

LoGRA reduces LLM reinforcement learning memory by up to 45.7% using low-rank gradient sketches and KL step control.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.