The authors express each collective through its input and output data layouts plus a copy or reduction operation.
Purlin decouples GPU collective coordination from hardware data movement
A shared orchestration protocol lets seven collective operations reuse coordination logic while adopting GPU-specific copy and reduction mechanisms.
Big Tech
Osayamen Jonathan Aimuyo · Swapnil Gandhi · Christos Kozyrakis
Stanford University · NVIDIA
Research Digest··2 min read
Aimuyo, Gandhi and Kozyrakis present Purlin, a framework for collective communication among GPUs connected within a high-bandwidth scale-up domain.
Why this paper
From NVIDIA and Stanford University · Released code
In one line
Purlin decouples collective orchestration from datapath, enabling up to 5.14x latency speedup and 4.50x bandwidth improvement.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (4 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§