Leto speeds up LLM training recovery by reusing state on surviving hardware

The system retains model and process state on working GPUs and preinitializes replacement trainers, avoiding checkpoint reloading and recomputation.

Academic
Geon-Woo Kim · Joon Ha Kim · Daehyeok Kim

The University of Texas at Austin

Research Digest··2 min read
Leto is a fault-tolerant training system that leverages surviving hardware after hardware-operable failures (HOFs) to recover quickly.

The authors identified that for hardware-operable failures, the state needed to resume training can be retained or prepared outside the active trainer while remaining on the same hardware.

Why this paper

From The University of Texas at Austin

In one line

Leto recovers LLM training from hardware faults faster by retaining state and preinitializing processes on surviving GPUs.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (gpus 64)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.