The authors collected training checkpoints from 1,000 training runs for each of two model families: LLMs (16M to 1B parameters) and VLMs (2M to 357M parameters).
Surrogate benchmarks enable systematic evaluation of scaling analysis methods across model types
Authors introduce ScAn-Bench with thousands of checkpoints to test data acquisition and extrapolation for LLMs and VLMs
Academic
Artin Sermaxhaj · Nastaran Alipour · Donat Sinani · Johannes Hog · Neeratyoy Mallik · Jenia Jitsev · +1 more
University of Freiburg · Zuse School ELIZA · Juelich Supercomputing Center (JSC), Research Center Juelich (FZJ)
Research Digest··2 min read
The authors present two surrogate benchmarks, ScAn-Bench-LLM and ScAn-Bench-VLM, built from 4524 and 8024 checkpoints of language and vision-language model pipelines, respectively.
Why this paper
From University of Freiburg and 2 others · Released code
In one line
ScAn-Bench introduces surrogate benchmarks and the first systematic evaluation of scaling analysis methodology across LLMs and VLMs.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§