Surrogate benchmarks enable systematic evaluation of scaling analysis methods across model types

Authors introduce ScAn-Bench with thousands of checkpoints to test data acquisition and extrapolation for LLMs and VLMs

Academic
Artin Sermaxhaj · Nastaran Alipour · Donat Sinani · Johannes Hog · Neeratyoy Mallik · Jenia Jitsev · +1 more

University of Freiburg · Zuse School ELIZA · Juelich Supercomputing Center (JSC), Research Center Juelich (FZJ)

Research Digest··2 min read
The authors present two surrogate benchmarks, ScAn-Bench-LLM and ScAn-Bench-VLM, built from 4524 and 8024 checkpoints of language and vision-language model pipelines, respectively.

The authors collected training checkpoints from 1,000 training runs for each of two model families: LLMs (16M to 1B parameters) and VLMs (2M to 357M parameters).

Why this paper

From University of Freiburg and 2 others · Released code

In one line

ScAn-Bench introduces surrogate benchmarks and the first systematic evaluation of scaling analysis methodology across LLMs and VLMs.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.