Small-Sample LLM Rankings Often Overstate Evaluation Reproducibility

A self-audit of eight models found that repeated outputs, model rankings and headline conclusions were sensitive to sampling and analysis choices.

Industry

Skelf Research

Research Digest··2 min read
Sarkar repeatedly tested eight open model variants on prompt-structure inference, preserving 293 intermediate representations and auditing both output consistency and the resulting rankings.

The author evaluated eight open model variants from five families, ranging from 8 billion to 675 billion parameters, on recovering structured representations from prompts.

Why this paper

From Skelf Research

In one line

LLM evaluation rankings often look definitive but are highly unstable under resampling and defensible analysis changes.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.