The authors tested 9 LLMs (including Llama, Mistral, and GPT series) on 6 standard benchmarks: MMLU for MCQA, GSM8K for math reasoning, HumanEval for code generation, and FLORES-200 for machine translation, among others.
Current date in system prompts skews LLM evaluation results
The authors show that LLM accuracy varies by up to 14% on math reasoning and reshuffles leaderboards when only the date changes, making the effect larger than other known sources of non-determinism.
Academic
Mario Sanz-Guerrero · Minh Duc Bui · Manuel Mager · Katharina von der Wense
Johannes Gutenberg University Mainz · Universidad Iberoamericana · University of Colorado Boulder
Research Digest··3 min read
Sanz-Guerrero et al.
Why this paper
From Johannes Gutenberg University Mainz and 2 others
In one line
Hidden dates in system prompts cause LLM evaluation scores to fluctuate by up to 14% and reorder model rankings.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (4 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§