non-determinism
Same inputs give different outputs across runs, confounding testing and evaluation.
- Papers
- 4
- Released code
- 1
- First seen
- Sept 2026
- Latest
- Sept 2026
4 papers in the last two months, against 0 in the two before.
The papers
Most central to this idea first, not most recent.
- Top Universitycs.AI
LLM judges as measurement instruments fail reliability tests on shared endpoints
University of Sheffield, Ranplan Wireless Network Design Ltd. · Sept 2026
- Top Universitycs.CR
Unverified state text can flip calibrated model decisions
Institute of Information Engineering, Chinese Academy of Sciences, University of Chinese Academy of Sciences · Sept 2026
- Big Techcs.CLcode
Cut-point replay turns recorded agent failures into regression tests
Microsoft · Sept 2026
- Big Techcs.AI
Memory guidelines help language-model agents succeed consistently across repeated runs
IBM Software Innovation Lab, IBM Research · Sept 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.