consequence-aware evaluation
Evaluation that assesses model performance by considering the real-world impact or risk of errors, not just standard accuracy.
- Papers
- 3
- Released code
- 1
- First seen
- Aug 2026
- Latest
- Aug 2026
3 papers in the last two months, against 0 in the two before.
The papers
Most central to this idea first, not most recent.
- Top Universitycs.CLcode
Semantic scores overstate language-model safety in air traffic control
ATMRI, Nanyang Technological University (NTU) · Aug 2026
- Top Universitycs.AI
Financial agents can cite rules while still attempting prohibited trades
HKUST, HKBU · Aug 2026
- Industrycs.AI
Medical AI models often give right answers with disconnected reasoning
Eindhoven University of Technology, Dana-Farber Cancer Institute · Aug 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.