llm as a judge
- Papers
- 6
- Released code
- 2
- First seen
- Aug 2026
- Latest
- Sept 2026
6 papers in the last two months, against 0 in the two before.
The papers
Most central to this idea first, not most recent.
- Independentcs.AI
Deterministic guardrails expose gains that fool LLM-based judges
Sept 2026
- AI Startupcs.AIcode
Goal pressure exposes wide gaps in security agents’ scope adherence
dreadnode · Sept 2026
- Industrycs.CRcode
Cyber agents hide critical weaknesses behind successful attack workflows
National Research Council, University of Windsor · Sept 2026
- Top Universitycs.CL
LLM agents struggle to infer wellbeing from long-term wearable data
Dartmouth College, University of Cambridge · Aug 2026
- Top Universitycs.LG
Embedding retrieval ranks matching words above shared underlying structure
Independent, MIT CSAIL · Sept 2026
- Industrycs.CL
Self-evolving loop synthesizes high-quality multimodal training data
vivo AI Lab · Aug 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.