Choosing speech encoder layers on evaluation data inflates depression scores

Across two clinical speech datasets, nested validation removed optimism caused by testing many encoder layers and changed which representations appeared strongest.

Independent
Paula A. Perez-Toro · David Gimeno-Gómez · Daniel Rückert · Andreas Maier
Research Digest··3 min read
Perez-Toro and colleagues examine a common evaluation shortcut in speech-based depression detection: selecting an encoder’s best-performing layer using the same cross-validation results reported as final performance.

The authors evaluated layer selection on DAIC-WOZ, a clinical interview corpus used for depression detection, using repeated nested cross-validation across five families of pretrained speech encoders.

Why this paper

Independent

In one line

Choosing speech encoder layers on evaluation folds inflates depression-detection AUC, while nested selection removes the bias and favors a compact affect-prosody representation.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.