What they did
The authors retrospectively studied 29,116 and 7,691 adults meeting Sepsis-3 criteria at two hospital systems. Their model used 43 routinely recorded variables across a 72-hour treatment window to generate a continuous hourly severity score.
Rather than treating each patient-hour as independently associated with survival or death, they used mortality as a trajectory-level ranking signal and allowed risk credit to vary across time. Evaluation used a permanent 20% test holdout, clinical vignettes, Spearman correlations, and patient-level bootstrap resampling for uncertainty intervals.
Key findings
- Within every baseline SOFA-2 stratum, non-survivors scored 1.19–1.64 points higher than survivors on the 0–10 scale; separation also persisted within strata defined by lactate, mean arterial pressure and creatinine.
- Within-patient score changes correlated with changes in lactate at Spearman ρ = 0.39 across 1,854 patients. Associations with mean arterial pressure and creatinine were weaker.
- Models trained at different institutions achieved 70–77% of the corresponding same-site correlation, indicating substantial but incomplete cross-site agreement.
- External within-patient correlations were 0.54 and 0.59, compared with estimated same-site ceilings of 0.92 and 0.90. The learned score also correlated with established severity indices, while null controls remained near zero.
Why it matters
A continuously updated learned score could represent changes in sepsis severity more finely than legacy indices built from fixed variables, weights and discrete thresholds. The trajectory-ranking approach is also useful when only patient-level outcomes are available, because it avoids imposing the same mortality label on every state during a hospital course.
Caveats
This was a retrospective study, so it does not establish that displaying the score improves treatment decisions or patient outcomes. Mortality is an indirect and care-dependent supervision signal, cross-institutional agreement remained below same-site performance, and validation was limited to two hospital systems; prospective testing, calibration studies and evaluation across broader populations are still needed.