In May I started a visiting research collaboration with Prof. Hamdi Al-Jamimi at King Fahd University of Petroleum and Minerals, on sepsis clinical decision support.
Hospitals are starting to deploy models that recommend a sepsis treatment together with an explanation of why, for example this patient's lactate or this blood-pressure trend. The explanation is what a clinician reads. It is rarely checked against what the model actually did.
The test
Take a piece of evidence the explanation names, perturb it, and check whether the recommendation shifts the way the explanation predicts. If the explanation says lactate drove the decision and changing the lactate changes nothing, the explanation is not describing the model. That recommendation gets flagged as unreliable, however accurate the underlying prediction is.
Data, and the step I care about most
Cases are labelled with the Sepsis-3 consensus criteria. The plan is MIMIC-IV and the PhysioNet 2019 Sepsis Challenge data first, then HiRID, a Swiss ICU cohort.
The last step is the one I care most about. A model whose accuracy transfers from US hospitals to a Swiss one may not carry its explanations with it, and almost no sepsis AI work checks that.