Here are more excerpts from that first paper (2009):
[Page 2]: "In general, phonetic variability represents one adverse factor to accuracy in text-independent speaker recognition. Changes in the acoustic environment and technical factors (transducer, channel), as well as
“within-speaker” variation of the speaker him/herself (state of health, mood, aging) represent other undesirable factors. In general, any variation between two recordings of the same speaker is known as session variability [111, 231]. Session variability is often described as mismatched training and test conditions, and it remains to be the most challenging problem in speaker recognition. "
[Page 16]: "Any variation in different utterances of the same speaker, as characterized by their supervectors – be it due to different handsets, environments, or phonetic content – is harmful."
I encourage everyone to read the entire paper, but certainly screaming versus spoken would qualify as severe "within-speaker" variation and have a corresponding effect on accuracy.