Really? You've actually read and understood papers in the field? You know what, I really don't believe this to be the case. I read a bunch last night, following citations from the bibliography in
http://cs.joensuu.fi/pages/tkinnu/webpage/pdf/speaker_recognition_overview.pdf
and not one of them mentioned this "fact". Do you have any citations?
Here's a related snippet from that paper:
"For a given speaker, the supervectors estimated from
different training utterances may not be the same especially when these training samples come from different
handsets. Channel compensation is therefore necessary
to make sure that test data obtained from dierent channel (than that of the training data) can be properly scored
against the speaker models. For channel compensation
to be possible, the channel variability has to be modelled explicitly. The technique of joint factor analysis
(JFA) [110] was proposed for this purpose"
In other words, the problem of the speaker sounding radically different from one sample to another is expected, and the modelling is built around this assumption.
As I said, you are just making things up. So far, you have displayed zero knowledge of the subject.