It would be great to do a double blind test of the technology.
From what that paper that was linked to earlier (which seemed like a survey paper) said, the technology is pretty much the application of some standard statistical tools to speech data.
There were two parts in there that I know nothing about: the first is the preprocessing of the data. All of these models take vectors as input, so the continuous sound waves of speech have to be in some way converted to vectors of "features" that the models work on. Not being an EE and not working in this field, I don't know a ton about this stuff, but the stuff in that paper didn't seem too outre or voodooish. The second part was all the stuff about supervectors: I know nothing at all about this stuff.
But I've used the basic methods in other domains for years, they're definitely not voodoo.