Here is a link to a youtube video giving a demonstration on how Easy Voice works.
pt 1
pt 2
pt 3
ETA: apparently I can't count because it's three videos.![]()
I saw those. They don't actually tell me how it works. I want to know what kind of models they used, how feature selection was done, what kind of standardization was done, how one evaluates the quality of a model once built, etc.
I suppose one could do the equivalent of generating a confusion matrix at very small scale by running a small number of samples through the software, and seeing how the numbers work out.
But, for instance, when he says in the video "anything over 85% is very good", I don't really know what that means without seeing what the accuracy of the thing over some larger corpus of examples was.
If what they did was build statistical models, then they had to have already done this, so it would be easy enough to just say.
Perhaps what they did was nothing like what was in that paper from Finland, and that isn't relevant at all, but then it would be nice to know what exactly the software does.
All of this isn't to say the software doesn't work. But it would be a lot easier to say if they simply made some of this information available.