• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Trayvon Martin, Vigilante Justice

Status
Not open for further replies.
It would be great to do a double blind test of the technology.

From what that paper that was linked to earlier (which seemed like a survey paper) said, the technology is pretty much the application of some standard statistical tools to speech data.

There were two parts in there that I know nothing about: the first is the preprocessing of the data. All of these models take vectors as input, so the continuous sound waves of speech have to be in some way converted to vectors of "features" that the models work on. Not being an EE and not working in this field, I don't know a ton about this stuff, but the stuff in that paper didn't seem too outre or voodooish. The second part was all the stuff about supervectors: I know nothing at all about this stuff.

But I've used the basic methods in other domains for years, they're definitely not voodoo.
 
The survey revealed that decisions were made in 34.8% of the comparisons

So, they could make no decision at all most of the time.

(1) Only original recordings of voice samples were accepted for examination,
unless the original recording had been erased and a high-quality copy was still
available.

Need two original recordings.

(1) The observed identification and elimination errors probably represent the
minimum error rates expected under actual forensic conditions, since
investigators are not always correct in their evaluation of a suspect’s
involvement, due to limited physical evidence, faulty eyewitness statements, etc.
(2) The stated results should only be considered valid when compared with
examiners having the same qualifications and using the same comparison
procedures.

The error rates given for the times they could decide are minimums, and comparisons to this study are likely not valid.
 
Last edited:
From what that paper that was linked to earlier (which seemed like a survey paper) said, the technology is pretty much the application of some standard statistical tools to speech data.

There were two parts in there that I know nothing about: the first is the preprocessing of the data. All of these models take vectors as input, so the continuous sound waves of speech have to be in some way converted to vectors of "features" that the models work on. Not being an EE and not working in this field, I don't know a ton about this stuff, but the stuff in that paper didn't seem too outre or voodooish. The second part was all the stuff about supervectors: I know nothing at all about this stuff.

But I've used the basic methods in other domains for years, they're definitely not voodoo.
I honestly didn't think they are. However, there is nothing wrong with testing existing technologies. Testing this one specifically using double blind protocol would address any concerns for it. But I'm happy to accept any expert consensus that was based on empirical data.
 
Here's a paper I came across that discusses the current problems in Speech Recognition research.
It seems that this situation has all the hallmarks of what the paper discusses as causing problems with the accuracy of current technology.
Please point to the discussion in that paper that refutes the validity of recorded voice recognition. I don't see it.

9. Summary
We have presented an overview of the classical and new methods of automatic text-independent speaker recognition. The recognition accuracy of current speaker recognition systems under controlled conditions is high. However, in practical situations many negative factors are encountered including mismatched handsets for training and testing, limited training data, unbalanced text, background noise and non-cooperative users.
There's a lot of discussion of technician error and less than ideal equipment in many cases. Sure, a lot of forensic recordings are going to be from every Tom Dick and Harry's little recording devices. But that would not be applicable to a company that has established its expertise in this area and 911 recording equipment.
 
So - Owens is pimping his own software, that may never have actually been used in court. And the paper does not disclose the relationship between Owen and the software.

Can I be a little skeptical of this ?

Hardly anybody else seems to think you can.

Let's face it. People want answers, and they want them now. Not just any answers, but the ones that support their preconceived beliefs, which are, of course, Right.

This is the majority behavior on probably the only remaining site where skepticism even occasionally happens.
 
So, they could make no decision at all most of the time.
That would be correct. And as I noted in my post above, forensic recordings come in all shapes and sizes. In general, 911 dispatch services have sophisticated recording equipment. That leaves the quality of the phone calls. A company with a reputation to uphold and no stake in the result (the news would be marketable no matter how the decision went) would likely say if the audio was of too poor quality to be certain.
 
This company's CV seems well established.

Unless, of course, one's confirmation bias prefers that GZ was honest.

I'd feel a lot better about Easy Voice Biometrics if they had something on their website that said something about how it works.

I understand that their customer base is law enforcement, but having _no technical information at all_ on their website is a little sad.

A whitepaper or something would be nice. I have no confirmation bias that GZ was honest, and I'm not saying that I think this stuff is voodoo.

But I have no idea how Easy Voice Biometrics works, or how accurate it is, and the company that makes it chooses not to make any information available about this.

Perhaps the paper that was posted earlier was not representative of what's out there in the field, but if it is, this is fairly technical stuff, and I'd want to know what it was doing if I were going to trust it for anything.

Software is a technical artifact, and we shouldn't have to take it on faith that it works, at least if it's going to be used in a court of law.
 
Without seeing the actual data and models, I couldn't tell how good these models are. But this general field is not pure ******** on the order of polygraphs.
What these guys claim to be doing certainly is.

It's one thing to identify a voice recorded in a studio, quite another under the conditions they claim to have done it.
 
A company with a reputation to uphold and no stake in the result (the news would be marketable no matter how the decision went) would likely say if the audio was of too poor quality to be certain.
Yeah, a private company with goods to sell has absoutely no interest in getting their names in every major news media in the country. :rolleyes:
 
I honestly didn't think they are. However, there is nothing wrong with testing existing technologies. Testing this one specifically using double blind protocol would address any concerns for it. But I'm happy to accept any expert consensus that was based on empirical data.

I don't think blinding is an issue here. Standard statistical measures are.

In general, what you do with this kind of modeling is break your data into pieces. In the old fashioned way, you would break it into 3 pieces, for training, testing, and validation (you can do n-fold cross-validation, but that's not relevant for now).

You build the model using 1/3 of the data, testing it against the testing data as you go. When you're all done, you run a test against the validation data.

There are all kinds of measures you can use. The simplest, in the case of a 2 class classifier (Martin's Voice/Not Martin's Voice), you output a confusion matrix, which looks like this:

| Martin |Not-Martin
predicted Martin | X | Y
predicted Not-Martin | Z | Q

Where X are the instances of Martin's Voice your predicted correctly, Y are instances of someone else's voice you thought were Martin's etc.

There are measures you can report, like precision (x /(x+z)) etc.

In the field I work in, when I build a model, I do this, and report this before ever using the model.

I feel pretty uncomfortable about this kind of modeling being used without such reporting (if this is in fact what the software in question does: the stuff in the academic paper from Finland that was linked earlier certainly did).
 
What these guys claim to be doing certainly is.

It's one thing to identify a voice recorded in a studio, quite another under the conditions they claim to have done it.

Honestly, I can't say that, since I haven't seen how the models work. It may be a perfectly reasonable application of standard statistical methods. It may not.
 
Honestly, I can't say that, since I haven't seen how the models work. It may be a perfectly reasonable application of standard statistical methods. It may not.
It really doesn't matter how their models work. This assclown claims to have isolated the scream in that recording, simply not possible. And it's a crappy quality recording, and a scream is not the same as a speaking voice.

Pure, utter woo.
 
I don't think blinding is an issue here. Standard statistical measures are.

In general, what you do with this kind of modeling is break your data into pieces. In the old fashioned way, you would break it into 3 pieces, for training, testing, and validation (you can do n-fold cross-validation, but that's not relevant for now).

You build the model using 1/3 of the data, testing it against the testing data as you go. When you're all done, you run a test against the validation data.

There are all kinds of measures you can use. The simplest, in the case of a 2 class classifier (Martin's Voice/Not Martin's Voice), you output a confusion matrix, which looks like this:

| Martin |Not-Martin
predicted Martin | X | Y
predicted Not-Martin | Z | Q

Where X are the instances of Martin's Voice your predicted correctly, Y are instances of someone else's voice you thought were Martin's etc.

There are measures you can report, like precision (x /(x+z)) etc.

In the field I work in, when I build a model, I do this, and report this before ever using the model.

I feel pretty uncomfortable about this kind of modeling being used without such reporting (if this is in fact what the software in question does: the stuff in the academic paper from Finland that was linked earlier certainly did).
Oh, I agree. My point is about compelling evidence for the lay person. If you conduct a double blind study it's easy to understand and compelling. That's all. When the OJ trial announced their DNA findings the jury didn't buy it. Had they also conducted a very simple double blind study where someone matches a target with a sample and then reveals that it was a match to OJ there wouldn't have been all of the head scratching about the science. KISS is perhaps the most important aspect in these types of cases.

Otherwise I agree with you.
 
Last edited:
It really doesn't matter how their models work. This assclown claims to have isolated the scream in that recording, simply not possible. And it's a crappy quality recording, and a scream is not the same as a speaking voice.

Pure, utter woo.

Hmm. I guess I'm just not as certain as you about that.

Without knowing how the sound is converted into feature vectors, I just have no idea what goes on at all, and it would probably take me somewhere between a week and a month of reading and screwing around with code to get to the point where I was confident I understood it, if I wasn't doing anything else, and even then I would not be confident I knew all the details. And I work in a fairly closely related discipline.

This is why I think it's kind of reasonable to expect the forensic software maker to at least have a whitepaper explaining what it does (there are ways around not exposing whatever their "secret sauce" is, software companies do this all the time). At least then, people in closely related disciplines could look at it and have a chance at being able to validate whether it was any good.

This specific field is something I never saw in grad school in computer science: I'm really not sure where the people are who do research in it. Maybe ee?

The idea that someone is going to come in as an "expert" and make a pronouncement about the technical details without a real technical explanation is alien to me: I guess that's how technical stuff in legal cases works though, since you can't expect judges and jury members to be competent in random technical fields that may crop up.
 
Last edited:
Yeah, a private company with goods to sell has absoutely no interest in getting their names in every major news media in the country. :rolleyes:
Right, no private company in the world could possibly be legit or care about their reputation and certainly a news company doesn't know how to hire anyone with forensic expertise. I'll see your :rolleyes: and raise you two. :rolleyes: :rolleyes:
 
It really doesn't matter how their models work. This assclown claims to have isolated the scream in that recording, simply not possible. And it's a crappy quality recording, and a scream is not the same as a speaking voice.

Pure, utter woo.
You do know, I hope, that whatever version of the tape you heard, typical news media get a clean copy directly from the 911 dispatch.
 
Status
Not open for further replies.

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom