• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Trayvon Martin, Vigilante Justice

Status
Not open for further replies.
I just want to know where Mr Owen got the 40+ points of comparison (provided that the media reports of 48% match are true) from the snippet of questioned audio the he claims would be required for eliminating Zimmerman as the donor - or even the 20 points he would need for a positive elimination assuming a 0% match rate.
Read my last post.
 
Based purely on the definitions of the words, the older spectrographic analyses would look purely at the audio analysis.

Adding a biometric component one would match certain sounds to the shape of the throat and larynx. So one could then compare different sounds.

Say for example throat/mouth/larynx A results in a range of sounds and you can make a model of those sounds. You also take throat/mouth/larynx B and make a similar map.

Now you compare the two and you see where they overlap and where they are distinctly different.

Then you do the same with a very large sample of throat/mouth/larynx recordings.

What you come out with in the end is a pattern that reflects the physical differences between individual anatomy reflected on an audio wave screen. Just like fingerprints don't rely on matching one print to another which with 6 billion people on the planet would be useless, different parts of the print are categorized. You can then eliminate various prints as not matching right away because they don't have certain characteristics. You don't have to match the whole print.

TM and GZ have different shaped throats/mouths/larynxes. If the voice identification/recognition software has identified certain key points in an audio recording that are recognizable as matching or not matching a certain biometric then spectrographic and biometric techniques would indeed be very different. The former would rely on simple matching while the latter would rely on matching the person's anatomical differences as reflected by the spectrographic print out.
And how does that work when there's been at least 1 and maybe 2 layers of lossy digital audio compression - one from the cordless/cell phone and one from the recording device at the 911 call center?

Every time you compress the signal, data is lost. And not just a little bit of data, a lot of data! More than 90% of it as a matter of fact.
 
Last edited:
Here are more excerpts from that first paper (2009):

[Page 2]: "In general, phonetic variability represents one adverse factor to accuracy in text-independent speaker recognition. Changes in the acoustic environment and technical factors (transducer, channel), as well as
“within-speaker” variation of the speaker him/herself (state of health, mood, aging) represent other undesirable factors. In general, any variation between two recordings of the same speaker is known as session variability [111, 231]. Session variability is often described as mismatched training and test conditions, and it remains to be the most challenging problem in speaker recognition. "

[Page 16]: "Any variation in different utterances of the same speaker, as characterized by their supervectors – be it due to different handsets, environments, or phonetic content – is harmful."


I encourage everyone to read the entire paper, but certainly screaming versus spoken would qualify as severe "within-speaker" variation and have a corresponding effect on accuracy.

Yes, but then they go on to talk about how various dimensionality reduction schemes reduce this problem. I noted this earlier. You have taken this quote out of context.

Having looked at a few papers, the thing they seem more concerned with is the kind of systematic differences caused by different headsets or microphones. But, again, this has not been presented as an insurmountable problem.
 
And it looks like I'm right.

Biometrics

The entire article explains the difference but you can read page 3 if you want to cut to the chase. (The article is copy paste protected).

And how does that work when there's been at least 1 and maybe 2 layers of lossy digital audio compression - one from the cordless/cell phone and one from the recording device at the 911 call center?

Every time you compress the signal, data is lost. And not just a little bit of data, a lot of data! More than 90% of it as a matter of fact.
See the intro on page 1.
 
Last edited:
The stuff Owen is doing really does seem as much art as science. That's one of the things that makes the machine learning approach seem attractive: there is still some art in the preprocessing, but overall it's a lot easier to quantify. I would imagine that the intelligence community also likes the idea of looking at a bunch of phone calls in an automated way to see if they can find a particular person's voice, which wouldn't be practical on Owen's approach.

The big open question none of use knows the answer to is exactly how good the machine learning stuff is and whether it's admissible in court. It looks like at least some software based on this approach has been ruled to be.
Agreed,

To put a finer point on it, we don't know if this particular person's software and specific technique are going to pass (or have already passed) the Frye test...

And we don't know if by the trial more rigorous testing using previously accepted methods will confirm or confound these results.

And last, we don't know if the defense would be able to convince the jury to believe the witness narratives, or some other bit of evidence that may not have even been released yet.
 
Again with ignoring the point of the paper which was to point out all the problems and then NOTE THAT NEWER TECHNOLOGIES OVERCOME THESE LIMITATIONS.

It's obvious you didn't read or understand the paper, or the phrase "it remains." The excerpt from page 16 specifically relates to supervectors.

From the summary:

"It remains a great challenge in the near future to understand what features to exactly look for in speech signal."

"However, many research problems remain to be addressed, such as human-related error sources, real-time implementation, and forensic interpretation of speaker recognition scores."

"There are also many other factors that have impacts on the speaker recognition performance. We should also address human-related error
sources, such as the eff ects of emotions
, vocal organ illness, aging, and level of attention."

These obstacles have not been overcome. These are the goals for the future direction of the technology.
 
And based on the level of skepticism being directed at the audio analysis we should accept this new "evidence" because... ?

You shouldn't just accept it. You should take it with a giant grain of salt. Just like the 'enhanced audio' of what GZ said on the phone. It's not conclusive.

I think we all knew it was a leak, and we would still have the paramedic records, police photos, etc to document injuries, and this video was not terribly important.
 
Yes, but then they go on to talk about how various dimensionality reduction schemes reduce this problem. I noted this earlier. You have taken this quote out of context.

Having looked at a few papers, the thing they seem more concerned with is the kind of systematic differences caused by different headsets or microphones. But, again, this has not been presented as an insurmountable problem.

It's not out of context. It's reiterated in the summary: "We should also address human-related error sources, such as the eff ects of emotions..."

The phrase "should also address" means it has not been addressed, nor resolved.
 
Shocking. All of the internet doctors already informed us there was no visible evidence.

http://www.internationalskeptics.com/forums/imagehosting/thum_500444f74b7a3cd5dd.jpghttp://www.internationalskeptics.com/forums/imagehosting/thum_500444f79e97d55e72.jpg

And based on the level of skepticism being directed at the audio analysis we should accept this new "evidence" because... ?

You shouldn't just accept it. You should take it with a giant grain of salt. Just like the 'enhanced audio' of what GZ said on the phone. It's not conclusive.
If two experts in the field claimed it was evidence of a gash then it tentatively would fulfill the preponderance of evidence standard for me. I wouldn't accept it as conclusive but I would form the opinion that there was a gash.
 
Last edited:
Agreed,

To put a finer point on it, we don't know if this particular person's software and specific technique are going to pass (or have already passed) the Frye test...

And we don't know if by the trial more rigorous testing using previously accepted methods will confirm or confound these results.

And last, we don't know if the defense would be able to convince the jury to believe the witness narratives, or some other bit of evidence that may not have even been released yet.

Yes, I agree completely.
 
It's not out of context. It's reiterated in the summary: "We should also address human-related error sources, such as the effects of emotions..."

The phrase "should also address" means it has not been addressed, nor resolved.
It's a bit out of context, since the next section of the paper was about using dimensionality reduction to address the problem.

Actually it seems that addressing those issues is pretty much what people have been doing the last few years, although they seem to be more concerned with channel problems than speaker related problems. It should also be noted that this particular field is moving really quickly right now (everything to do with machine learning is: it's a somewhat fashionable, high visibility domain), so people have been addressing theses issues for a few years.

But this particular issue certainly hasn't dominated the papers, nor has it been mentioned in every paper (or even more than a couple of the ones I looked at), and certainly the within speaker issue hasn't been dwelled on a lot: mostly the channel problem. Although to be fair to wildcat, it really looks like we were talking about a completely disjoint set of papers on essentially different topics.

This paper is kind of interesting:

http://www.icsi.berkeley.edu/pubs/speech/larastoll-dissertation-final.pdf

According to this, there are particular _speakers_ whose voices intrinsically cause a lot of errors. Not their emotional state or whatever, but just their voice. How the hell would anyone know from a single test if a particular speaker was one of these?

Ultimately, we will have to see if this data is presented with some numbers to back up its performance.

And to all of the armchair quarterbacks pretending to know how this stuff works and claiming they _know_ the screaming is a problem. You don't. This is a complex discipline involving a ton of math, and a bunch of signal processing, and nobody here is an expert in it (me included: I at least am an expert in something similar enough that I can tell how inexpert I am). This is something that needs to be judged by someone who knows the field well. Nobody here does (and I'm pretty doubtful that the guys in question do). Not one person on this thread. Not even close. So how about we quit pretending and let the thing work out using the old adversarial technique?
 
Last edited:
Sorry, this is a different paper. I'm referring to Cylinder's posts not this one.

I'll revise my reply after walking my dogs. In the meantime, Owen words are worth noting so I'll leave them up.
It's obvious you didn't read or understand the paper, or the phrase "it remains." The excerpt from page 16 specifically relates to supervectors.

From the summary:

"It remains a great challenge in the near future to understand what features to exactly look for in speech signal."

"However, many research problems remain to be addressed, such as human-related error sources, real-time implementation, and forensic interpretation of speaker recognition scores."

"There are also many other factors that have impacts on the speaker recognition performance. We should also address human-related error
sources, such as the effects of emotions
, vocal organ illness, aging, and level of attention."

These obstacles have not been overcome. These are the goals for the future direction of the technology.
I read it and quoted from it You are referring to a different paper and I stand by my quote.

Owen is saying the courts may not be up to speed with the advances in the field because past cases dealt with older technologies. The conclusion continues:
Court decisions reviewing the early voice identification cases may not be relevant to present day cases because the older decisions were based on less sophisticated procedures. Most of the courts which have rejected admission have been aware of continuing work in this field and have specifically left the door open as to future admissibility.
Proper presentation and explanation of the research pertaining to spectrographic voice identification analysis will allow the courts to better understand the accuracy and reliability of the spectrographic voice identification method. When the research is properly presented, the studies show that properly trained individuals, using standard methodology, produce accurate results.
Got that last line? Let me make sure you caught that: PRODUCE ACCURATE RESULTS
 
Last edited:
And to all of the armchair quarterbacks pretending to know how this stuff works and claiming they _know_ the screaming is a problem. You don't.
Yes, we do. Because the very experts you cite tell us it is.

It is you who is disagreeing with the experts, not me. The experts tell us you have to compare like to like, even the same phrase verbatim.

Not a single expert in any published paper claims to be able to compare a scream to a speaking voice.
 
Again with ignoring the point of the paper which was to point out all the problems and then NOTE THAT NEWER TECHNOLOGIES OVERCOME THESE LIMITATIONS.
This is where the discussion went off track. My bad. My point is still valid but I'll look at the 2009 paper in a bit.
 
New york times article

http://www.nytimes.com/2012/04/02/u...review-of-ideals.html?_r=1&src=me&ref=general

A 7 page article about the events.

In the attempt to make it a more interesting and detailed narrative, they do things like this, though:

Hey, we’ve had some break-ins in my neighborhood,” Mr. Zimmerman said to start the conversation with the dispatcher. “And there’s a real suspicious guy.”
This guy seemed to be up to no good; like he was on drugs or something; in a gray hoodie. Asked to describe him further, he said, “He looks black.”


BTW - they say TRUCK not SUV
 
Status
Not open for further replies.

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom