A company noticed, more than a year later, that its card processor had been advancing its receivables from the very start: the credit due in three days landed in the account the next day, and the fee charged for that advance was steep. The business owner insisted he had never authorized it, and the lawsuit was filed. The processor, in its defense, produced its best weapon: the recording of the call in which the arrangement had supposedly been approved.
The audio seemed perfect. An accredited technician opens the call, identifies himself by employee number, and states that the client wants to advance receivables. The processor's attendant asks to confirm with the account holder. Seconds later, a voice takes over the line, identifies itself as the owner, and approves the transaction. The forensic speaker comparison examined the recording against the business owner's authentic voice samples and concluded: that voice was not his. That finding alone dismantled the charges. But the examination went further and found the detail that turns the case into a lesson: the voice impersonating the owner displayed the speech characteristics of the technician himself, present on the same call, who had identified himself at its start. The impostor and his "witness" were the same person. In scenarios like this, the improper charges are subject to restitution, as a rule doubled (Consumer Protection Code, art. 42, sole paragraph).
What speaker comparison examines
Forensic speaker comparison confronts the questioned speech with a known sample of the speaker, combining two dimensions. In perceptual analysis, the trained examiner studies vocal quality, articulatory patterns, rhythm, speech habits. In acoustic analysis, the physical parameters of the signal are measured: fundamental frequency, formants (the resonances that each person's vocal tract imprints on sounds), characteristics that cannot be consciously controlled. It is the combination of the two readings, documented and reproducible, that sustains a forensic conclusion.
Similarity is not identity
The most common error in voice reports that arrive for rebuttal is a conclusion built solely on vocal similarity. Similar voices exist by the thousands, and a good impressionist fools any audience by reproducing a famous actor's timbre. What the impressionist cannot reproduce are the parameters he neither hears nor controls: the formant structure of his own vocal tract, the articulatory micro-details, the physical signature of the instrument that produces the speech. The examiner without specific training in phonetics judges by ear, like the impressionist's audience, and commits in the report the same mistake the audience commits in the theater.
"Artificial intelligence can assist the analysis, but whoever signs the report answers for it civilly and criminally. And who is held accountable if the wrong conclusion came from the AI? The expert is not permitted to be wrong."
The tool helps; the expert answers
Voice comparison software exists, some of it with artificial intelligence features, and it has a legitimate place as support for the examination. What does not exist is a expert report outsourced to the machine. Every automated answer must be validated by someone who masters the technique, because systems make mistakes: they classify as positive what is false, and vice versa. The expert who signs a conclusion he cannot verify assumes a risk he does not control, and the law does not forgive him: false expert testimony is a crime (Penal Code, art. 342), and civil liability for the report is personal. The tool does not take the stand in place of the person who signed.
The asymmetry the court needs to understand
Voice examination does not deliver a binary answer, and this is the part that most surprises those who commission it. The conclusion is graded in levels, which inform the court how much of the speaker's characteristics the examined material carries. And there is a fundamental asymmetry: when the questioned voice does not display the natural characteristics of the suspect's speech, the examiner can exclude categorically. The reverse does not exist. However much the characteristics converge, a positive identification never reaches absolute affirmation: the voice is biometric data, but its individualization power differs from that of a fingerprint or a DNA test. The honest report says "very high compatibility"; the reckless report says "it is him, without a doubt". Distrust the second.
The triage that saves the case
A share of the audio files that reach the laboratory lacks the technical condition to sustain an examination or serve as evidence, and discovering that early is the difference between strategy and disaster. The party that files suit betting on an unusable audio file discovers the problem at the worst moment: during the court-ordered examination, with the case already underway. That is why preliminary analysis is essential, and it has a double utility: it identifies what is useless, and it recovers what seems lost. Many apparently unusable recordings can be processed and filtered until an examination becomes feasible, although that is not the rule. Factors that weigh in favor:
- Sufficient net speech from the questioned speaker (below roughly 10 seconds, a viable examination is rare);
- Recording quality and channel compatible between the questioned material and the reference sample;
- A voice reference contemporaneous with the questioned recording;
- Absence of overlapping voices and of dominant noise in the critical passages.
The other side of the triage is the collection of the reference sample, which weighs as much as the questioned audio. The ideal reference is collected in a controlled recording, combining reading and spontaneous speech, long enough to reveal the speaker's natural range of variation and, whenever possible, through a channel compatible with the questioned recording: a telephone voice compares better with a telephone voice. A poor reference produces a poor examination, however good the questioned material may be.
In the card processor case, the voice decided everything: it undid a debt built on an internal fraud and pointed to its author. That is what speaker comparison does when conducted with method: it turns "my word against the recording" into a demonstrated fact, with the limitations declared and the conclusion sized exactly to what the technique allows one to affirm.
