Home // FORENSIC UNITS // DOCUMENT, VOICE AND IMAGE

Forensic Phonetics and Voice Examination

The voice is not a fingerprint, and that is precisely why method matters: specialist training in Forensic Speaker Identification (CEFAC), a hybrid perceptual-acoustic examination and probabilistic conclusions built to withstand adversarial scrutiny.

What voice examination answers

Nine examinations, one precise vocabulary

The market confuses identification, comparison, verification and speaker profiling. Choosing the right examination decides the value of the evidence; each one answers a different question.

01

Speaker comparison

The central task: the examination compares the questioned sample with the known sample of the person being compared and concludes in terms of probabilistic support, on the 9-point scale (from -4 to +4, with zero as inconclusive), assessing similarity and typicality: how alike the features are and how common they are in the population.

02

Speaker identification

An umbrella term: placing an unknown speaker inside or outside a set of candidates. In a closed-set scenario, the speaker is among them; in an open-set scenario, he may not be, and the expert report states that difference before concluding.

03

Comparison of interlocutors

Who is who within the same dialogue: wiretaps, group audio messages and recorded meetings with multiple voices have each utterance attributed to the corresponding speaker.

04

Speaker verification (voice biometrics)

One-to-one authentication against an enrolled model: access logic, not evidentiary logic. Commercial biometrics is not to be confused with forensic examination, and the expert report explains why.

05

Speaker profiling

When there is no suspect, the speech describes the speaker: sex, age range, regional origin, level of education and possible pathologies, on a sociophonetic basis.

06

Forensic transcription

Faithful, standardised transcription: pauses, hesitations, overlaps and unintelligible passages marked, with phonetic hypotheses documented and a degree of reliability declared for each passage. Nothing to do with commercial or automatic transcription.

07

Audio authenticity and integrity

Cuts, insertions and edits leave traces: spectral discontinuity, jumps in background noise, recompression, codec signature. The path runs both ways: proving the editing or documenting the absence of traces, thereby supporting integrity.

08

Synthetic and cloned voice (audio deepfake)

Examination with declared methodology: vocoder artefacts, anomalous prosody, absence of the physiological micro-variations of the human voice (natural jitter and shimmer) and digital traces in the file itself.

09

Complementary examinations

Intelligibility enhancement that never creates content; analysis of non-vocal sound sources (gunshots, impacts, sequence of events) and reconstruction of the audibility of a scene; collection of voice exemplars under a directed protocol; rebuttal of another expert's report, drafting of questions to the expert and party-appointed expert support.

From the trained ear to the spectrogram: the examination of the voice
How the examination is carried out

Hybrid perceptual-acoustic method, the international standard

The examination combines the trained ear and the instrument, following the practice recommended by the IAFPA and by the ENFSI Best Practice Manual, conducted with specialist training in Forensic Speaker Identification (CEFAC).

Auditory-perceptual analysis
  • Voice quality, using Laver's VPA protocol
  • Accent and idiolect: the personal marks of speech
  • Speaking rate and articulation rate
  • Characteristic disfluencies and hesitations
  • Prosodic patterns and lexical choices
Acoustic-instrumental analysis
  • Spectrograms and spectral measurements
  • Formants F1 to F4, in long-term distributions
  • Fundamental frequency (f0): mean, variability and long-term measures
  • Jitter, shimmer and VOT
  • Prosody and rhythm measurements

When the material allows, perceptual-acoustic analysis is complemented by automatic speaker comparison systems, always under human supervision and interpretation: the machine computes, the expert concludes. And every examination goes through quality control: double-checking of findings, documentation that allows reproduction by another expert and declared attention to the mitigation of contextual bias.

The conclusion is expressed as a likelihood ratio: the expert report does not declare "it is the person"; it quantifies how far the evidence supports the same-speaker hypothesis against the different-speaker hypothesis, on a 9-point verbal scale. That is the way of concluding that withstands adversarial scrutiny.

The workflow in phases
ScreeningFeasibility assessment of the material, at no cost, with an answer within 48 business hours
AuthenticityVerification of the integrity and quality of the recording before any comparison
ExemplarsDirected collection of voice exemplars, in person or by video conference
AnalysisAuditory-perceptual and acoustic-instrumental examination of the samples
Expert reportConclusion expressed as a likelihood ratio, on the 9-point scale, with documented method
Court supportTechnical defence of the report in court, clarifications and answers to questions to the expert
What sets it apart

The audio is preserved from the start: phonetics and computer forensics under the same signature

Before the voice examination, the forensic extraction of the original audio straight from the device (iOS and Android), before the application recompresses it, with hashing, binding to the device and documented chain of custody. The examination begins at the source of the evidence, not at the copy that survived being forwarded.

Methodological grounding: IAFPA · ENFSI Best Practice Manual for Forensic Speaker Comparison.

Declared limits

What forensic phonetics does not assert

A declared limitation is a strength, not a weakness: it is what separates science from guesswork, and it is where fragile reports collapse.

A categorical voice report is a fragile report

Naive superposition of spectrograms and absolute certainties do not survive adversarial scrutiny. A technical rebuttal examines the method of the opposing report and points out, step by step, where it does not hold up.

Forensic audio analysis workstation: spectrograms under analysis and recording in an acoustically treated room
AUDIO AND VOICE · Analysis workstationAcoustically treated room
Material required

What the examination needs in order to begin

DurationBelow roughly 10 seconds of net speech there is little speaker information; the desirable range is between 20 and 60 seconds per sample. The screening states frankly what the material allows
QualityReasonable signal-to-noise ratio and minimal overlapping speech; extreme noise may make the examination unfeasible
ChannelTelephone compared with telephone whenever possible: the channel shifts formants. The origin of the recording (application, wiretap, voice recorder, CCTV) must be stated
OriginalForwarding through WhatsApp recompresses the file and erases metadata. Send only via cloud services (Drive, WeTransfer or a secure link), never inside the conversation itself; if unavoidable, "send as document". Better still: extraction of the raw file straight from the device, with chain of custody
Voice exemplarsDirected collection, with spontaneous speech and reading, replicating where possible the conditions of the questioned sample
LawfulnessAn ambient recording made by one of the participants in the conversation is lawful as evidence, under STF (Brazilian Federal Supreme Court) general repercussion Theme 237, RE 583937; wiretaps require the conditions set out in Law 9,296/1996 (the Brazilian Wiretapping Act)
Who it is for

Where voice examination decides cases

Criminal

Wiretaps, recordings and exclusion of a suspect

Integrity of wiretaps and attribution of voices (who is who), covert recordings, threats sent by audio message and exclusion of a suspect: negative proof is also a result, and it has already cleared innocent people of accusation.

Employment

Audio in labour disputes

Harassment and offensive remarks in messaging groups, authenticity of a meeting recording, denial of authorship of audio attributed to an employee or employer.

Family

Custody disputes and family conflicts

Audio in custody disputes and allegations of parental alienation, threats between former spouses, verification of tampering in home recordings.

Corporate

Compliance and voice fraud

CEO fraud using a cloned voice, authenticity of meeting recordings, internal investigations and validation of evidence from the whistleblowing channel.

Frequently asked questions in this area
Is it possible to identify someone by voice alone?
Not as a "fingerprint". The examination expresses degrees of probabilistic support for the competing hypotheses, on the 9-point scale, and that is how the evidence withstands adversarial scrutiny.
How much audio is needed?
A practical rule: less than about 10 seconds of net speech rarely makes the examination viable; the desirable range is between 20 and 60 seconds per sample. The screening, at no cost, answers that for your specific case.
Is WhatsApp audio usable?
It is, with a caveat: forwarding compresses the file and erases metadata. What sets the laboratory apart is the extraction of the raw audio straight from the device, with chain of custody: phonetics and computer forensics under the same signature.
Can it be proved that the audio was edited? And that it was NOT?
Both examinations exist. Editing leaves spectral and compression traces; the documented absence of those traces, under a declared method, supports the integrity of the recording.
What if the voice is disguised, or belongs to a twin?
These are real limits, declared in the report: disguise, a degraded channel and close kinship make the examination harder and may lead to an inconclusive result. A serious report says so in plain words.
Can an AI-cloned voice be detected?
Yes: the examination looks for synthesis artefacts, anomalous prosody and the absence of the physiological micro-variations of the human voice, together with the digital traces in the file itself.
Does an inconclusive result mean the examination was useless?
No: an inconclusive result defines what the evidence supports and frequently brings down an opposing report that was far too categorical. Knowing the limit of the evidence is also a procedural advantage.
Is a privately commissioned report valid in court? Who can produce one?
It is valid as a technical opinion and as party-appointed expert support, under the CPC (Brazilian Code of Civil Procedure). The examination requires specific training: at VALLIM, specialist training in Forensic Speaker Identification (CEFAC) combined with 25 years of computer forensics.
Completed cases
ANONYMISED CASE

Speaker identification in a covert recording within a family dispute

Comparative phonetic analysis concluded against identification, excluding the person as the author of the questioned speech.

Complete multimedia evidence

The voice is one part. The rest of the evidence can be examined too.

Send your audio for a feasibility assessment

The initial screening is free of charge and states frankly what the material allows: duration, quality, channel and the real chances of reaching a conclusion. Answer within 48 business hours.

Request a free assessment
Representative cases

The level of work you are engaging

Before deciding, it is worth seeing what has already come through this laboratory: cases described without identifying the parties, in the format of challenge, method and result.

See all representative cases →

The technical evidence your case requires. The authority courts respect.

Initial feasibility consultation at no cost. Reply within 24h on business days.
Request an Examination