Home // BLOG // FORENSIC PHONETICS

The audio was a montage: authenticity of recordings in the age of AI

From the electoral deepfake to spliced passages in family court: how the examination separates the intact recording from the fabricated one.

Forensic Phonetics · April 22, 2026 · 8 min read

Authenticity of recordings: the waveform under examination

The election was only days away when the audio began circulating through the WhatsApp groups of a small town in the interior of Mato Grosso do Sul. In it, the mayor, running for reelection and a major rural producer in the region, hurled heavy insults at the town's poorest residents. The material was devastating at the worst possible hour: a man of means despising the voters who would decide his future. There was just one problem. That voice never left a throat.

The forensic examination established that the audio was not an edited recording: it was a complete fabrication, a synthetic voice generated by artificial intelligence. The key to the analysis was what was missing. Human speech carries the natural reverberation of the phonatory instrument: the rib cage, the vocal tract, the whole body imprint organic, irregular, living variations on the sound. The synthetic audio under examination exhibited practically uniform frequency, with variations of a deterministic nature, the pattern of a generation process, not of a human being speaking. The report was presented to the Electoral Courts, the material was recognized as false, and the fraud did not change the outcome: the attacked candidate was reelected by a wide margin.

The other kind of fake: the montage by editing

Not every fraudulent audio is born in a voice generator. In a family court dispute, husband and wife presented recordings against each other: she filed audio of his insults; he filed audio in which she attacked him even more severely. Each denied the authenticity of the other's material. The examination answered both. The audio presented by the wife was intact. The audio presented by the husband did, in fact, contain the wife's voice, but it was a collage: passages recorded at different moments, reassembled in the order that served his narrative. The sequence had no internal logic, and the analysis identified the cut points and restructured part of the material based on the discontinuities of the background noise, which changes from one environment to another and betrays the splice. The one who forged the evidence answered for bad-faith litigation (CPC, arts. 79 to 81, Brazilian Code of Civil Procedure).

"Human speech reverberates in the body that produces it. The synthetic voice has the regularity of a machine. And the montage, however good, leaves the splice where the background noise changes."

What the authenticity examination looks for

Unlike speaker comparison, which asks who spoke, the authenticity examination asks whether the recording is intact and original. The work moves through layers: the continuity of the signal (cuts, splices, abrupt spectral transitions), the coherence of the acoustic environment (reverberation, background noise, microphone distance, constant or not), the traces of recompression and reprocessing of the file, the container's metadata and, at the current frontier, the patterns that separate organic speech from synthetic speech. Each finding is documented reproducibly: another examiner, retracing the path, reaches the same findings.

Time testifies as well. The file's metadata records creation and modification dates that must square with the story told in the record: an audio supposedly recorded in January, in a file created in March, has some explaining to do. And the format version, the generating application, and the chain of recompressions tell how many tools the material passed through before becoming "evidence".

The forwarded audio loses its anchor

A good share of the audio files that become evidence arrives by forwarding: someone received it, re-forwarded it, saved it, and filed it. That material can still be analyzed as audio, but it loses what matters most in court: the anchor of its origin and the chain of custody. Where did it come from, through how many hands did it pass, what did each resend recompress and erase? The opposing party can question exactly that, and defeat the use of the material through the impossibility of demonstrating its provenance.

The mature answer unites two forensic practices. Computer forensics extracts and preserves the evidence on the originating device, with hash values and a documented chain of custody: the audio ceases to be a loose file and becomes an object traceable to its source. Over that preserved material, forensic phonetics conducts the authenticity examination and, when the case requires, the comparison with the voice reference of the person to whom the speech is attributed. One practice demonstrates that the audio is that audio; the other, what the audio is.

From recording to paper: the transcript with forensic value

Once the recording is authenticated, one step remains that tends to be underestimated: turning it into text usable in the record. Forensic transcription is not the commercial transcript any automated service delivers. It is faithful and standardized: it registers pauses, hesitations, overlapping voices, and simultaneous speech; it marks the unintelligible passages instead of inventing what "seems" to have been said; and it documents the reading hypotheses with the degree of reliability of each passage. The difference matters because automatic transcription fails exactly where the case is decided: in the whispered name, the amount spoken over background noise, the sentence cut short. The text that goes into the case file must carry the same honesty as the examination that sustains it: say what is heard, declare what is not heard, and never fill gaps with supposition.

What to preserve when the audio matters

  • The device that originally received or recorded the audio, without deleting the conversation;
  • The complete conversation in the app, not just the exported file;
  • No resending: forwarding recompresses the file and discards metadata;
  • A record of who sent it, when, and in what context;
  • Engagement of the forensic expert before any processing or home-made "enhancement" of the sound.

The age of artificial intelligence did not invent fake audio; it industrialized it. The handcrafted montage of the family court case and the complete fabrication of the electoral case are two ends of the same problem, and the answer to both is the same: authenticity is not decided by ear, it is decided by examination. Between the audio that circulates and the audio that proves, there is a laboratory along the way.

VALLIM

Adriano Vallim

Forensic expert specializing in digital crimes, working across computer forensics, handwriting and document examination, and forensic phonetics. He combines technical, academic and institutional credentials that place him among the most complete references in the field in Brazil. See the full background →

Read next

The technical evidence your case requires. The authority courts respect.

Initial feasibility consultation at no cost. Reply within 24h on business days.
Request an Examination