Getting a usable recording
Most clips that come back unclear were analysable when they were made and were damaged on the way. The handling matters more than the equipment.
The one-line version. Keep the original, move it without re-encoding it, change nothing about it, and send the longest continuous stretch of one person speaking. Everything below is that rule with the reasons attached.
Why handling matters this much
Evidence of synthesis lives in fine detail: the exact shape of the spectrum, small discontinuities where a system reconstructed part of the signal, the way a background behaves. Every process that makes a file smaller or cleaner works by deciding which detail is not worth keeping — and it is broadly the same detail. A codec is not being careless; it is doing precisely its job, which happens to be destroying the thing we measure.
This is also why likely human is the weaker verdict. A clip stripped of its evidence looks exactly like a clip that never had any. Reading your result covers that asymmetry.
Six steps
Keep the original
Before anything else, save the file as it arrived. Export the voicemail from your carrier’s app, save the voice note out of the conversation, copy the recording off the device. Keep that copy untouched and work from duplicates.
Do not forward it through a messaging app
Sending audio through a chat app re-encodes it on the way out, often to a smaller format than it arrived in. Forwarding it twice compounds that. Move the file by a route that copies bytes: a cable transfer, a cloud drive upload, or an email attachment.
Do not clean it up
Noise reduction, normalisation, volume boosting and enhancement all remove or reshape exactly what is being measured. Worse, some of them leave artefacts of their own that a detector has to account for. Submit the file rough. It is meant to be rough.
Do not re-record it through a speaker
Playing the clip aloud and capturing it on a second phone is the most common way a usable file becomes an unusable one. It adds a room, a microphone and another codec on top of what was already there. If the audio is trapped in an app that will not export, say so rather than working around it this way.
Send the longest continuous stretch, not the worst sentence
The instinct is to trim to the damning line. Resist it. Analysis needs continuous speech from one person; a three-second extract carries far less than twenty seconds of the same speaker saying something unremarkable. Ten to thirty seconds is the sweet spot. If you must trim, cut at silences and keep the passage whole.
Write down where it came from
Note when it arrived, on what number or account, who else received it, and every step you took with the file. Provenance frequently settles a question that the audio alone cannot, and it is the part that becomes impossible to reconstruct a week later.
If you are recording something deliberately
When you have the chance to capture rather than salvage, record uncompressed or near-uncompressed if the device allows it, keep the microphone close to the source, and avoid speakerphone. A phone recording an incoming call from its own line will always beat a phone recording a speaker across a table.
Recording law differs by jurisdiction, and some places require every party to consent. Check your position before you record. [VERIFY: verify the position in your markets]. An unlawfully made recording is usually worth nothing in the proceeding you wanted it for.
The audio file inspector will tell you the duration, sample rate, channel count and size of a file entirely in your own browser, which is a quick way to see whether a clip is worth submitting at all.
Questions people ask
How long should the clip be?
Ten to thirty seconds of one person speaking. Below about five seconds the analysis will usually return unclear rather than guess.
My only copy came through a chat app. Is it worthless?
Not worthless, but weakened. Submit it, and read the source-quality row on the result carefully — it will tell you how much the file had left.
Which format is best?
Whatever the original is. Converting to a “better” format after the fact adds a step without recovering anything. WAV, MP3, M4A, OGG and WEBM are all accepted, up to 25 MB.
Should I trim silence from the start and end?
You can, but there is little to gain and something to lose — room tone at the edges is itself informative. Leave it if you are unsure.
[VERIFY: written by — real name, real role]. Reviewed [VERIFY: date]. Corrections: method@aivoicedetctor.com.
Reviewed