What we will publish here
There are no posts yet. Rather than fill this page with placeholders, here is what the blog is for, what it will not carry, and the pieces we are positioned to write first.
Nothing published yet. When the first post goes up it will appear here with a date and a named author. Until then this page is a statement of intent, not a catalogue.
What it publishes
Measurement results
Rates we have actually measured, with the dataset described and the date attached — including the ones that moved the wrong way after a retrain. A number that only ever improves is a marketing number.
Method changes
What changed in the engine, why, and what it did to the error rates in both directions. Each post links the model version it describes, so an old result stays readable against the method that produced it.
Analysis of new generators
When a synthesis system ships, what it sounds like, what it does to our detection rate, and how long our attribution holds up against it. Dated, because the answer changes.
What it does not publish
Trend pieces
No commentary on where the industry is heading. If we have not measured it, we have no more standing to write it than anyone else does.
Fear marketing
Voice fraud is a real harm and it does not need help from us to sound alarming. Posts describe what happened and what to do about it, at the size it actually is.
Accuracy claims without a dataset
No figure appears here without the set it was measured on, how that set was built, and when. That rule applies to our numbers and to anyone else’s we quote.
The first posts we intend to write
These are drawn from work already on this site. None of them exists yet. Titles will change; the subjects probably will not.
The false positive rate nobody publishes
Why detectors advertise the catch rate and never the rate at which real speech gets flagged, what the trade-off between them actually costs, and what ours is.
Why re-recorded audio defeats detection
Play a generated clip through a speaker, capture it on a phone, and the file acquires a genuine recording chain. This is our hardest open problem and we do not have a good answer to it.
What “unclear” means, and why it is a result
A third verdict is not a failure to answer. It is a statement about the recording rather than the speaker — and the reason we do not bill for it.
How attribution decays after a vendor ships
Naming the generator degrades within weeks of a new model release, while the synthetic-or-human verdict holds up far longer. What that means for how much weight to put on the generator field.
Why phone audio is the hardest case
Narrow-band codecs discard the detail detection depends on, and most published accuracy figures were never measured on it. What we changed to work on the audio people actually receive.
What an evaluation dataset has to contain to be honest
Held out from training, balanced across recording conditions rather than only across labels, and stocked with real human speech carrying real defects. Anything less produces a number that cannot be argued with.
Real-time voice conversion is a harder problem than text to speech
A live person speaking through a conversion model carries real prosody and real timing. Our rates on it are lower, and this post is where we would say by how much.
What a retrain did to our error rates
A before-and-after on a single engine version, both directions, by audio condition — written the same way whether the result improved or got worse.
When there is something to read
[VERIFY: newsletter signup, or delete this section. If you keep it, say how often it sends, what it contains, and that it is only sent when there is a result worth sending — then honour that.]
Corrections to anything published here go to corrections@aivoicedetctor.com. A correction is appended and dated, never edited in silently.
Reviewed