The institution sets the policy. The student pays for the mistakes.
Before this goes anywhere near a misconduct procedure, be clear about the distribution of harm. A missed case of cheating costs a mark. A false accusation costs a student a year, a visa, a reference, or a degree they earned — and there is nothing they can produce to answer it.
Where audio has become assessable
Spoken work is no longer only the language department’s problem. Recorded oral submissions, podcast and presentation assignments, remote vivas, spoken components of language qualifications, recorded reflective pieces on placement, and admissions interviews conducted over a video platform — all of it now sits in an environment where a competent synthetic voice takes minutes to produce and costs nothing.
Institutions have been here before with written work, and the lesson from that round is the relevant one. Detection tools were adopted quickly, treated as more certain than they were, and the resulting false accusations fell hardest on students who were already least able to contest them: international students, students writing in a second language, students without an academic family to tell them to appeal. Whatever you do here, do it having learned that.
What this cannot do for you
It cannot establish that a student cheated. It analyses a recording. It has no view on who submitted it, who was in the room, or whether the work is the student’s own.
It cannot identify a speaker. No matching against a known sample of the student’s voice. A recording of somebody else reading the student’s script is not a case we can detect at all.
It cannot verify authenticity. Likely human means no known signature was found. It is the weaker verdict and must not be recorded as a clearance.
It cannot be relied on for the audio students actually submit. Phone microphones, shared kitchens, laptop noise suppression, a file compressed by the platform you asked them to upload to, and speech in a second language all reduce reliability. [VERIFY: state measured performance by language, or say plainly that it is untested]
It cannot be answered by the student. This is the one that should shape your policy. There is no document, no metadata, no witness and no re-recording that proves a voice recording is genuine. A procedure that presents a score and invites the student to rebut it has invited them to do something impossible, and the panel will read their failure to do it as consistent with guilt.
It cannot make the finding. A named human decision-maker weighs the analysis against everything else and is accountable for the outcome.
The asymmetry, in the terms your registrar will recognise
A missed case of misconduct is a mark awarded that should not have been. It is bad, it is bounded, and the institution absorbs it.
A false accusation is a student in a formal procedure they cannot win by producing evidence. Depending on their circumstances, the consequences run from a capped mark to a withdrawn offer, a lost placement, a compromised student visa, or an expulsion attached to their name permanently. They will also carry the memory of an institution telling them a machine said they were dishonest. The rate at which that happens is not zero, and every point of sensitivity you gain increases it.
The practical conclusion is not that the tool is useless. It is that the threshold for look into this and the threshold for allege misconduct must be different numbers, and the second one should be high enough that no case ever rests on the score.
Design that reduces the need for detection
Most of the value here is in assessment design rather than analysis, and academics will find this the more familiar argument.
Make it interactive
A short live viva on submitted work — two unscripted questions — distinguishes understanding from a generated recording more reliably than any model, and it is defensible because a human made the judgement.
Assess the process
Require a draft, an outline, or a recorded working note alongside the final submission. Fabricating a plausible development history is far harder than fabricating an output.
Anchor to context
Ask for reference to something only that cohort experienced: a seminar discussion, a set text read that week, a site visit. Generation systems have no access to it.
If you use it anyway, use it like this
- Publish the policy before the assessment. Students should know that submissions may be analysed, by what, and what a result does and does not mean. Retrospective use of a tool nobody was told about is the fastest route to a successful appeal.
- Never let a score be the allegation. Open a case on academic grounds — a viva that does not match the submission, an inconsistency in the work — and let the analysis sit inside that case as one item.
- Give the student everything. The report, the engine version, the confidence, the audio conditions and the stated error rate. Withholding the technical detail from the person it is used against is indefensible.
- Record the audio quality. Where the recording was poor, that fact belongs in the panel papers next to the verdict, because it is usually the explanation.
- Track outcomes by group. If flags cluster by language, nationality or fee status, you have a problem with the tool, not with those students.
- Give a route to re-submit. Offering a supervised re-recording is a far better remedy than a hearing, and it resolves most genuine cases in an afternoon.
Academic misconduct procedures, data protection duties towards students and rules on automated decision-making differ by jurisdiction and by institution type. [VERIFY: verify with your own legal and academic governance teams] Nothing here is legal advice.
Questions from academic registries
Can we screen every submitted recording?
You can, and it is the deployment we would most warn against. At cohort scale, even a small false positive rate produces a steady flow of accusations against students who cannot answer them, and the administrative pressure will be to treat each flag as meaningful because it arrived.
What about admissions interviews?
Applicants are the most vulnerable group here: no relationship with the institution, no appeal route worth the name, and often the worst connections. Any use in admissions needs disclosure at application and a human review of every negative signal, and we would not run it as a screen.
A student says our tool is wrong. What now?
Assume it might be. Offer a supervised re-recording, check the audio conditions on the original, and look for academic evidence independent of the audio. If the case stands only on the verdict, it does not stand.
Should students be allowed to check their own audio first?
It is a reasonable thing to permit, and it makes the tool a shared instrument rather than an instrument used on them. Be clear that a pre-check is not a guarantee, since a later analysis on a differently compressed copy can differ.
Read the limitations before the policy meeting
If one paper goes to the committee, make it that one rather than this one.
Reviewed