Teaching students to work with AI
How to tell if a student used AI
Detectors don’t work well enough to accuse anyone. Here’s what the false positive rates actually mean for your class, and what to do instead of guessing.
Start with the arithmetic, because it’s the part that changes minds.
Suppose a detector is 99% accurate - far better than anything on the market claims in independent testing. You run it across 120 submissions in a large section. Even at that accuracy, you should expect roughly one innocent student to be flagged. Not because the tool is broken, but because that’s what a 1% false positive rate means when you apply it 120 times.
Now consider what you do with that flag. You cannot tell which of your flagged students is the innocent one. You are holding a number, not evidence.
A detector output is a probability estimate about a population. An academic integrity case is a claim about one person. The first cannot substantiate the second.
It’s worse than the arithmetic suggests
Real-world accuracy is lower than the marketing numbers, and the errors aren’t randomly distributed:
- Non-native English writers are flagged far more often. Detectors key on low “perplexity” - text that is grammatically plain and predictable. That describes a lot of careful second-language academic writing. Several institutions have disabled detection tools for precisely this reason.
- Neurodivergent and formulaic writers get flagged. A student taught to write in rigid five-paragraph structure produces exactly the statistical signature these tools look for.
- Light editing defeats them. Paraphrasing tools, a few manual rewrites, or asking the model to “write like a tired sophomore” drops scores substantially. The students most likely to be caught are the ones who did the least to hide it, which inverts what you want from an enforcement mechanism.
- They can’t show their work. When a student appeals, you have a percentage and no explanation. That is not a position you want to defend in front of a committee.
What to do instead
1. Ask them to walk you through it
A ten-minute conversation is more diagnostic than any tool. Ask why they structured the argument that way, what they cut, which source they found least convincing and why. A student who did the work answers easily and often enjoys it. A student who didn’t usually says so within a few minutes.
This scales worse than a detector and works far better. Reserve it for the cases that actually matter.
2. Assess the process, not just the artifact
If the submission is a finished file, you’re inspecting the one thing AI produces perfectly. If the assignment captures decisions along the way - what the student chose, what they rejected, how they responded when challenged - you’re looking at something a chatbot can’t hand over in one turn.
This is the structural fix, and the only one that scales. AI-resistant assignment design goes into the mechanics.
3. Build in a moment of live thinking
One in-class paragraph, one short oral defence, one question that can only be answered by someone who was in the room for the week-six discussion. You don’t need to proctor everything. You need one anchor point per unit where you see the student think.
4. Make disclosure easy and non-punitive
Most AI use in your course isn’t the dramatic kind. It’s a student who asked for an outline and isn’t sure whether that counts. Give them a low-friction way to say so, and you’ll learn more about actual usage patterns in one term than detection would tell you in five. Policy language for this.
Signals worth noticing (and their limits)
None of these is proof. All of them are reasons to start a conversation rather than a case.
- Citations that don’t exist. Still the most reliable signal. Models invent plausible references. Check two.
- Confident wrongness about your specific course. A polished paragraph that misstates the framework you spent two weeks on is more telling than any style marker.
- A sudden change in voice. Useful only if you have earlier writing from the same student, which is an argument for collecting a short in-class sample in week one.
- Generic examples where specific ones were required. If the assignment asked about the case you assigned and the answer could be about any case, something is off.
- Absence of the messy bits. Real student thinking has hedges, false starts, and a paragraph that doesn’t quite land. Uniform competence across a whole submission is unusual.
If you do decide to raise it
Lead with curiosity, not accusation: “I’d like to talk through your process on this one.” Document what the student says. Never open with a detector score - it puts you on ground you can’t hold, and it turns a teaching conversation into an adversarial one before you’ve learned anything.
Common questions
Are any detectors reliable?
Not reliably enough to accuse an individual. They can be mildly informative in aggregate - if a third of a section trips the threshold, something about the assignment is worth rethinking. That’s a signal about your assignment design, not about any particular student.
What about tools that track typing or revision history?
More informative than text analysis, because they capture process rather than inferring it. They also raise real surveillance concerns and can be defeated by retyping. Useful as one signal among several, not as a verdict.
My institution requires me to use a detector. Now what?
Use it as a triage flag that prompts a conversation, never as evidence in itself. Document the conversation, not the score. And push, where you can, for the policy to say the same thing - many institutions have revised theirs after seeing the false positive data.
Isn’t this just giving up on integrity?
The opposite. Detection was never the integrity mechanism - it’s a recent and poor substitute for one. The durable mechanisms are assessment design, a clear policy, and knowing your students’ work well enough to notice when something doesn’t fit.
See what this looks like in a real assignment
Kova turns a sentence into immersive, auto-graded coursework - AI conversations students have to argue their way through, models they build and defend, transcripts scored on the reasoning rather than the artifact.
The full guideHow Conjure builds itRequest a demo