Teaching students to work with AI
How to grade students’ use of AI
If you allow AI, you need a way to grade how well it was used. Here’s a six-criterion rubric, what each level looks like in practice, and how to collect evidence without reading every transcript.
Allowing AI creates a grading problem that a normal rubric doesn’t solve. Two students submit comparable work. One interrogated the model, caught a bad assumption, and rebuilt the analysis. The other accepted the first output and reformatted it. Your rubric scores them the same, because it was built to evaluate an artifact.
What follows is a rubric for the other thing - the quality of the student’s reasoning while working with the tool. It works alongside your existing criteria, not instead of them.
Six criteria
1. Framing the problem
Did the student give the model enough to work with? A vague prompt produces generic output, and a student who accepts generic output usually never noticed the problem was theirs.
- Strong: supplies context, constraints, and what a good answer would need to contain.
- Developing: asks a reasonable question but omits the constraints that matter.
- Weak: pastes the assignment prompt verbatim.
2. Verification
Did the student check anything? This is the single highest-value habit to grade, because it’s the one that transfers to every context they’ll work in later.
- Strong: checks specific claims against a source, recomputes a number, notices a citation that doesn’t exist.
- Developing: checks the conclusion but not the steps that produced it.
- Weak: no evidence of checking anything.
3. Catching the error
Models make confident mistakes. Did the student find the one in their output? If you seed the assignment with a scenario where the model reliably goes wrong, this becomes easy to assess and very informative.
4. Pushing back
Did the student challenge an answer, or accept the first one? Look for follow-ups that test the reasoning - “what would make this wrong,” “you assumed X, what if not.”
- Strong: challenges a specific claim and changes position based on the answer, or holds position with a reason.
- Developing: asks follow-ups that request more detail rather than testing the logic.
- Weak: single exchange, no follow-up.
5. Judgment about what to keep
What did the student do with what they got? Wholesale adoption and wholesale rejection are both failures of judgment. You want evidence of selection - this part was useful, this part was wrong, this part I rewrote because it missed the point of the assignment.
6. Honest accounting
Does the disclosure match what actually happened? A student who says “I used it to check my structure” and whose transcript shows they generated the whole thing has committed a much clearer violation than an ambiguous one about AI use generally. This criterion is what makes the others enforceable.
Weighting it
In most courses these six should carry less weight than your subject-matter criteria - somewhere between 15% and 30% of the assignment. The point isn’t to make the course about AI. It’s to make sure the part of the work that AI touched is still being assessed on something.
Collecting the evidence
The rubric is straightforward. The hard part is seeing enough to apply it, without committing to reading 90 chat transcripts.
The disclosure note
Cheapest option and surprisingly effective. Three sentences at the end of the submission: what tool, what you asked for, what you changed. Grades criteria 5 and 6 directly, and gives you a starting point for the rest.
The annotated excerpt
Ask for one exchange the student found useful and one they found wrong, with a sentence on each. Two short excerpts instead of a whole transcript. This surfaces criteria 2, 3 and 4 with maybe ninety seconds of reading per submission.
The reflection question
One question on the assignment: where was the model most confidently wrong, and how did you know? A student who can answer it did the work. A student who can’t will say something vague, which is itself the answer.
The captured conversation
If the AI conversation happens inside your course platform, the transcript arrives with the submission and can be scored automatically against the criteria above. That’s the version that scales to large sections - how automated transcript grading works.
Whichever you pick, tell students the rubric in advance. The rubric is the teaching. A student who knows they’ll be graded on catching the model’s errors goes looking for errors, which is the entire point.
Three pitfalls
- Rewarding volume of interaction. Twenty shallow exchanges are worse than three sharp ones. Grade the quality of the questions, not the count.
- Penalising efficient use. A student who got what they needed in one well-constructed prompt did better than one who took fifteen turns to get there. Criterion 1 should reward this, not punish it.
- Grading the transcript instead of the work. The artifact still matters. This rubric supplements your subject criteria; it doesn’t replace them.
Common questions
Isn’t this a lot of extra grading?
It is if you ask for full transcripts. It isn’t if you ask for a three-sentence disclosure and one annotated excerpt. Start with the cheap version and add depth only where the course warrants it.
What if students fabricate the disclosure?
Some will. But a fabricated disclosure is a clean, specific violation - much easier to act on than a vague suspicion about AI use in general. You’ve converted an unenforceable rule into an enforceable one.
Does this work for group work?
Yes, and it gets better. Ask the group to record one disagreement about whether to accept an AI output and how they resolved it. That single artifact tells you more about the group’s reasoning than the deliverable usually does.
How do I grade this in a class of 200?
Disclosure note plus one reflection question, scored on a three-point scale. Anything requiring you to read a transcript won’t survive at that size unless the platform scores it for you first.
See what this looks like in a real assignment
Kova turns a sentence into immersive, auto-graded coursework - AI conversations students have to argue their way through, models they build and defend, transcripts scored on the reasoning rather than the artifact.
The full guideHow Conjure builds itRequest a demo