AI FOR TEACHERS8 min read

AI grading: what to automate, and what you should never delegate

LS
Laxman Shah
building TeachRoster ·

Every conversation about AI grading collapses into the same two camps. One says marking is drudgery and should be automated to zero. The other says marking is teaching and must never be touched. Both are describing something real, and both are describing a different half of the pile on your desk.

The useful question is not whether to automate grading. It is which specific marking task you are talking about, because the honest answer changes completely between them.

Tier one: automate it completely

Multiple choice. True or false. Fill in the blank with a single accepted answer. Matching. Anything where a correct response can be written down in advance.

There is no argument for marking these by hand, and the argument is not about time. A machine is more consistent than you are. It does not get more generous at question 12 of 35, it does not get harsher after a bad meeting, and it does not skip a line and mark the wrong row. Teachers are excellent at judgement and mediocre at repetitive comparison, particularly at nine in the evening.

this tier is not “AI” in any meaningful sense — it is string comparison with better marketing

Tier two: let it suggest, you decide

Short answers. One-paragraph explanations. Show-your-reasoning questions. Here a model can genuinely help, and here it is also confidently wrong often enough that unsupervised marking is indefensible.

What makes this tier workable is a confidence level attached to every suggestion. It tells you where to spend the attention you have.

A short answer in the TeachRoster grading inbox with the student's answer, an AI suggestion of 3 out of 3 marked high confidence, a one-line reason, and editable feedback with accept or score-it-yourself buttons
The suggestion, the confidence, the reason, and the buttons to disagree with all three.

In practice that means: accept the high-confidence suggestions on answers that are plainly right or plainly blank, and read every low-confidence one properly. That is not a compromise, it is triage, and it is what you already do when you flick through a stack looking for the ones that need thought.

The failure mode nobody advertises

Automated marking is not randomly wrong. It is wrong in a specific, predictable direction: it penalises the student who solved the problem a different way.

The student who reasons about the ratio instead of applying the procedure you taught. The student whose English is a third language and whose phrasing is unusual but whose maths is perfect. The student who skips three steps because they can. These answers look further from the expected answer, and they get marked down.

The students most likely to be marked unfairly by a machine are the interesting ones.

Which is precisely why the low-confidence pile is where you should spend your time, and why a system that hides the confidence from you is worse than one that shows it.

Tier three: never hand this over

Some marking is not a measurement task at all. Automating it does not save time, it deletes the thing being done.

  • Work where the comment is the point. A first draft of a piece of writing. The grade is nearly irrelevant; what the student does next depends entirely on which two sentences you chose to respond to.
  • Anything measuring effort or progress. A model sees this submission. It does not know that this student has produced nothing for three weeks and that this paragraph is a breakthrough.
  • High-stakes final marks. Reports, certifications, anything a student or a parent can appeal. Not because the model is necessarily worse, but because you must be able to explain and defend every one of those numbers.
  • The student you are worried about. If a name in the queue makes you pause, that is you noticing something. Read it yourself.

A working rule

Automate the marking that is comparison. Assist the marking that is judgement, and keep the final say. Do the marking that is a conversation yourself, on purpose, and protect the time to do it by having got the first two tiers off your desk.

The point of automating tier one is not to spend less time on marking. It is to spend the same time on the twelve pieces of work where your attention changes something.

What to demand of any tool that does this

  • Tell you what was auto-graded and what was suggested. Those are not the same event.
  • Show a confidence level, and make the low-confidence work easy to find.
  • Let you override any score in one action, without a dialogue arguing with you.
  • Never send anything to a student until you press return.
  • Be clear that student work is not used to train public models.

If a tool cannot answer those five, it is not asking to save you time. It is asking you to trust it, which is a different request entirely.

More from the blog