Marking 120 Essays in a Weekend: The Honest AI Workflow
A step-by-step marking workflow that cuts a 120-essay pile from eleven hours to about four - including the parts of marking you must never hand over.
A pile of 120 essays takes most secondary teachers between nine and twelve hours to mark properly. I have got mine down to roughly four, and the feedback pupils receive is measurably more specific than it was before. What follows is the exact sequence, including the parts where I stopped using AI because it made the work worse.
Be clear about the claim: this is not automated marking. Every grade in my workflow is awarded by me. What changes is where my hours go - away from retyping the same six comments and towards the decisions only a teacher can make.
Before you touch a device
Two preparation steps do most of the work.
Write your rubric as if a stranger had to apply it. Vague criteria such as good analysis produce vague feedback whether a human or a machine drafts it. Replace them with observable evidence: makes a claim, supports it with a named example, explains how the example supports the claim. This single rewrite improved my feedback quality more than any tool did.
Mark ten scripts by hand first. Choose a spread: two strong, two weak, six middling. You are calibrating yourself and simultaneously producing the anchor examples the rest of the workflow depends on. Skip this and everything downstream drifts.
Step 1: Build a comment bank from your own marking
Take the ten hand-marked scripts and paste the comments you actually wrote into a document. Then ask an assistant to do something narrow:
Here are 38 feedback comments I wrote while marking Year 10 history essays.
Group them into themes. For each theme, give me:
- the shortest version of the comment that keeps the specificity
- one "next step" sentence phrased as an action the pupil can take tomorrow
Do not invent comments I did not write.
That last line matters. Unconstrained, you get generic edu-speak. Constrained to your own words, you get a reusable bank in your voice. Mine collapsed to nine themes and lives in a document I have reused for three years.
Step 2: Batch by rubric criterion, not by pupil
This is the change that saved the most time, and it has nothing to do with AI. Instead of reading each essay once and judging six criteria simultaneously, read the whole pile for one criterion at a time.
- Pass one: thesis and structure only.
- Pass two: use of evidence.
- Pass three: analysis and conclusion.
Three fast focused passes beat one slow anxious pass. Your standards stay far more consistent across the pile, which is the thing moderation meetings always catch.
Step 3: Use AI for the description, never the judgement
For each essay I paste the text with a tightly bounded request:
Rubric criterion: "Uses at least two specific pieces of evidence and explains
how each supports the argument."
Essay below. Do three things only:
1. Quote every piece of evidence the pupil used.
2. For each, say whether an explanation follows it in the same paragraph.
3. Do not award a grade, level or score. Do not praise the essay.
Essay: [...]
What comes back is an evidence inventory. I read it in about fifteen seconds, glance at the essay to confirm, and award the level myself. The inventory catches things tired eyes miss at essay ninety - particularly evidence that is present but unexplained, which is the most common reason a script sits one band lower than the pupil expected.
Step 4: Draft the comment, then rewrite the first sentence
I select the two most relevant comment-bank themes and ask for a short paragraph addressed to the pupil. Then I always rewrite the opening sentence by hand, naming something specific from their script.
Your paragraph on rationing was the strongest - the ration-book detail did real work. No system can produce that sentence, because it requires having read this pupil's essay and remembered what the class did in October. It is also the sentence pupils read most carefully.
Step 5: The five checks before anything goes out
Non-negotiable, and they take about twenty minutes for the whole pile.
- Grade check. Every level was set by me, in my own pass, before I read any drafted comment.
- Name check. No feedback names the wrong pupil or the wrong task. Batch errors happen.
- Tone check. Nothing sarcastic, nothing that could be read as a comment on the pupil rather than the work.
- Accuracy check. Any subject claim in the feedback is one I would defend in a parents evening.
- Outlier check. I reread in full every script where my level and the evidence inventory disagreed. That is usually three or four essays, and it is always time well spent.
Where I stopped using AI
Borderline scripts. Anything on a grade boundary gets full human reading, twice, ideally with a colleague. The cost of being wrong is a pupil's target grade.
Anything emotionally loaded. A creative piece about a bereavement, a personal statement, a pupil who has just come back from a long absence. Drafted feedback on these reads as hollow, and pupils notice.
First drafts in a redrafting cycle. These need me to know what the pupil did last time. Context I hold in my head is the whole value of the comment.
Anything I would not show the pupil. If I would be uncomfortable saying a tool helped me draft this feedback, and I checked and edited all of it, the tool should not be in that part of the process. I say exactly that to my classes, once, at the start of the year.
What the time actually looks like
| Stage | Before | After |
|---|---|---|
| Calibration and comment bank | 0 | 75 min (once per task type) |
| Reading and levelling 120 scripts | 8 hr | 2 hr 30 (three focused passes) |
| Writing feedback | 3 hr | 45 min |
| Checks | 0 | 20 min |
The comment bank is a one-off cost that pays back on every future set. The honest total for a repeat task is a little under four hours.
FAQ
Is this allowed under my school's policy? Usually yes, because pupil grades remain teacher-awarded, but check two things specifically: whether pupil work may be pasted into an external tool at all, and whether your data protection officer requires names removed first. Strip names by default.
Does feedback quality drop? In our department it rose, for an unglamorous reason: there was enough time left to write a specific opening sentence for every pupil. Previously the last thirty essays got three words.
What about handwritten scripts? Photograph them and use the tool only for the evidence inventory. Transcription errors make anything more ambitious unreliable.
How do I convince a sceptical head of department? Offer to run it on one set alongside their normal marking and compare a sample of ten for consistency and specificity. The batching-by-criterion change alone usually wins the argument, and it involves no technology at all.