Most generative AI tool selections in education start with a feature grid — chat tutor, quiz generator, essay feedback, lesson planner — and end with whichever product has the most ticks. That comparison assumes the features are deterministic courseware, which they are not. A generative model produces a plausible next token, and a plausible next token in a maths worked solution can be a wrong step delivered with complete confidence. The feature exists; the guarantee does not.
So the useful evaluation question is not “what can this tool do?” but “what does it look like when this tool is wrong, and would we notice before a learner did?” That reframing changes the entire procurement exercise, because it makes measurement — not feature coverage — the thing you are buying.
We should be direct about the limits of our own position here. TechnoLynx has not delivered a generative AI deployment inside a school or university, so nothing below is an outcome claim about education results. What we can contribute is evaluation method: how to structure a pilot, what to log while it runs, and which vendor statements are substantiable in principle. That method transfers from other domains where probabilistic output meets a workflow that assumed correctness.
What does generative AI actually do well in an education setting, and where does it fail?
The failure modes cluster into three shapes, and they are not equally severe.
The first is confident factual error. A model asked to generate ten questions on stoichiometry will produce ten well-formed questions, some fraction of which have wrong answer keys. Nothing in the output signals which ones. This is the failure mode that scales worst, because a wrong answer key propagates to every learner in the cohort simultaneously.
The second is register and reading-level mismatch. Ask for material pitched at fourteen-year-olds and you will frequently get prose that is syntactically fine and lexically two years too hard. Unlike factual error, this is usually visible on a skim — which makes it cheap to catch and therefore lower-risk.
The third is inconsistency across a cohort. The same rubric applied to thirty essays by the same model does not produce thirty comparably calibrated marks, because the model has no persistent standard between calls — only whatever the prompt and the essay itself supply. In a school this shows up as unfairness complaints; in higher education, where marking consistency is an appeal-relevant property, it is a governance problem. Our colleagues working on assessment-integrity questions treat this as the harder half of the problem, and it is developed further in AI in higher education: what it means in practice.
What generative models do genuinely well is first-draft production against a specification a human already holds — a lesson-plan skeleton, a set of paraphrases at different reading levels, a bank of distractors for multiple-choice items, a rewrite of an explanation in a different framing. Every one of those is a case where the human retains the correctness judgement and the model absorbs the typing. The capability and limitation framing behind this — why generation is strong on form and weak on ground truth — belongs to the broader generative AI practice rather than to education specifically.
What a term-long pilot should measure
A pilot that reports usage counts has told you that people clicked. It has not told you whether the output was good, and it cannot support an accept/reject decision. Three measurements can.
Factual-error rate on a held-out question set. Build a set of items — 100 is workable — where the correct answer is already known and the model has not seen your key. Generate against it, mark the output, and record the error rate per subject and per task type. This is the only number that distinguishes tools on the axis that actually matters. Expect it to vary sharply by subject: symbolic and multi-step reasoning tasks fail at a different rate than summarisation or paraphrase.
Teacher review time per generated item. Measured in minutes, logged by the reviewing teacher, separated into “accepted as-is”, “edited”, and “discarded”. If review time per item plus editing time exceeds authoring time from scratch, the tool is a cost with a nicer interface.
Proportion of generated material accepted without edit. The clean acceptance rate, tracked weekly across the term. A flat or falling curve tells you the tool is not adapting to your context; a rising one tells you your staff are learning to prompt it, which is a real gain but one that lives in your people rather than in the licence.
From those three you get the figure worth comparing across vendors: cost per accepted artefact, not cost per API call or per seat. A tool with a low token price and a 40% acceptance rate is more expensive than it looks.
Pilot instrumentation checklist
| What to log | Granularity | Why it matters |
|---|---|---|
| Prompt and full model output | Every call | Without the input you cannot reproduce or diagnose a bad output |
| Model and version identifier | Every call | Vendors update models silently; a quality shift with no version record is uninvestigable |
| Reviewer verdict (accept / edit / discard) | Every generated item | Source of the acceptance rate |
| Review time in minutes | Every generated item | Source of the review-cost figure |
| Error type when discarded | Every discard | Separates factual error from register mismatch from formatting |
| Held-out set score | Weekly or per model change | Detects regression after a vendor update |
| Learner-facing exposure flag | Every item | Tells you which outputs reached students unreviewed |
| Data residency and retention setting | Once, plus on any change | Evidence for your data-protection assessment |
Two operational notes. Log the model version even when the vendor abstracts it away — if they cannot tell you what version served a given response, you have found a substantiation limit, and that is itself a finding. And instrument the review step from day one; retrofitting review-time capture mid-term destroys comparability across the pilot.
Which vendor claims can be substantiated, and which cannot?
Not every claim is dishonest, but claims differ in whether evidence for them can exist at all. Sort them before the demo, not after.
| Vendor claim | Substantiable? | What to ask for |
|---|---|---|
| “Accuracy of X% on subject Y” | Yes, conditionally | The question set, its provenance, the marking method, and the model version. If the set is theirs and unreleased, treat the number as marketing, not measurement |
| “Aligned to the national curriculum” | Yes | A mapping document at topic granularity, and whether alignment is enforced in generation or checked afterwards |
| “Saves teachers N hours per week” | Rarely | The baseline task, who measured it, cohort size, and whether review time was counted. Uncounted review time is where most of these numbers come from |
| “Student data is not used for training” | Yes | The contractual clause, the retention period, the sub-processor list, and the data region |
| “Detects AI-generated coursework reliably” | No | False-positive rate on non-native-speaker writing, and what the institution does when the tool is wrong. Detection scores are probabilistic and not evidence of misconduct on their own |
| “Improves learning outcomes” | Almost never at tool level | The study design and whether there was a control group. Outcome effects rarely isolate to one piece of software |
| “Personalises to each learner” | Depends | Whether personalisation reads a persistent learner model or only the current session’s context window |
| “Certified educator training available” | Not a fitness signal | Nothing about training availability speaks to error rate on your subjects. Useful for adoption; irrelevant to evaluation |
The detection row deserves emphasis because it is the claim schools most want to be true. Stylometric and perplexity-based detectors return a likelihood, not a determination, and they misfire disproportionately on writing that is already atypical — non-native speakers, heavily scaffolded writers, students using assistive tools. An institution that treats a detector score as proof has built a disciplinary process on a probabilistic signal it cannot audit. The workable response is assessment design that makes provenance observable — in-class components, drafts, viva-style checks — rather than post-hoc detection.
Keeping human review without erasing the saving
The obvious tension: review is what makes generated material safe, and review is what consumes the time the tool was meant to give back. Blanket review of everything usually cancels the benefit outright.
The resolution is to tier review by consequence rather than by volume. Items with a single verifiable correct answer — answer keys, worked solutions, factual summaries — need checking every time, because that is exactly where confident error hides and where the cost of being wrong is highest. Items where the model is producing form rather than truth — a paraphrase at a different reading level, a lesson-plan skeleton, a set of discussion prompts — can move to sampled review once your acceptance-rate data shows the sample is stable. And anything that will be seen by learners without a teacher in the loop, or that contributes to a summative mark, keeps mandatory human sign-off regardless of how good the numbers get.
That last line is not a hedge. A model with a 3% factual-error rate is a useful assistant and an unacceptable unsupervised marker, and the difference is entirely about who bears the consequence of the 3%.
Data questions to answer before adoption, not during
These are unglamorous and they gate everything else. Where is student data processed, and under which jurisdiction. Whether prompts and outputs are retained, for how long, and whether retention can be switched off contractually rather than by a settings toggle. Which sub-processors see the data. Whether the vendor can produce a deletion confirmation for a named learner. Whether the model provider’s terms permit training on your submissions, and whether the vendor’s contract with you is stronger than their contract with the model provider.
A tool that scores well on error rate and cannot answer these is not a candidate. In our experience across regulated domains, this is the stage where procurement timelines actually break — not on capability, on data provenance — and starting it after a successful pilot wastes the pilot.
Where this leaves the decision
Evaluating generative AI for education is mostly an exercise in building a reference you can measure against, then holding vendors to the subset of their claims that evidence can reach. The tooling landscape will keep moving; the held-out question set, the review log, and the cost-per-accepted-artefact figure will still be the things that let you compare next year’s options to this year’s.
For how this fits the wider picture of what AI is and is not doing in classrooms and institutions, see our overview of AI in education, which maps the surrounding problem classes this evaluation method sits inside.
One thing we cannot resolve from the outside: whether a term is long enough. A single term gives you an acceptance curve and an error rate, but not the answer to whether a cohort taught partly with generated material ends up differently prepared. That question needs a longer instrument than a pilot, and nobody selling a tool is in a position to answer it for you.
Frequently Asked Questions
How should a reader evaluate ‘generative ai for education’?
Start from the learning task and its failure mode rather than the feature list. Build a held-out question set where you already know the correct answers, run each candidate tool against it, and log the factual-error rate, the teacher review time per item, and the proportion of output accepted without edit. Those three numbers, gathered over a term, support an accept/reject decision; feature counts and usage statistics do not.
What does a generative AI pilot in a school or university measure over a single term?
Three things: factual-error rate on a held-out question set, teacher review time per generated item split into accepted/edited/discarded, and the share of material accepted without edit tracked weekly. Combine them into cost per accepted artefact so tools are comparable on the same axis. Log prompts, outputs and model versions throughout, because a silent vendor model update can shift quality with no other trace.
Which vendor claims about generative AI for education can be substantiated, and which cannot?
Accuracy figures, curriculum-alignment mappings and data-handling commitments are substantiable if the vendor will supply the question set, the mapping document and the contractual clauses. Time-saving figures usually are not, because review time is rarely counted in them. Learning-outcome claims almost never isolate to a single tool, and reliable AI-detection claims are not substantiable at all — detectors return likelihoods, not determinations.
How do you keep a human review step in place without erasing the time saving?
Tier review by consequence, not by volume. Anything with a single verifiable correct answer — answer keys, worked solutions, factual summaries — gets checked every time. Form-oriented output such as paraphrases or lesson skeletons can move to sampled review once acceptance-rate data is stable. Anything that reaches learners unsupervised or contributes to a summative mark keeps mandatory sign-off regardless.
What data protection and student-data questions need answering before a tool is adopted?
Where student data is processed and under which jurisdiction; whether prompts and outputs are retained and for how long; whether retention can be disabled contractually rather than by a settings toggle; the sub-processor list; whether deletion for a named learner can be confirmed; and whether the underlying model provider’s terms permit training on your submissions. A tool that cannot answer these is not a candidate, whatever its error rate.
Evaluation criteria that separate prototype from production
Request the vendor’s retention policy, hallucination audit trail, and per‑seat cost breakdown before any pilot. Generative AI Education rewards teams that measure first and argue later — start with the smallest instrumented slice and let the numbers settle the design.