The failure does not look like a bad document. It looks like a funded programme that promised to shorten the certification path and delivered a faster drafting layer instead — which is a real gain, described as the wrong thing. By the time an OEM reviewer or an external assessor asks which of your generated artefacts they are supposed to accept as a safety argument, the scope claim has already been made in a steering deck, and walking it back costs more than the automation saved.
So state the boundary before someone else states it for you. Document automation belongs to the drafting and reconciliation layer; certification is the adjudication step, and adjudication is not a document-production problem. That is the whole distinction, and everything below is a consequence of it.
Why does an AI-assembled evidence pack not substitute for a certification case?
Because the two artefacts have to survive different things.
A drafted supplier-compliance document has to survive reconciliation against its source. A reviewer opens the generated clause, follows the reference back to the supplier questionnaire, certificate, or test report that produced it, and confirms the two agree. That check is mechanical, repeatable, and — this is the point — automatable in its supporting evidence, because the ground truth is a document the supplier already sent you.
A homologation submission or a functional-safety case has to survive an assessor’s judgement against a standard, with a named accountable engineer behind each argument. There is no source artefact that contains the answer; the answer is a professional determination that the evidence assembled is sufficient for the hazard in question. Generation models are good at producing text that resembles such a determination. That resemblance is the hazard, not the feature.
Two claims worth extracting from this, both from what we see in regulated documentation work rather than from a published benchmark:
- Drafting automation ends at the point where a human decides that evidence is sufficient; certification work is almost entirely that decision. The percentage of a certification case that is text assembly is small, and it is the least interesting part.
- A document workflow is drafting-automatable when its output is a draft a reviewer will check, and adjudication-only when its output is a conclusion a reviewer will rely on. The test is about what the downstream reader does with the artefact, not about how hard the text is to write.
We use that second sentence as an intake filter, and it settles most arguments in one meeting.
Which side of the line does each document task fall on?
The line is drawn per workflow, not per department. A single supplier-onboarding process usually straddles it.
| Document task | Side of the line | What makes it so |
|---|---|---|
| Supplier questionnaire responses drafted from a governed answer library | Drafting | Output is a draft; every answer reconciles to a prior evidenced submission |
| Multi-vendor evidence normalisation and field matching | Reconciliation | Output is a comparison view; disagreements are surfaced, not resolved |
| Compliance evidence pack assembly with per-clause source references | Drafting + reconciliation | Output is a pack a reviewer audits by trace, not by prose |
| Gap identification against a requirement schema | Reconciliation (flag only) | The flag is a prompt for a human; the sufficiency call is not made |
| “This supplier’s evidence is adequate for the requirement” | Adjudication | A conclusion a reviewer relies on; accountable person required |
| Functional-safety argument (hazard → requirement → evidence sufficiency) | Certification | Survives assessor judgement, not source reconciliation |
| Homologation submission sign-off | Certification | Named accountable engineer; regulatory consequence attaches to the signature |
The rows above the double break are where automation earns its keep — and where the throughput numbers that justify the programme actually come from. The rows below it are where a generated artefact creates liability without creating capability.
Note what the table does not say: certification work gets no benefit from automation at all. It does. Retrieval over prior submissions, consistency checking across a requirement set, and change detection when a supplier revises an input all reduce the assessor-facing workload. What is excluded is the adjudication itself. Keeping that distinction crisp is the same governance carveout that applies to AI-assisted regulated documentation in any vertical — the mechanism is not automotive-specific, only the standards are.
The accountability chain looks different on each side
On the drafting side, accountability is distributed and recoverable. If a generated clause misstates a supplier’s material declaration, the trace shows which source field and which revision produced it, the reviewer who accepted it, and the version state at acceptance. The remedy is a corrected draft and a re-review. Nobody’s professional sign-off is implicated, because nobody signed the draft as a determination.
On the certification side, accountability is singular and non-delegable. A named engineer asserts that the safety argument holds. That assertion cannot be produced by a pipeline, cannot be inherited from a model’s confidence score, and cannot be reconstructed after the fact from logs. This is why “the AI drafted it and I reviewed it” is an acceptable answer for a supplier evidence pack and an unacceptable one for a hazard analysis conclusion.
The practical consequence for tooling: on the drafting side you invest in traceability infrastructure, because the trace is what the reviewer audits. This is the same reasoning we develop in keeping traceability when supplier compliance documents are AI-generated. On the certification side you invest in making the assessor’s reading easier — indexing, cross-referencing, revision visibility — and you deliberately do not automate the conclusion.
Warning signs the scope has drifted
These are the signals that a programme has been sold against a certification claim it cannot support. Any two together are worth a scope review before the next funding gate.
- The business case is expressed as time-to-certification rather than as supplier onboarding cycle time or reconciliation throughput.
- No one can name, per workflow, whether the output is a draft or a determination. The classification exists nowhere as a written artefact.
- The pilot corpus includes safety-case fragments “to show the model can handle the hard stuff”.
- A reviewer’s role has quietly become approving batches rather than reconciling individual assertions to source.
- The generated artefacts have no per-clause source reference, so the only available audit is reading the prose — which means the pack cannot be defended by trace at all.
- Someone outside the programme has started describing the tool as “the compliance system” rather than “the drafting pipeline”.
The last one is the cheapest to catch and the most predictive. Language drift in an internal deck precedes scope drift in a review by about one quarter, in our experience.
How to phrase the boundary without undercutting your own automation
The instinct is to hedge, and hedging is what makes an OEM reviewer suspicious. The stronger move is to volunteer the boundary and then be specific about what sits inside it: we automate drafting and reconciliation for supplier-compliance evidence, with per-clause traceability to the source supplier input; sufficiency determinations and certification arguments remain with named engineers and are not model-generated. That sentence has never cost us a conversation. It usually shortens one, because it answers the question the reviewer was going to ask on slide fourteen.
It also changes the measurement conversation. A drafting-layer programme is measured on the share of document workload correctly classified as drafting-automatable versus adjudication-only, on traceability completeness across the automated portion, and on remediation cycles avoided after an OEM finding. It is not measured on certification readiness, and a programme that reports against certification readiness has already conceded the scope argument it was going to lose. The wider question of which workflows qualify in the first place is worked through in our analysis of automotive supplier-compliance document automation and in the sibling piece on where to draw the AI boundary between drafting assistance and compliance adjudication.
One uncertainty we have not resolved: as assessors themselves begin using retrieval and consistency tooling to read submissions, the reconciliation layer may become the surface they interrogate directly rather than the pack. If that happens, the boundary does not move — but the audience for your traceability artefacts does, and it is not obvious yet what that changes about how they should be built.
Frequently Asked Questions
What does the distinction between drafting automation and safety-critical automotive certification mean in practice?
Drafted documents face reconciliation audits against source data; certification arguments face assessor scrutiny against normative standards. The first check is mechanical and supportable by automation; the second is a professional determination with a named engineer behind it. Automation stops at the point the determination begins.
Which supplier-compliance document tasks fall on the drafting side of the boundary, and which are adjudication? Questionnaire drafting from a governed answer library, multi-vendor field normalisation, evidence-pack assembly with per-clause source references, and gap flagging are all drafting or reconciliation — their output is something a reviewer checks. Statements that a supplier’s evidence is adequate, hazard-analysis conclusions, and homologation sign-off are adjudication, because their output is a conclusion someone relies on.
Why can’t an AI-assembled evidence pack substitute for a functional-safety or homologation case? The evidence pack’s ground truth is a set of documents the supplier already sent, so its correctness is checkable by trace. A safety case has no such source artefact — the argument that the evidence is sufficient for the hazard is created by the accountable engineer, not extracted from an input. A model can produce text that resembles that argument, which is precisely why the substitution is unsafe.
What does the human accountability chain look like on each side of the line? On the drafting side accountability is distributed and recoverable: the trace shows source field, revision, transformation, and the reviewer who accepted it, and a defect is fixed with a corrected draft. On the certification side accountability is singular and non-delegable — one named engineer asserts the argument holds, and that assertion cannot be inherited from a pipeline or reconstructed from logs afterwards.
What warning signs indicate a document-automation programme has been scoped against a certification claim it cannot support? The business case is stated as time-to-certification rather than onboarding cycle time or reconciliation throughput; no written classification exists of which workflows produce drafts versus determinations; safety-case fragments appear in the pilot corpus; reviewers approve batches instead of reconciling assertions; and generated artefacts carry no per-clause source reference. Internal language drift — “the compliance system” instead of “the drafting pipeline” — usually arrives first.
Automating templates versus generating compliance arguments
Templates with fixed structure and variable parameters rarely trigger full tool qualification; free-text generation of safety reasoning always does. Revisit it when your workload shifts.