Audit the trace, not the prose. Once a compliance evidence pack is assembled with model assistance, reading the output for factual errors tells you whether the text is plausible; it does not tell you whether the pack survives a reviewer asking where a specific number came from. Those are different audits, and only one of them is the one an OEM performs.
The pattern we see when document automation lands in a supplier-compliance function is a predictable shift in the question being asked. For the first quarter or two, the question is throughput: can we turn twelve supplier submissions into an evidence pack in a day instead of a week. After the first OEM review cycle, the question becomes defensibility. The remediation loop that follows a compliance finding is the largest single cost item in this workflow, and it is almost always triggered by an unanswerable provenance question rather than a wrong sentence.
What does auditing AI-assisted compliance evidence actually involve?
It involves treating each assertion in the pack — not each document — as the audit object. An assertion is any statement a reviewer could challenge: a material declaration, a process capability figure, a date of certification, a scope-of-supply boundary. For each one, the audit asks four questions in order.
- Which supplier input does this derive from? A specific artifact, a specific field within it, not “the supplier’s submission”.
- What transformation was applied? Direct extraction, unit conversion, aggregation across revisions, or model-drafted paraphrase.
- Who accepted it, and against what? The named reviewer, and the source they reconciled against.
- What version state was in force at acceptance? The revision of the supplier input at the moment the assertion was accepted.
If all four resolve, the assertion is auditable. If any one does not, it is an unresolved-provenance item and belongs on an exception list rather than in the pack. This is deliberately a stricter test than “is the statement true” — a true statement with no resolvable source is still a finding waiting to happen, because the reviewer cannot verify it without redoing your work.
Sampling audit versus trace audit
Both have a place. The failure is using the cheaper one where the expensive one is required.
| Output-sampling audit | Traceability audit | |
|---|---|---|
| Audit object | Generated prose | Each assertion’s provenance record |
| Question answered | Does the text read correctly and consistently? | Can every claim be resolved to a source input and an acceptance event? |
| Coverage | A sample (commonly 5–15% of sections) | Complete; expressed as traceability completeness |
| Catches | Fluency errors, formatting drift, contradictions within the pack | Orphaned assertions, stale supplier revisions, silent gap-filling, rubber-stamp review |
| Misses | Assertions with no source at all | Prose quality, internal contradiction |
| Sufficient when | The output is a draft a reviewer will check line by line | The output is, or supports, a compliance assertion the OEM will rely on |
The distinction maps cleanly onto the boundary between drafting and adjudication, which we treat as a design decision rather than an emergent property of the tooling — see drafting assistance versus compliance adjudication for where that line belongs. Sampling is adequate on the drafting side. On the assertion side, sampling is a measurement of the wrong thing: a 10% sample of a pack with 400 assertions leaves 360 unexamined provenance links, and the reviewer’s challenge will land on one of them.
The pre-submission audit checklist
This is the internal pass, run before the pack leaves the building. It is deliberately mechanical, because the parts that need judgement come later.
- Traceability completeness. Share of generated assertions with a resolvable source link. Report the number, not an adjective. Anything below 100% needs a named reason per gap.
- Unresolved-provenance count. Absolute count per evidence pack, listed item by item with the reason each is unresolved.
- Revision currency. For every referenced supplier input, compare the revision cited in the trace against the current revision in the source system. Any mismatch is flagged, not silently refreshed.
- Transformation classification. Each assertion tagged as extracted, derived, or drafted. Drafted assertions carry the highest review weight because the model contributed content rather than moving it.
- Acceptance evidence. Each assertion has a named accepting reviewer and a record of what they reconciled against.
- Gap behaviour. Confirm the pipeline failed loudly on missing inputs. A pack with no gaps and no exception list is more suspicious than one with eight documented gaps.
- Reviewer response drill. Pick three assertions at random and time how long it takes to answer “where did this number come from” from the trace alone. This is the metric that predicts how the OEM session will go.
The last item is the one teams skip and the one that most reliably exposes a broken trace. Mean reviewer time to answer a source-of-truth challenge is measurable in your own conference room, weeks before anyone external asks. In our experience, when that number sits in the tens of minutes, the trace exists on paper but not in a form a human can traverse under pressure.
What belongs to you and what belongs to the reviewer
The OEM reviewer is not running your checklist. Their audit is narrower and sharper: they select assertions that matter to their risk model, follow the trace you provided, and record a finding where it breaks. They are auditing your evidence of control, not rebuilding it.
That split has a practical consequence. Your internal audit must be complete across all assertions, because you do not know which ones they will pick. Their audit is deliberately partial, because a partial audit of a complete trace is a valid test — and a partial audit of an incomplete trace produces a finding on the first pull. The asymmetry is the whole reason the internal pass cannot be a sample.
Two categories sit clearly on your side and never transfer: classification of transformation type, and the honesty of the exception list. Two sit on theirs: sufficiency of the underlying supplier evidence against their requirement, and the judgement call on whether a documented gap is acceptable.
Human review that is not a rubber stamp
A signature field proves someone clicked. It does not evidence review. What distinguishes the two is whether the acceptance record captures the reconciliation act rather than the approval act.
We treat three properties as the minimum. The reviewer’s view must present the generated assertion and its cited source field side by side, so acceptance requires looking at both. The record must capture what was reconciled, not just that acceptance occurred. And a meaningful share of reviews must produce rejections or corrections — a queue with a near-perfect acceptance rate across hundreds of assertions is evidence that the review step is not functioning, which is exactly the inference an experienced reviewer will draw (an observed pattern across regulated-document engagements, not a published benchmark).
Where the trace itself is the thing being engineered rather than audited, the mechanics of binding source-to-assertion at generation time are covered in keeping traceability when supplier compliance documents are AI-generated. The audit pattern here is the reviewer-facing read of the same evidence, and it is part of the traceability and validation scope in our engineering services.
Assertions with no resolvable source, and inputs revised after generation
Two situations break naive audits, and both have deterministic handling.
An assertion the model produced with no resolvable source input is not a traceability defect to be papered over — it is content the pipeline invented, however plausible. The correct handling is removal from the pack and, if the underlying statement is needed, re-derivation from a real supplier input with a fresh acceptance. Retro-fitting a source link to text that was not derived from it is the one move that converts a process gap into a misrepresentation.
Supplier inputs revised after a document is generated are more common and more insidious. The generated assertion is now bound to a superseded revision. The audit trail needs to show three things: the revision in force at generation, the revision in force now, and either a re-acceptance against the new revision or an explicit statement that the delta does not affect the assertion. Silent refresh — regenerating against the new input and keeping the old acceptance record — destroys the acceptance chain while appearing to improve currency.
Frequently Asked Questions
What does auditing AI-assisted compliance evidence for an OEM reviewer mean in practice?
OEM reviewers must verify that AI-generated compliance documentation maintains traceability to source requirements and human oversight at critical decision points. It means auditing each assertion’s provenance record rather than reading the generated document for errors. For every challengeable statement in the pack, you confirm the source supplier input, the transformation applied, the accepting reviewer, and the revision state at acceptance. The output of the audit is a completeness figure and an exception list, not a sign-off.
What is the minimum trace record each generated assertion must carry to be auditable?
Four fields: the specific source artifact and field it derives from, the transformation class (extracted, derived, or drafted), the named reviewer who accepted it with what they reconciled against, and the revision of the source input in force at acceptance. Any assertion missing one of the four is an unresolved-provenance item.
How do you distinguish an output-sampling audit from a traceability audit, and when is each sufficient?
A sampling audit examines a subset of generated prose for correctness and consistency; a traceability audit examines every assertion’s link back to source and acceptance. Sampling is sufficient where the output is a draft a reviewer will check line by line. Where the output is or supports a compliance assertion the OEM relies on, only the complete traceability audit answers the question they will ask.
Which checks belong to the internal pre-submission audit versus the OEM reviewer’s own audit?
Yours must be complete across all assertions — traceability completeness, revision currency, transformation classification, acceptance evidence, and gap behaviour — because you cannot predict which assertions they will select. Theirs is deliberately partial: they follow your trace on the assertions that matter to their risk model and judge whether the underlying supplier evidence is sufficient.
How do you evidence human review and acceptance without turning it into a rubber stamp?
Present the generated assertion and its cited source field together so acceptance requires looking at both, and record what was reconciled rather than only that approval happened. A review queue with a near-perfect acceptance rate across hundreds of assertions is itself evidence the step is not working.
How do you handle assertions the model produced that have no resolvable source input?
Remove them from the pack. If the statement is genuinely needed, re-derive it from an actual supplier input and record a fresh acceptance. Attaching a source link after the fact to text that was not derived from it converts a process gap into a misrepresentation.
What does the audit trail need to look like when supplier inputs are revised after a document is generated?
It must show the revision in force at generation, the current revision, and either a re-acceptance against the new revision or an explicit finding that the delta does not affect the assertion. Regenerating silently against the newer input while keeping the original acceptance record breaks the chain the reviewer is auditing.
The open question we have not seen settled anywhere is how much of the trace an OEM should be handed by default. Complete provenance for every assertion is auditable but overwhelming; a summary invites the challenge you cannot answer. Where that line sits is still being negotiated programme by programme.
Reviewer checklist for AI-generated compliance artifacts
Start by verifying that every AI-drafted claim maps to a human-reviewed source document, then trace the prompt chain and check for hallucinated citations or softened liability language.