A perception robustness audit measures how a model behaves on the production driving distribution. It does not certify the vehicle as safe, and the two are not interchangeable inputs to a release decision. The audit produces engineering validation evidence — per-scenario-class failure rates, edge-class findings, sensor-variance coverage — that a safety assessor can consume. It does not produce the safety case, the hazard analysis, or the residual-risk argument, and no audit report can promise zero edge-case failure in the field.
That distinction sounds pedantic until you watch it fail. The pattern we see is a team that commissions an audit late, reads the pass result as a clearance, and arrives at the homologation gate with a folder of validation metrics where an assessor expected work products from a functional-safety process. Nothing in the folder is wrong. It is simply the wrong category of document, and the calendar cost of discovering that at the gate is measured in release slips rather than engineering days.
Why does treating a robustness audit as certification fail?
Because the two activities answer different questions and are owned by different parties.
An audit asks an empirical question: on the distribution we characterised, how often and under what conditions does this model fail? The answer is a measurement with a stated scope. It has boundaries — the weather buckets sampled, the sensor rigs covered, the edge classes enumerated — and those boundaries are part of the result, not a caveat bolted on at the end.
A safety case asks an argumentative question: given everything we know, why is the residual risk of this system acceptable, and who is accountable for that judgement? That argument is built from hazard analysis, a safety concept, verification against derived safety requirements, and traceability from hazards to evidence. Standards work — ISO 26262 for functional safety of E/E systems, and SOTIF for hazards arising from performance limitations rather than faults — defines the shape of those work products. An audit supplies evidence into that argument; it never authors the argument.
The conflation is easy to make because both activities produce documents that look like validation. The difference is that one is a measurement anyone competent can reproduce, and the other is a claim someone signs.
Where engineering validation ends and the safety case begins
The cleanest way to hold the boundary is to write it down before the statement of work is signed, per artefact rather than per phase.
| Artefact | Robustness audit produces it | OEM / Tier 1 safety case owns it |
|---|---|---|
| Characterisation of the production driving distribution | Yes — route, weather, lighting, road-class and sensor-rig mix used for test-set construction | Reviews it for representativeness against the intended operational domain |
| Per-scenario-class failure rate | Yes — measured, with sample counts per class | Consumed as verification evidence against performance requirements |
| Edge-class enumeration and coverage gaps | Yes — including explicit statement of what was not sampled | Feeds triggering-condition analysis |
| Sensor-mounting and calibration variance results | Yes | Maps to relevant hazards and degradation modes |
| Drift instrumentation and monitoring plan | Yes — as an engineering deliverable | Accepts or rejects it as a field-monitoring measure |
| Hazard analysis and risk assessment | No | Yes |
| Derived safety requirements and safety concept | No | Yes |
| ISO 26262 / SOTIF work products | No | Yes |
| Residual-risk acceptance and sign-off | No | Yes — named individual, named role |
| Homologation submission | No | Yes |
Read the two right-hand columns as a division of labour, not a hierarchy. The audit column is where measurement lives. The safety-case column is where accountability lives. A vendor who blurs the columns is offering something they cannot deliver, and an internal team that blurs them is deferring a conversation that gets more expensive every sprint.
The engineering discipline behind that division is not automotive-specific — it is the same scope hygiene we apply to reliability work generally, applied here to a workload where the consequences of scope confusion are regulatory. Our broader treatment of how perception robustness gets measured against production conditions sits in automotive perception robustness and the production driving distribution, which is where the measurement methodology itself is developed.d.d.d.
Warning signs the boundary is slipping
These are the signals worth acting on early, whether they come from a vendor or from your own programme.
- The statement of work uses “certification”, “compliance”, or “approval” to describe an audit deliverable, rather than “validation evidence” or “measurement report”.
- No named individual owns residual-risk sign-off. If you cannot name the person, the safety case does not exist yet.
- The audit scope was written by the model team alone, with no input from the safety or homologation function about what evidence they will need to cite.
- The proposal claims edge cases will be “covered” or “handled” rather than enumerated, sampled, and reported with known gaps.
- Aggregate accuracy is offered as the headline result, with per-scenario-class breakdowns available “on request”.
- Nobody has asked what happens to the audit conclusions when the model is retrained.
Any one of these is recoverable in a week if caught during scoping. All six together, discovered at a gate review, is a re-plan.
What to commission alongside the audit
If homologation is the eventual destination, the audit is one workstream of three, and they should be scoped in the same conversation.
- The robustness audit itself — test-set construction against the characterised production distribution, per-scenario-class failure rates, edge-class coverage with stated gaps. This is the engineering measurement.
- The safety-case evidence map — a short document, owned by the safety function, that states which hazards and derived requirements each audit output will be cited against. Written before the audit runs, it changes what the audit measures. Written after, it exposes gaps you now have to close under time pressure.
- The monitoring and drift plan — because an audit is a point-in-time measurement and the field distribution moves. Our validation evidence work packages this as a standing harness rather than a one-off report; see Production AI Monitoring Harness for how the instrumentation side is scoped.
In our experience, the second item is the one teams skip, and it is the cheapest of the three. A two-page evidence map produced early is what converts an audit from a reassuring document into a citable input.
Worth stating plainly: none of this makes the audit less valuable. It makes the audit usable. A measurement with honest boundaries can be cited in a safety argument. A measurement dressed as a clearance cannot be cited by anyone, because an assessor’s first question is what it did not cover — and a document that claims to have covered everything has no answer.
The same scope discipline applies to the vision engineering underneath the audit, which we cover across our computer vision practice.
Frequently Asked Questions
What does it mean in practice that a perception robustness audit is not safety certification?
Confusing robustness audits with safety certification leads teams to misallocate validation resources and misrepresent system capabilities. Stripped down, Perception Robustness Audit is the following. Perception Robustness Audit is one of those terms that hides a simple idea. Stripped down, Perception Robustness Audit is the following. Perception Robustness Audit works like this. Perception Robustness Audit is one of those terms that hides a simple idea. Stripped down, Perception Robustness Audit is the following. Perception Robustness Audit is one of those terms that hides a simple idea. Stripped down, Perception Robustness Audit is the following. Perception Robustness Audit works like this. Perception Robustness Audit is one of those terms that hides a simple idea. Stripped down, Perception Robustness Audit is the following. Perception Robustness Audit is one of those terms that hides a simple idea. Stripped down, Perception Robustness Audit is the following. Perception Robustness Audit has one honest answer. Perception Robustness Audit works like this. For What a Perception Robustness Audit Is specifically, in real use, it means the audit output is a measurement report, not an approval. It tells you how the model performed on a defined slice of the production driving distribution and what that slice excluded. Certification is a judgement about acceptable residual risk that an OEM or Tier 1 makes using that measurement plus hazard analysis, safety requirements, and process work products the audit does not produce., validation ends at the measurement and its stated scope: test-set provenance, per-scenario-class failure rates, coverage gaps, sensor-variance results. The safety case begins where those numbers are argued against identified hazards and derived safety requirements, and it ends with a named person accepting residual risk. The handover point is the evidence map that says which measurement supports which requirement.
What evidence from a robustness audit can be reused as an input to a functional-safety or SOTIF argument, and what cannot? Reusable: distribution characterisation, per-scenario-class failure rates with sample counts, edge-class enumeration, calibration and mounting variance results, and the monitoring plan. Not reusable as-is: any aggregate score without scenario breakdown, and any statement about overall system safety. The audit also cannot supply hazard analysis, the safety concept, or traceability to safety requirements — those are safety-process outputs.
Who signs off on residual risk if the audit does not, and what do they need from the audit to do it? The OEM or Tier 1 safety function signs, through a named role rather than a committee. What they need is traceable evidence: where the test data came from, how it maps to the intended operational domain, failure rates per scenario class, and an explicit list of what was not covered. Undocumented gaps are what block a sign-off, not high failure rates that are known and bounded.
What should a team commission alongside a robustness audit if homologation is the eventual goal? Three things in one scoping conversation: the audit, a safety-case evidence map written by the safety function before the audit runs, and a drift-monitoring plan for after release. The evidence map is the item most often skipped and the cheapest to produce; it changes what the audit measures rather than merely documenting it afterwards.
How do we scope an audit statement of work so it is not misread internally as a certification deliverable? Name the deliverable as validation evidence, list the specific artefacts it contains, and add an explicit non-scope section covering hazard analysis, safety work products, residual-risk acceptance, and homologation. Have the safety function countersign the scope. If the word “certification” appears anywhere in the document, it should be in the exclusions.
What are the warning signs that a vendor is implying certification it cannot provide? Language that promises approval, compliance, or “handled” edge cases rather than measured failure rates with stated gaps; an aggregate score as the headline result; no named residual-risk owner; and a scope written without safety-function input. A vendor confident in their measurement will volunteer its boundaries unprompted.
Which leaves the question worth taking into your next scoping meeting: if your audit came back clean tomorrow, could you name the person who would sign the residual-risk argument, and the document they would cite?
Audit vs. Certification: Why Confusion Persists
Regulators care about process traceability and worst-case guarantees; robustness audits deliver statistical profiles and failure budgets. The teams that do tend to ship the boring, correct version first.