Two artefacts, two owners, two different questions. The perception validation package answers whether the perception component behaves as specified. The regulatory safety case answers whether the system is safe to release. Confusing them is expensive in both directions, and the bill usually arrives mid-release.
The pattern we see most often is the first one: a perception lead is handed a request for “the safety documentation” and starts writing argumentation the perception team has no standing to make — claims about hazard mitigation at vehicle level, about residual risk acceptance, about whether the driver monitoring fallback is adequate. None of that is a perception-component question. The reverse failure is quieter and often costlier. A safety engineer receives a test summary with strong numbers, no stated operational domain, no traceable dataset provenance, and no failure taxonomy, and has to rebuild component evidence from scratch before the argument can be assembled at all.
What does the boundary between a perception validation package and a regulatory safety case actually mean?
It is a boundary of claim, not of format. The perception package makes claims about a component’s measured behaviour within stated conditions and states the limits of those claims. The safety case makes claims about the system’s acceptable risk in a defined operational context, and it consumes component claims as evidence.
The practical test is what a reviewer’s question is asking. When a reviewer asks “does this show the perception component behaves as specified?”, the package must answer. When the reviewer asks “does this show the system is safe?”, the package should decline — and name what it feeds instead. A package engineered around the second question is a package that has quietly absorbed argumentation its authors cannot defend under scrutiny.
That refusal is not evasion. It is the single most useful sentence a validation package can contain, because it tells the safety team exactly which part of their argument is now supported and which part is still open.
The routing table
The following split is what we fix before packaging begins on automotive perception work. Reviewer questions get routed once, not renegotiated per review round.
| Question a reviewer asks | Answered by the perception package | Answered by the safety case |
|---|---|---|
| What operational design domain does the measured performance cover? | Yes — stated as scope, with the conditions actually tested | Restated as the ODD the system claims, which may be narrower |
| What is the detection performance on pedestrians at dusk? | Yes — per-slice metrics with dataset provenance | No — consumed as evidence |
| Is the false-negative rate acceptable for this release? | No — the package states the rate and its measurement conditions | Yes — acceptability is a system-level risk judgement |
| Which hazards does this component contribute to? | Partially — the package supplies the hazard-to-test mapping it was given | Yes — hazard identification and allocation is safety-case work |
| Is the fallback strategy sufficient when perception degrades? | No — the package documents degradation behaviour only | Yes |
| Was the training and evaluation data independent? | Yes — dataset manifests, version hashes, split rationale | No — consumed as evidence |
| Does the system meet the applicable standard? | Never | Yes |
| What changed since the previous release, and what was re-measured? | Yes | Yes, as an impact argument on the existing case |
The middle column is deliberately the longer one. Component evidence is where the perception team has authority; system argumentation is where it does not. A package that answers a question in the right-hand column has overreached, and that overreach is what gets discovered late — usually when the OEM’s safety function reads a sentence the supplier cannot substantiate.
What the package has to hand over
Handover quality determines whether the safety team builds an argument or rebuilds evidence. The following six items are what makes component evidence usable upstream rather than merely present:
- A scope statement naming the conditions under which the measurements hold — sensor configuration, resolution, frame rate, weather and lighting bands, object classes in scope.
- Dataset and scenario provenance with version identifiers, so a claim can be re-derived rather than re-trusted.
- Metric definitions with thresholds and rationale — a number without its definition is not evidence, because two teams will compute recall differently on the same data.
- A failure-mode log with disposition — what was searched for, what was found, what was fixed, what remains open.
- Stated limits and known blind spots, phrased as claims the package explicitly does not make.
- A traceability index from each section to the artefact behind it.
Recorded this way, operational domain assumptions and failure modes become safety-case inputs directly. Recorded as prose narrative, they become interview material — and the interview is where review rounds multiply.
The economics of getting it wrong
A stated boundary removes the largest single source of rework we see in perception sign-off: deliverables re-scoped after review because the reviewer’s question sat on the other side of the line (an observed pattern across our automotive validation engagements, not a benchmarked figure). Teams that fix the boundary before packaging tend to keep sign-off to one review round instead of the two or three spent renegotiating what the deliverable was supposed to contain.
There is a second, less obvious return. Evidence sections that are not entangled with system-level argumentation stay reusable across model updates, because they do not change when the vehicle configuration changes. Argumentation, by contrast, is configuration-specific — bind the two together and every retrain drags the safety narrative along with it. The invariant-structure discipline that makes a package survive retraining is the same discipline that keeps the boundary clean; we cover the structural side of that in the package structure that survives model updates.
Where the boundary moves
The line is not fixed across programmes, and three variables move it.
Automation level. In a Level 2 driver-assist release, the driver is the fallback, and a great deal of weight sits on the component behaving as specified. In a Level 3 release, the system-level argument carries far more of the load — the same component evidence now supports a much larger claim, so the package’s stated limits are read more adversarially. The evidence does not change; the scrutiny applied to what it does not cover does.
Organisational split. When the perception supplier and the vehicle integrator are different companies, the boundary is also a contractual interface. The supplier owns the component package and its limits; the integrator owns the safety case and the risk acceptance. Where this goes wrong is when the integrator’s requirement document asks for safety-case content from a supplier who has no visibility of the vehicle-level hazard analysis. Naming the artefact owner in the statement of work costs one paragraph and saves a re-scope.
Governance depth. Some OEM reviews want more than component evidence but less than a safety argument: change control, independent auditability, ownership of post-SOP model updates. That is a governance-grade evidence layer that feeds a safety case rather than substituting for one, and it sits alongside the validation package rather than inside it. The transition between the two is covered in what an OEM governance review asks beyond release validation.
Signs the line is in the wrong place
Three symptoms show up before the cost does. First, the package contains a sentence starting “the system is safe because” — component evidence has crossed into argument. Second, the safety team asks for a metric definition or a dataset split that the package never recorded — the handover is incomplete and evidence is about to be rebuilt. Third, reviewer questions in a single meeting land in both columns of the table above with no visible routing — nobody agreed the boundary, so it is being negotiated live, at the most expensive possible moment.
Correcting mid-release is not catastrophic, but it is not cheap. Pulling argumentation out of a package means re-issuing it; adding missing provenance means re-running measurements whose configuration was never recorded. In our experience the second is the worse of the two, because unrecorded conditions cannot be reconstructed after the fact — the data campaign has to happen again.
The wider engineering problem this sits inside — how evidence is structured so a reviewer can act on it — is what we cover across our work on production AI reliability, where perception validation is one lens on a broader question: what does a system have to demonstrate about itself before someone signs for it?
So the boundary question is not paperwork. It is a question about authority: which claims can this team defend, and which ones belong to someone with a wider view of the vehicle? Answer that first, and the package writes itself around it.
Frequently Asked Questions
What does the boundary between a perception validation package and a regulatory safety case mean in practice?
Perception validation packages document algorithmic performance, while regulatory safety cases argue system-level risk acceptance under legal frameworks. Perception Validation Package vs behaves predictably once you see the mechanism. Perception Validation Package vs Regulatory Safety makes this clear: it is a boundary of claim rather than of document format. The perception package states what the component was measured to do, under which conditions, and what it does not cover; the safety case argues that the system’s risk is acceptable and uses the component evidence as input. In practice the line shows up in reviewer questions — “does the component behave as specified?” belongs to the package, “is the system safe?” does not., the package answers questions about measured behaviour: operational domain covered, per-slice metrics, dataset provenance, degradation behaviour, and what changed since the last release. It should refuse acceptability judgements, hazard allocation, fallback sufficiency, and standard conformance — all of which require a vehicle-level view the perception team does not hold. A clean refusal that names the receiving artefact is more useful than an answer the organisation cannot defend.
What does the perception package have to hand the safety case team so the argument can be built without re-deriving component evidence?
Six things: a scope statement fixing the tested conditions, dataset and scenario provenance with version identifiers, metric definitions with thresholds and rationale, a failure-mode log with disposition, explicitly stated limits, and a traceability index from each section to its underlying artefact. With those, the safety team assembles an argument. Without them, it rebuilds measurements.
How should stated limits, operational domain assumptions and known failure modes be recorded so they are usable as safety-case inputs?
Record them as claims and non-claims rather than narrative prose — the conditions under which each number holds, and the conditions the package explicitly does not speak to. Failure modes belong in a log with disposition (fixed, mitigated, open, accepted) rather than in a discussion section. This format lets a safety engineer lift the item straight into an argument; prose has to be interviewed out of the author.
Who owns each artefact when the perception supplier and the vehicle integrator are different organisations?
The supplier owns the perception validation package and the accuracy of its stated limits; the integrator owns the safety case and the risk acceptance, because only the integrator sees the vehicle-level hazard analysis. Problems begin when a statement of work asks the supplier for safety-case content it has no visibility to produce. Naming the owner per artefact in the contract is the cheapest possible fix.
How does the boundary shift between a Level 2 driver-assist release and a Level 3 release?
The evidence the package owes does not change, but the weight placed on the system-level argument does. At Level 2 the driver is the fallback, so component conformance carries much of the load; at Level 3 the same component evidence supports a far larger claim, and the package’s stated limits are read much more adversarially. Expect deeper scrutiny of what the package does not cover, not of what it does.
What are the signs the boundary has been drawn in the wrong place, and what does the correction cost mid-release?
Watch for argumentation language inside the package, safety-team requests for provenance the package never recorded, and review meetings where questions from both sides of the line arrive unrouted. Pulling argumentation back out means re-issuing the package. Restoring missing provenance can mean re-running measurements whose configuration was never captured — which is the expensive correction, because unrecorded conditions cannot be reconstructed afterwards.
Two documents with overlapping but distinct mandates
Perception packages demonstrate technical capability under defined conditions; safety cases argue that those capabilities, combined with fallback mechanisms and operational constraints, meet harm thresholds regulators will accept. The teams that do tend to ship the boring, correct version first.