A validation pack and a regulatory submission defend two different claims to two different readers. The pack defends “this model performs acceptably here, on this distribution, under this ground-truth protocol.” A submission defends “this device, with this intended use, is safe and effective for market.” Those are not two lengths of the same document. They are two arguments, and the decision worth making early is which artefacts serve both, which serve only one, and who each one is written for.
The failure we see most often is a pack that grew. A team builds the evidence a hospital reviewer asked for, then keeps adding — a risk section here, a literature summary there — on the theory that a fatter document is a safer document. What they end up with is a pack too heavy to circulate through procurement and a submission still missing its spine: the design-history and risk-management documentation that no volume of validation statistics substitutes for. Two failures from one avoidable ambiguity.
Which audience is adjudicating what?
The divergence point is not document length. It is the decision the reader is making.
A hospital or health-system reviewer is adjudicating fitness in their environment. They want to know whether performance established elsewhere holds on their scanner fleet, their acquisition protocols, their case mix, their prevalence — and whether the ground truth the model was scored against was constructed in a way they would accept. Their question is portability.
A regulator is adjudicating intended use, risk classification and device claims. They want the statement of what the device is for, the risk analysis behind that statement, the clinical evaluation strategy that supports it, and the plan for what happens after it is on the market. Their question is whether the claim itself is admissible.
Both readers will look at the same performance numbers. They will read them for different reasons, and that changes framing more than content. The same stratified performance table that a site reviewer reads as “does this hold for us” is read by a regulator as evidence supporting or bounding an intended-use statement. Authoring once and reframing twice is cheap. Rebuilding is not.
The three-bucket sort
The practical decision reduces to sorting every artefact into one of three buckets. In our experience this sort takes an afternoon and saves months of duplicated evidence work.
| Artefact | Bucket | Primary audience | Framing note |
|---|---|---|---|
| Validation-set construction protocol | Shared | Site reviewer + regulator | Site reviewer reads it for distribution match; regulator reads it as evidence design |
| Ground-truth adjudication evidence (readers, conditions, disagreement resolution) | Shared | Both | Same records; submission adds the procedural control context |
| Performance reporting, stratified by scanner / protocol / subgroup | Shared | Both | Pack leads with the site’s strata; submission leads with intended-use population |
| Intended-use statement | Submission-only | Regulator | A pack that asserts an intended use is making a device claim |
| Risk management file | Submission-only | Regulator | Not derivable from validation statistics |
| Clinical evaluation strategy | Submission-only | Regulator | Positions the evidence against the claim |
| Post-market surveillance plan | Submission-only | Regulator | Population-level obligation, not local operations |
| Design history / change-control records | Submission-only | Regulator | The spine a grown pack never acquires |
| Site-specific portability evidence | Pack-only | Site reviewer | Meaningless to a regulator; decisive locally |
| Deployment drift telemetry framed for local operations | Pack-only | Site reviewer + local ops | Same pipeline as surveillance, different audience and cadence |
| Local workflow and integration evidence (PACS, reading order, escalation path) | Pack-only | Site reviewer | Operational fitness, not device safety |
The one row worth pausing on is drift telemetry. The pipeline that produces it is a single pipeline — the same monitoring of input distribution and output score behaviour against named validation strata. But a site-facing pack presents it as an operational control: which shifts are watched, what the thresholds are, who adjudicates a flagged case at this site, how often. A post-market surveillance plan presents the same instrumentation as a population-level commitment: what constitutes a reportable signal across all deployments, and what the vendor will do about it. Build the telemetry once; write the section twice. We treat that as the default assumption rather than a discovered optimisation, and it is developed further in our work on post-deployment drift evidence in a clinical imaging validation pack.
What the pack deliberately does not attempt
A validation pack is not a partial submission, and saying so in the pack is protective rather than an admission of a gap. It makes no intended-use claim, asserts no risk classification, and does not present itself as a regulatory determination. That carve-out is what keeps the pack honest: it lets the document say precisely what it can defend — measured performance under a stated protocol on a stated distribution — without implying a market authorisation nobody has granted.
The same discipline applies to shared artefacts that live in both worlds under different framing. HIPAA and GxP workflow evidence is the clearest case: the lawful basis under which validation data was assembled, the named readers who adjudicated it, and the infrastructure the telemetry moves through are all load-bearing in a pack and in a submission, but the pack frames them as site-acceptable handling while a submission frames them as procedural control. We work through those seams in where a clinical validation pack meets HIPAA and GxP workflow evidence. The broader question of how reliability evidence is structured for production systems sits under our production AI reliability work.
Sequencing when a regulatory path is likely but unconfirmed
This is the common commercial position: the team expects to pursue a regulatory path eventually, has not committed to one, and needs revenue from research or non-diagnostic deployments now. The sequencing rule that holds up is to author shared artefacts to submission quality and pack format, and to defer submission-only artefacts until the path and jurisdiction are named.
Concretely, that means the validation-set construction protocol, adjudication records and performance report are written with the traceability a regulator would expect — versioned, dated, attributable — but laid out for a site reviewer who has forty minutes. Nothing in them needs rewriting later; only re-sectioning. Meanwhile the intended-use statement, risk file and clinical evaluation strategy wait, because writing them before the intended use is settled produces documents that must be discarded, and worse, produces claim language that leaks into the pack.
What signals a pack is becoming an accidental submission?
Four signals are reliable enough to use as a checklist:
- The pack has acquired an intended-use or indications-for-use section.
- Reviewers stop returning comments and start returning silence — the document is no longer readable inside a procurement cycle.
- Sections are being added in response to hypothetical regulator questions rather than actual site-reviewer questions.
- The pack contains risk language (“mitigations”, “residual risk”, “hazard analysis”) without a risk management file behind it, which is the worst of both states: submission vocabulary with no submission structure.
Any one of these is a prompt to re-run the three-bucket sort. Two or more usually means the document needs splitting rather than trimming.
The scoping question that comes first
Drawing this boundary is the first scoping step of a validation-pack engagement, not a late refinement. It determines what gets authored, in what order, and to whose standard — and it lets a team answer honestly, before a regulatory path opens, which artefacts already exist and which are genuinely missing. The parent discussion of what a regulated clinical deployment requires from its evidence set is in clinical imaging validation pack contents.
The uncertainty we do not pretend to resolve: the shared/submission-only split above is stable across the engagements we have seen, but the exact line moves with jurisdiction and risk class. What does not move is the discipline of naming which bucket each artefact belongs to before writing it. Which of your existing artefacts are you unsure about?
Frequently Asked Questions
What does the boundary between a validation pack and a regulatory submission mean in practice — where exactly does one end and the other begin?
Regulatory agencies draw a hard line between validation documentation and formal submission dossiers—understanding this distinction prevents costly rework. When applied to Where the Validation Pack Ends and, the boundary is the claim being defended. A validation pack defends measured performance on a stated distribution under a stated ground-truth protocol; a submission defends that a device with a declared intended use is safe and effective for market. The pack ends the moment a document starts asserting intended use or risk classification., shared: validation-set construction protocol, ground-truth adjudication evidence, stratified performance reporting. Submission-only: intended-use statement, risk management file, clinical evaluation strategy, post-market surveillance plan, design-history records. Pack-only: site-specific portability evidence, locally framed drift telemetry, and workflow-integration evidence.
Which audience is each artefact written for, and how does that change the way the same performance evidence is presented? A site reviewer adjudicates fitness on their own population and workflow, so the pack leads with the strata that match their scanner mix and case mix. A regulator adjudicates intended use and device claims, so a submission leads with the intended-use population. Same numbers, different ordering and different emphasis.
What does a validation pack deliberately not attempt to do, and why is that carve-out protective rather than a gap? The pack makes no intended-use claim, asserts no risk classification, and offers no regulatory determination. Saying so keeps the document defensible — it commits only to what measurement supports — and prevents claim language from implying a market authorisation nobody has granted.
How should a team sequence the work when a regulatory path is likely but not yet confirmed? Author the shared artefacts to submission-grade traceability but in pack format, and defer the submission-only artefacts until the path and jurisdiction are named. Writing an intended-use statement before intended use is settled produces documents that get discarded and claim language that leaks into the pack.
How does post-deployment drift telemetry sit differently in a site-facing pack versus a post-market surveillance plan? The instrumentation is the same pipeline. The pack presents it as a local operational control — which shifts are monitored, what the thresholds are, who adjudicates a flagged case at this site. A surveillance plan presents it as a population-level commitment across all deployments, with reportable-signal definitions and vendor obligations.
What are the signals that a validation pack is being over-built into an accidental submission? An intended-use or indications section appearing; reviewers going quiet because the document no longer fits a procurement cycle; sections added for hypothetical regulator questions rather than real site questions; and risk vocabulary used without a risk management file behind it.
Documentation handoff: two artefacts, two audiences
A validation pack demonstrates technical control; a regulatory submission demonstrates compliance with a legal framework — and the boundary sits exactly where your internal evidence stops being sufficient for an auditor.