A validation report is not a clearance. Both produce documents, both use the word “validation”, and in imaging AI projects under pilot pressure the two get collapsed into one thing — usually in the slide that goes to procurement. The engineering side answers whether a model holds on the buyer’s actual image distribution and whether drift will be visible after deployment. The regulatory side answers whether the intended use, risk classification, and quality system satisfy a notified body or the FDA. Those are different questions with different owners, and the failure mode we see most often is a team that has answered one well and assumes it has answered both.
We build the engineering validation layer: the measured evidence pack and the monitoring design. We do not produce regulatory submissions or device clearances, and being explicit about that early is what makes the evidence useful rather than misleading.
What does the boundary look like in practice?
Take a segmentation or triage model going into a pilot at a hospital group with four scanner models across two vendors, mixed reconstruction kernels, and a decade of protocol drift in the archive. The engineering question is whether performance measured on a curated research cohort survives contact with that distribution — per-scanner, per-protocol, per-subgroup — and what monitoring will tell you when it stops surviving. That question has a technical answer you can produce in weeks and defend with numbers.
The regulatory question is categorically different. It asks what the software’s intended use statement says, what risk class that intended use implies, whether the quality management system behind the development process is documented and auditable, what the clinical evaluation strategy is, and how post-market surveillance is formalised. None of that is answered by a stronger AUC. A model can be excellently validated in the engineering sense and hold no regulatory standing whatsoever — and, less obviously, a cleared device can still fail on a specific buyer’s scanner mix, which is precisely why the engineering layer keeps mattering after clearance exists.
The two documents also fail in opposite directions when confused. Teams that treat the eval pack as a clearance proxy walk into procurement making a claim they cannot support. Teams that make the opposite error — assuming no engineering validation is worth building until a regulatory pathway is scoped — stall for months and then discover that the dossier needs evidence they never collected.
Where each artefact stops
| Question | Engineering eval evidence pack | Regulatory submission |
|---|---|---|
| Does performance hold on this buyer’s scanner and protocol mix? | Yes — measured, stratified | Not its job |
| Are subgroup and site-level gaps quantified? | Yes | Referenced, not generated |
| Will post-deployment drift be visible, and by what signal? | Yes — monitoring design is part of the pack | Expects surveillance to exist |
| What is the declared intended use and risk class? | Out of scope | Core requirement |
| Is there an auditable quality management system behind development? | Partially — engineering process records only | Core requirement |
| Clinical evaluation and labelling | Out of scope | Core requirement |
| Legal authority to market or use clinically | No | Yes, if granted |
| Who signs | Engineering lead, with vendor | Regulatory affairs / legal manufacturer |
The bottom row is the one that decides everything above it. Only the right-hand column carries market authorisation, and no amount of measured performance moves a claim from the left column to the right.
Can validation evidence feed a regulatory pathway?
Usually yes, and that is the practical argument for building it early — but only if three things are true at the time of collection.
First, the data provenance has to be traceable: which cohort, which sites, which acquisition parameters, with the inclusion and exclusion logic written down before the numbers were produced rather than reconstructed afterwards. Second, the test set has to be genuinely held out and the holdout discipline documented, because a reviewer’s first question is whether the reported figures leaked training data. Third, the analysis plan has to be pre-specified — stratification axes, primary metrics, and acceptance thresholds fixed in advance, not selected once results were visible.
An evidence pack built with those three properties is a defensible input to a clinical evaluation. One built without them typically has to be rebuilt from scratch, which is the rework cycle the boundary discipline exists to avoid. In our experience the rebuild is more expensive than doing it properly the first time, mostly because the original cohort access is hard to reacquire.
There is a related question about GxP scope. When an imaging deployment sits inside a GxP-regulated process — a pharma imaging workflow rather than a clinical diagnostic one — computerised system validation applies to the system as installed and operated, and that is a third artefact again, distinct from both columns above. It overlaps the engineering pack on qualification testing and overlaps the regulatory column on documentation formality, but it does not substitute for either. We treat it as a separate scoping conversation, and the governance and workflow evidence side of the project is where it lands.
What you can honestly claim with validation evidence alone
The safe formulation is narrow and specific: measured performance on this cohort and this scanner mix, with these stratified results and this monitoring coverage. That is a strong statement. It is stronger than most vendors bring to a pilot conversation, and it is the kind of statement that survives a clinical stakeholder reading it carefully.
What it does not license:
- “Validated for clinical use” — conflates measurement with authorisation
- “Regulatory-ready” — implies a pathway assessment nobody performed
- “Equivalent to a cleared device” — a substantial-equivalence argument belongs in a submission, not a slide
- Any claim of intended use beyond what the pilot protocol actually covers
Overstated clearance claims tend not to surface in engineering review. They surface later, in procurement, or with clinical affairs, and the escalation is expensive in a way that has nothing to do with model quality. Keeping the boundary explicit also shortens reviewer round-trips, because reviewers stop receiving evidence that answers a question they did not ask. Our broader position on what buyers actually inspect is developed in what medical imaging AI buyers check before a pilot, and the monitoring design that makes the drift half of the pack real is the Production AI Monitoring Harness, scoped to clinical imaging for life sciences work.
Frequently Asked Questions
What does the boundary between validation evidence and regulatory clearance mean in practice for a medical imaging AI project?
Under the hood, Validation Evidence vs Regulatory is this. In deployment, it means two separate deliverables with separate owners. Engineering validation measures whether the model holds on the buyer’s real scanner and protocol distribution and designs the monitoring that will reveal drift; regulatory clearance establishes intended use, risk class, quality-system compliance, and market authorisation. Neither one produces the other, and a project plan that names only one of them is incomplete.
What does an engineering eval evidence pack contain, and what does a regulatory submission require that it does not cover?
The eval pack contains stratified performance on the buyer’s cohort and scanner mix, subgroup and site-level breakdowns, documented data provenance and holdout discipline, and a monitoring design with defined drift signals. A submission additionally requires an intended-use statement, risk classification, an auditable quality management system, clinical evaluation, labelling, and post-market surveillance commitments — none of which follow from performance numbers.
Can validation evidence be reused as input to a regulatory pathway, and what has to be true for it to be reusable?
Usually yes, provided three conditions held at collection time: traceable data provenance with pre-written inclusion and exclusion logic, a genuinely held-out test set with documented holdout discipline, and a pre-specified analysis plan fixing metrics and stratification before results were seen. Evidence produced without those properties generally has to be rebuilt, and reacquiring the original cohort access is the expensive part.
What claims can a team legitimately make in procurement or pilot conversations on the strength of validation evidence alone?
Only the measured claim: performance on this named cohort and scanner mix, with these stratified results and this monitoring coverage. “Validated for clinical use”, “regulatory-ready”, and any equivalence-to-a-cleared-device framing all overstate what the evidence supports. The narrow claim is more persuasive to clinical stakeholders anyway, because it survives being read closely.
Regulators care about safety claims, not statistical perfection
A device with 10,000-sample validation may still face rejection if its intended use statement implies clinical benefits the evidence cannot support, while a 200-sample study with tightly scoped claims can sail through. If Validation Evidence vs Regulatory is on your roadmap, the next step is to map it onto your own constraints rather than copy a reference architecture.