Site two is where you find out whether you built a validation pack or wrote a results document. The pack that cleared the first hospital review gets resent, the reviewer reads it, and the first question back is not about the AUC. It is: how would this validation set have been built from our patients? That question has no answer in a results document. It has a short answer in a pack whose construction protocol was written down before any number existed.
This article walks one pack through that sequence — first site, second site, third site — and reports what survived the transfer unchanged, what had to be re-measured, and where portability genuinely stopped. What belongs inside the pack is a separate question; we treat the anatomy of the document in what a clinical imaging validation pack contains beyond the benchmark report. Here the subject is the pack in motion.
The two layers, separated before site two
The useful split is between the design layer and the measurement layer.
The design layer is the reusable part: the validation-set construction protocol (inclusion and exclusion criteria, scanner and vendor mix targets, patient-level partitioning), the ground-truth adjudication procedure (reader count, information conditions, disagreement resolution), the reporting structure (which slices are reported and in what order), and the drift-telemetry design (which strata are monitored against which thresholds).
The measurement layer is everything that is site-specific by definition: the actual cohort assembled at that hospital, the numbers produced on that hospital’s scanner fleet and acquisition protocols, the local prevalence, and the per-slice performance table that follows.
Teams that never separated those two layers re-litigate methodology at every site while the procurement clock runs. That is the whole failure mode, and it is structural rather than clinical — the pack has no seam along which local work can be swapped in, so every local question reopens the entire document.
What actually happened at each site
The sequence below is an observed pattern from clinical-imaging engagements where a validation pack was carried across multiple hospital reviews. The counts are engagement observations, not a published benchmark, and the question-round numbers are specific to these reviews rather than a rate you should expect to reproduce.
| Site | Review outcome | Question rounds | Methodology questions vs site-specific | What had to be redone |
|---|---|---|---|---|
| Site 1 (design site) | Cleared after full methodology defence | 4 | Mostly methodology — the protocol was being adjudicated for the first time | Everything: protocol written, adjudication procedure defined, reporting slices chosen |
| Site 2 | Cleared | 2 | Predominantly site-specific — one scanner-vendor gap, one prevalence question | Measurement layer only: local cohort assembled to the same protocol, slices re-reported |
| Site 3 | Cleared, one section escalated | 2 | Site-specific, plus one genuine protocol amendment | Measurement layer, plus an added exclusion criterion for a paediatric sub-population the protocol had not contemplated |
Two things fall out of that table. First, the drop from four rounds to two happened because the reviewer at site two was handed a structure that had already been defended somewhere else, so the conversation moved to their data rather than to our method. Second, the escalation at site three was not a failure of the pack — it was the pack doing its job, surfacing a population the construction protocol did not cover and forcing an explicit amendment rather than a silent extrapolation.
Why does the second reviewer ask different questions than the first?
Because the first reviewer is adjudicating a method and the second is adjudicating a fit. The first review asks whether the evidence-generation process is sound at all: who read the cases, how disagreements were resolved, whether the split leaked at patient level. Once that has been settled and written down, the second reviewer largely accepts the process and pushes on the part that is theirs — their scanner mix, their case mix, their prevalence. Sequencing matters here more than teams expect: the effort spent making the construction protocol reviewable at site one is what converts site two into a re-measurement exercise instead of a second methodology trial. We treat the protocol itself as an artefact in the validation-set construction protocol as a reviewable artefact.
Re-instantiating the protocol on a different scanner fleet
The mechanical work at site two was narrower than it sounds. The protocol specified target proportions across scanner vendors, field strengths and acquisition protocols, plus patient-level partitioning and a prevalence statement relative to the deploying population. Re-instantiating it meant querying the local PACS against those same criteria and reporting where the local fleet could not hit a target proportion — site two had no representation of one vendor present in the design cohort, so the pack carried an explicit coverage gap rather than a quiet substitution.
That declared gap is worth more in a review than a filled cell would have been. A reviewer who sees the gap named, bounded, and tied to a claim restriction reads the rest of the numbers with more confidence, not less. In our experience this is the single highest-leverage habit in site-to-site work: state the coverage the local cohort does not have, before someone asks.
Drift telemetry from site one also changed the conversation at sites two and three. Having roughly nine months of monitored score distributions and adjudicated flagged cases from a live deployment let the pack answer a prospective question with prospective evidence instead of a promise. The mechanics of that loop — which strata are monitored, who adjudicates a flag, how the result is versioned back — are covered in post-deployment drift evidence in a clinical imaging validation pack. Practically, it meant reviewer questions about degradation resolved in one round rather than three.
What travelled and what did not
| Pack section | Travelled unchanged? | Note |
|---|---|---|
| Validation-set construction protocol | Yes, with one amendment at site 3 | Design layer; the amendment was additive, not a rewrite |
| Ground-truth adjudication procedure | Yes | Reader count and disagreement resolution held across all three reviews |
| Reporting structure (slices, order) | Yes | Same table shapes, different values |
| Drift-telemetry design | Yes; data accumulated | Design carried; the evidence grew stronger per site |
| Local cohort and performance numbers | No — re-measured every time | Measurement layer by definition |
| Prevalence and claim scoping | No — re-earned per site | Any “performs acceptably here” claim is local |
| Workflow and PHI-handling evidence | No | Each site’s own lawful basis, BAA and infrastructure review |
Portability stopped at the claim itself. The pack’s structure is transferable; the sentence “this model performs acceptably on this population, in this workflow” is not, and should never be written as if it were. That distinction underpins how we approach production AI reliability generally — the reusable asset is the evidence machinery, not the conclusion it produced last time.
The other thing that did not travel: governance evidence. Site transfer reliably triggered a fresh HIPAA and workflow-evidence request alongside the performance pack, on each site’s own terms. Budget for it as a parallel track rather than an annex.
How the effort was measured
Three counts, recorded per review, and worth instrumenting from site one:
- Question rounds — discrete reviewer response cycles until sign-off.
- Question composition — each question tagged as methodology (already answered in the pack) or site-specific (genuinely new). A rising methodology share across sites means the pack’s structure is not holding.
- Portability failure rate — how often a reviewer rejected or escalated a section because the construction protocol did not map to their population. One escalation in three reviews here, and it produced a protocol amendment rather than a rejection.
Re-measurement effort versus re-design effort is the summary number procurement stakeholders care about. Across sites two and three, the bulk of validation work was re-running the measurement layer on local data. That is the whole return on writing the design layer down properly the first time.
Where this leaves an open question: we do not know how far the protocol amendment cadence flattens. Site three added a sub-population the protocol had not contemplated; whether site six adds another, or whether the protocol converges, is not something three reviews can tell you.
Frequently Asked Questions
ROI: what does a clinical imaging validation pack carried across site-to-site procurement mean in practice? How does validation translate when the same CT protocol runs at two different hospital sites? Clinical Imaging Validation Pack has one honest answer. Looked at closely, Clinical Imaging Validation Pack is this. Clinical Imaging Validation Pack has one honest answer. Clinical Imaging Validation Pack answers cleanly when you separate two things. Looked at closely, Clinical Imaging Validation Pack is this. Clinical Imaging Validation Pack has one honest answer. Looked at closely, Clinical Imaging Validation Pack is this. Clinical Imaging Validation Pack has one honest answer. Clinical Imaging Validation Pack answers cleanly when you separate two things. Looked at closely, Clinical Imaging Validation Pack is this. Clinical Imaging Validation Pack has one honest answer. Looked at closely, Clinical Imaging Validation Pack is this. Clinical Imaging Validation Pack has one honest answer. Clinical Imaging Validation Pack turns on one distinction. Clinical Imaging Validation Pack answers cleanly when you separate two things. It means the design layer — construction protocol, adjudication procedure, reporting structure, drift-telemetry design — is reused verbatim, and only the measurement layer is regenerated on local data. In the reviews described here, that took question rounds from four at the design site to two at each subsequent site.
Which sections of the pack travelled unchanged between sites, and which had to be re-measured on local data? The construction protocol, adjudication procedure, reporting structure and telemetry design travelled unchanged. The local cohort, all performance numbers, the prevalence statement and the site-scoped claim were re-measured or re-earned every time.
What did the second and third site reviewers ask that the first reviewer did not? They asked fit questions rather than method questions — scanner-vendor coverage, local prevalence, and at site three a paediatric sub-population the protocol did not cover. That revealed the pack’s structure was holding: the methodology defence had been absorbed, leaving only genuinely local ground to argue.
How was the validation-set construction protocol re-instantiated on a different scanner fleet and case mix without redesigning the methodology? The protocol’s criteria were run against the local PACS to assemble a new cohort to the same targets, and any target the local fleet could not meet was reported as a declared coverage gap with a corresponding claim restriction — no substitution, no redesign.
How did drift telemetry from the first deployed site strengthen the pack presented at the next site? It replaced a promise with prospective evidence: monitored score distributions and adjudicated flagged cases from a live deployment. Degradation questions then closed in one round instead of several.
How were procurement question rounds and rework effort measured across the three reviews? By counting discrete reviewer response cycles per site, tagging each question as methodology or site-specific, and tracking how often a section was escalated for population mismatch — the portability failure rate.
Where did portability genuinely stop — which claims had to be re-earned per site rather than carried? At the performance claim itself. “Performs acceptably on this population, in this workflow” is local by construction, as are prevalence statements and each site’s PHI-handling and workflow evidence.
Cross-site deployment changes everything
A validation pack that worked at Site A will surface entirely new failure modes at Site B, and your documentation strategy must anticipate that from day one. Revisit it when your workload shifts.