Worked Example: An LLM Procurement Evidence Pack Approved in One Committee Round

An outcome study of an LLM procurement evidence pack that closed committee approval in a single sitting

Worked Example: An LLM Procurement Evidence Pack Approved in One Committee Round
Written by TechnoLynx Published on 01 Sep 2026

The committee had already deferred once. That is the useful starting condition for this worked example, because a second deferral is not a scheduling problem — it is the point at which the buying team loses control of the timeline and the vendor’s renewal calendar starts driving the decision instead. The second attempt was approved in one sitting. What changed was not the model shortlist; it was the ordering of the artefact.

The first attempt had been a vendor comparison deck: three candidate models, public leaderboard positions, a pricing table, a recommendation on the last slide. The questions came live, and none of them were unreasonable. What is accuracy on our own tickets? What does it do when it is wrong? What does one decision cost at our volume? What happens when the vendor ships a new version in March? Each answer was available in principle and absent in the room, so the committee did the correct thing and asked for another round.

The second pack inverted the order. Instead of presenting work and answering questions afterwards, we enumerated the questions the committee was obliged to ask — accuracy on the buyer’s prompt distribution, failure behaviour at the buyer’s risk tolerance, cost-per-decision at the buyer’s real load, drift posture against vendor versioning — and built one traceable section per question. The divergence point between a one-round approval and a defer-and-return cycle is the first question a reviewer has to take on trust. Once that happens, the rest of the pack is being read sceptically rather than being read as evidence.

What was in the pack, and which objection each section retired

The pack was five sections and an appendix. It is worth being precise about which section carried the decision, because two of the five were never opened in the meeting.

Section Committee question it closed Did it move the decision?
Task accuracy on 1,140 sampled production prompts, scored against a written rubric “Is it accurate enough on our task?” Yes — the deciding section
Failure-mode catalogue: 9 reproducible triggers, each with severity and containment “What happens when it’s wrong?” Yes — one entry forced a scope change
Cost-per-decision model at measured load, with assumptions stated inline “What does this cost at our volume?” Partly — accepted after one assumption was challenged
Drift posture: re-validation trigger list and re-run scope per trigger “What happens at the next vendor release?” Yes — became an approval condition
Shortlisting rationale (leaderboards, suite scores) “Why these three candidates?” Never asked
Appendix: full prompt set, scoring sheets, reviewer identities “Show me behind the number” Opened once, for two rows

The shortlisting section is the interesting negative result. It had been the entire first attempt, and in the second round nobody asked about it. Public leaderboard positions were doing the job they can actually do — narrowing a field of candidates before real work starts — and once the pack answered the organisation-specific questions, the ranking was simply not what the committee was deciding on. We see this pattern regularly enough that it now shapes how we allocate effort: shortlisting evidence gets a page, not a deck.

How was task accuracy presented so the models could be compared like-for-like?

The accuracy section rested on one decision made before any model was run: the prompt distribution was sampled from the buyer’s own production traffic, not written for the evaluation. Twelve months of tickets were stratified by intent and by handling risk, and 1,140 prompts were drawn so that low-frequency high-risk intents were deliberately over-sampled relative to their natural rate — with the over-sampling factor stated in the table header so the committee could read headline accuracy and risk-weighted accuracy side by side.

All three candidates ran the identical prompt set, with identical system prompts, identical retrieval context and identical temperature, and were scored by the same two human reviewers against a four-level rubric written before scoring began. Inter-reviewer disagreement was 6.1% of items and every disagreement was listed rather than averaged away (project-specific measurement from this engagement, not a published benchmark). That last detail mattered more than the accuracy numbers themselves. A committee member’s first probe was whether the scoring was reproducible by someone else; the disagreement log was the answer, and it took thirty seconds.

The comparison table showed three numbers per model — overall rubric pass rate, pass rate on the highest-risk intent stratum, and pass rate on the long tail of rare intents. The spread between models on the overall figure was small. The spread on the rare-intent stratum was not, and that is where the recommendation actually came from. A single blended score would have hidden it, which is roughly the same failure a public leaderboard commits at larger scale — the reason leaderboard rankings stall procurement committees is that they average over exactly the distribution the buyer cares about.

Cost-per-decision, and the assumption that got challenged

Cost was expressed as cost per resolved decision, not cost per million tokens. The conversion needed four stated assumptions: mean input and output token counts measured from the accuracy run rather than estimated; the retry rate observed when the model’s confidence gate rejected an output; the proportion of decisions escalated to a human reviewer; and the load profile — peak-hour concurrency taken from the existing ticketing system, not a smoothed daily average.

The escalation rate was challenged, and correctly. The pack had used the escalation rate observed during evaluation, which was measured with the reviewers paying close attention; the committee’s operations lead argued live handling would escalate more. Rather than defend the number, we had already published a sensitivity band alongside it, so the answer was to read across the table to the pessimistic column, where the cost conclusion still held. Stating assumptions inline is what let a challenged number survive without a second round. Had the sensitivity band been in an appendix, the meeting would have gone to “come back with revised figures”.

Where the committee still pushed back

Two concessions were made before approval, and an outcome study that omitted them would be worthless.

The failure-mode catalogue contained an entry where the model produced a confidently-worded but incorrect entitlement statement — low frequency, high consequence, and containment relied on a downstream check that did not yet exist. The committee approved the model and refused the original scope: the highest-risk intent stratum was carved out of automated handling for the first quarter, pending the containment control being built and evidenced. That is a better outcome than approval-as-requested, but it is a scope reduction, and the pack is what surfaced it early rather than in production.

The second pushback was on drift. Our proposed re-validation trigger list was accepted, but the committee added a calendar floor — re-evidence every six months whether or not the vendor ships anything — on the grounding that a silent endpoint change would otherwise go unnoticed. The mechanics of that re-run scope are a discipline of their own, covered in re-validating an evaluation pack when the vendor ships a new version.

What the pack cost against what it avoided

Assembly took roughly nine working days: four on prompt sampling and scoring, two on the failure-mode reproduction, one on the cost model, one on drift, one on assembly and internal review. The avoided delay was one review cycle at minimum — the committee met monthly — and plausibly two, since the first deferral’s follow-up request had itself been under-specified. What was measured is the number of approval rounds; what was not measured is any counterfactual about deployment outcomes, and we are careful not to claim the pack made the model choice correct. It made the choice defensible on a stated date, which is a narrower and more honest claim.

The reuse value showed up four months later. When the vendor shipped a point release, the prompt set, rubric, reviewer instructions and cost model were re-run rather than rebuilt, and the re-validation consumed under two days for the affected sections. That amortisation is where the nine days pay for themselves, and it is the argument for treating the pack as a versioned artefact rather than a one-off procurement deliverable — the broader structure sits in our work on AI governance and trust.

One boundary is worth naming plainly. The pack applied benchmark methodology; it did not define it. Measurement construct, task taxonomy design and reproducibility standards belong to the LynxBenchAI methodology line, and conflating the two produces an artefact that satisfies neither a methodologist nor a committee.

What we still do not know from a single engagement is how much of the one-round result is transferable and how much came from a committee that had already been through a deferral and knew exactly what it wanted to see. A committee’s second sitting is a friendlier audience than its first. The honest open question is whether the same pack structure closes a first sitting — and that is a question about the committee’s prior experience, not about the evidence.

Frequently Asked Questions

ROI: what does a procurement evidence pack unblocking committee approval in a single round mean in practice? A common Worked Example question is worth clarifying. It means the measured outcome is approval rounds, not model quality. In this engagement a model-selection decision that was expected to take two to three monthly review cycles closed in one sitting, because every question the committee was obliged to ask already had a traceable section. Nine working days of assembly bought back at least one, probably two, monthly cycles.

What was in each section of the pack, and which specific committee question did each section close? Five sections: task accuracy on the buyer’s sampled production prompts, a nine-entry failure-mode catalogue with containment, a cost-per-decision model at measured load, a drift posture with re-validation triggers, and a short shortlisting rationale. Each mapped to one committee question — accurate enough on our task, what happens when it fails, what it costs at our volume, what happens at the next vendor release, and why these candidates.

Which pieces of evidence actually moved the decision, and which were prepared but never asked for? The accuracy stratification and the failure-mode catalogue carried the decision; the drift posture became an approval condition. The shortlisting section — leaderboard positions and suite scores, which had been the whole of the first failed attempt — was never opened. The appendix was opened once, for two disputed scoring rows.

How was task-specific accuracy on the buyer’s own prompt distribution presented so the committee could compare candidate models like-for-like? Prompts were sampled from twelve months of the buyer’s production traffic, stratified by intent and handling risk, with high-risk intents deliberately over-sampled and the factor disclosed. All candidates ran the identical set under identical conditions, scored by the same two reviewers against a rubric fixed before scoring, and results were reported as three separate figures rather than one blended score.

How was cost-per-decision at the buyer’s real load shown, and what assumptions had to be stated for the committee to accept it? Cost was expressed per resolved decision rather than per million tokens, which required four stated assumptions: measured token counts, observed retry rate, human escalation proportion, and peak-hour concurrency from the existing ticketing system. Each was published inline with a sensitivity band, which is why the challenged escalation-rate assumption did not trigger a second round.

Where did the committee still push back, and what was added or conceded before approval? Two concessions. The highest-risk intent stratum was carved out of automated handling for a quarter until a containment control for one failure mode existed and was evidenced. And the committee added a six-month calendar floor for re-evidencing on top of the proposed vendor-triggered re-validation list.

What did the pack cost to assemble against the review cycles it avoided, and how was the baseline reused at the next vendor version review? Assembly was roughly nine working days against at least one avoided monthly review cycle. Four months later the vendor shipped a point release and the prompt set, rubric, reviewer instructions and cost model were re-run rather than rebuilt, consuming under two days for the affected sections — which is where the original effort amortises.

How this evidence pack structure won approval

Committee sign-off accelerated once we presented cost per query alongside benchmark scores, included vendor contractual commitments in writing, and demonstrated fallback options for each shortlisted model. Everything else is detail.

Back See Blogs
arrow icon