Most LLM evaluation packs are organised around the work that was done rather than the decision they exist to support: a methods section, a results table, an appendix of prompts. That ordering is honest and it is also the reason so many packs come back from committee with a request for “a summary”. A committee reading a research write-up has to reverse-engineer answers to its own questions out of someone else’s narrative, and it will not finish that job in a sixty-minute slot.
Invert the ordering and the pack behaves differently. One section per question the committee is chartered to ask, with the evidence subordinated underneath each. The question is the heading; the numbers are the support.
What does structuring the pack around a committee’s questions mean in practice?
It means the table of contents is written before the evaluation is written up, and it is written from the committee’s charter rather than from the evaluation plan. In practice a technology approval committee asks a small, stable set of things about a model choice, and they recur almost verbatim across buyers:
- Is this model accurate enough on our task, measured on our own prompt distribution?
- What happens when it fails, and can we live with those failures?
- What does it cost at our real load, per decision rather than per token?
- What happens when the vendor ships a new version of the model?
- Who signed off on the test set, the thresholds, and the residual risk?
Those five become five fixed section headings. Everything the evaluation produced is then filed underneath the question it answers, and anything that answers no question is moved to an appendix or dropped. The mapping between evidence and decision criteria is work that has to happen somewhere; the only choice is whether it happens in the document or live in the room. Done in the room, it is done incompletely and under time pressure, and the model choice ends up defended from memory rather than from an artefact.
We see the same tell repeatedly in governance reviews: a pack with a strong methods section and a weak or missing “who signed off” section will generate follow-up requests regardless of how good the underlying measurement was. The evaluation was not the problem. The shape was.
The section skeleton
The skeleton below is what we assemble a procurement evidence pack into. It is deliberately boring and deliberately fixed — the reuse value comes from it not changing between model reviews.
| Committee question | Section heading | What sits under it | Explicit non-coverage statement |
|---|---|---|---|
| Accurate enough on our task? | Task accuracy on our prompt distribution | Named prompt set, sample counts, scoring rubric, per-intent breakdown, threshold and pass/fail | Which intents were under-sampled; which tasks were not tested at all |
| What happens when it fails? | Failure-mode catalogue at our risk tolerance | Reproducible triggers, severity per failure, containing control, residual risk | Failure classes observed but not characterised |
| What does it cost at our load? | Cost per decision under our load profile | Load assumptions, token accounting, latency at target concurrency, cost sensitivity | Load levels not exercised; pricing assumptions with an expiry |
| What when the vendor updates? | Drift posture and re-validation triggers | Trigger list, re-run scope per trigger, owner, cadence | Behaviour changes the trigger list would not catch |
| Who owns this decision? | Sign-off, ownership and audit trail | Named approver per section, test-set provenance, versioning, dated review | Approvals still outstanding at time of writing |
Two properties matter more than the exact wording. First, every section carries an explicit statement of what was not tested — a pack that only reports what it covered invites the committee to guess at the gaps, and committees guess pessimistically. Second, every claim in a section has a page reference back to the evidence, so a member who wants to check one number can do it without reading the whole document. Both of these are structural, not analytical; they cost nothing to add and they remove the two most common reasons for deferral.
The per-criterion scoring that lives inside each section — thresholds, weightings, pass/fail logic — is a separate concern with its own discipline, and we treat it as a scorecard nested under the section rather than as the section itself. For a fuller account of what belongs inside an evidence pack at all, rather than how it is ordered, see what sits beyond the leaderboard in an LLM evaluation evidence pack. The evidence requirements per section — which prompts, how many, scored by whom — are set out in the per-section evidence checklist.
Readability without weakening the evidence
Committees are mixed audiences. A general counsel, a CFO’s delegate and a platform architect read the same pack and need different depths from it. The structure resolves this better than plain-language rewriting does: each section opens with a two-to-three sentence answer in the committee’s own vocabulary, then a stated confidence and caveat, then the evidence. A non-technical reader stops after the first paragraph of each section and still has the complete decision. A technical reader continues.
What does not work is simplifying the evidence itself. Rounding a confidence interval away, collapsing a per-intent accuracy breakdown into a single headline number, or dropping the sampling method to save a page — each of those moves the pack’s weakest point from “hard to read” to “unable to withstand a question”. Keep the evidence intact and use layering, not deletion.
Where the structure shifts for a regulated buyer
The skeleton holds, but the weight redistributes. For a regulated buyer the sign-off and audit-trail section stops being the last section and becomes a load-bearing one: test-set provenance, third-party reproducibility of the test set, a named owner and documented control per failure mode, and a drift posture that commits to a re-evidencing cadence rather than an intention to re-run when the vendor updates. The accuracy and cost sections change less than people expect. What changes is who has to be able to verify them, and how long after the fact.
The reciprocal case is worth naming: an unregulated buyer who imports the regulated pack’s evidentiary standard usually stalls the procurement rather than strengthening it. Match the standard to the obligation.
Re-running the same skeleton on a new model version
The skeleton’s second job is comparability over time. Because the headings are fixed, a vendor version bump does not require a new pack — it requires the affected sections re-run and re-dated, with the prior figures kept alongside for comparison. In our experience the sections that move on a point release are task accuracy and the failure-mode catalogue; cost and sign-off often survive unchanged. Fixed headings are what makes that selective re-run legible to a committee that approved the previous version: they are reading a diff, not a new argument.
That is also the reusability argument for the structure. The same headings carry into the next model-vendor review, so the second comparison is like-for-like instead of re-argued from scratch. The measurable outcomes we track against this are approval rounds per model decision, the count of committee follow-up requests raised against a pack, and the time to re-run the pack against a new model version.
Where the pack stops and benchmark methodology begins
The pack’s structure is a decision artefact. It defends one buying decision to one committee, on one date. It is not a measurement methodology, and conflating the two produces a document that neither survives methodological scrutiny nor answers the approval questions. How a model should be scored — measurement construct, task taxonomy, sustained-load design, reproducibility — belongs to benchmarking methodology, which is LynxBenchAI’s territory rather than ours. We build the evidence pack; LynxBenchAI defines the methodology the pack applies. That division is the same one our wider work on AI governance and trust is built around, and keeping it visible inside the pack is itself a credibility feature — a committee can see which claims rest on the buyer’s own measurement and which rest on an external method.
The open question we have not resolved cleanly is how far a committee-shaped skeleton should bend to a committee that has never written its charter down. Where the questions are tacit, the first version of the pack ends up proposing them — and a pack that defines the criteria it is then judged against is a structure worth being careful with.
Frequently Asked Questions
What does structuring the pack around an approval committee’s questions mean in practice? Your approval committee’s questions should dictate how you organize every section of your LLM evaluation pack. It means writing the table of contents from the committee’s charter before writing up the evaluation, so each question the committee must ask becomes a fixed section heading with evidence filed underneath it. Anything that answers no committee question moves to an appendix. The mapping from evidence to decision criteria happens in the document rather than live in the meeting.
Which questions does an approval committee actually ask about an LLM choice, and how are they turned into fixed section headings? The recurring set is accuracy on our task, behaviour on failure, cost at our load, exposure to vendor version changes, and ownership of the sign-off. Each becomes a heading phrased in the committee’s vocabulary — “Task accuracy on our prompt distribution”, “Failure-mode catalogue at our risk tolerance”, and so on — and stays fixed across successive model reviews.
What belongs in each section — evidence, caveats, and the explicit statement of what was not tested? All three. Each section carries its evidence with a page reference, the caveats that bound it, and an explicit statement of what was not tested. Omitting the non-coverage statement invites the committee to guess at the gaps, and committees guess pessimistically.
How do you keep the pack readable for non-technical committee members without weakening the evidence underneath? Layer rather than simplify: open each section with a two-to-three sentence answer in the committee’s own language, then the confidence and caveat, then the full evidence. Never round away intervals, collapse per-intent breakdowns, or drop the sampling method — that converts a hard-to-read pack into one that cannot withstand a question.
How should the structure change for a regulated buyer, where sign-off and auditability sections carry more weight? The skeleton holds but the weight moves: sign-off and audit trail become load-bearing, requiring test-set provenance, third-party reproducibility, a named owner and control per failure mode, and a committed re-evidencing cadence. Accuracy and cost sections change less than expected; what changes is who must be able to verify them later.
How is the same section skeleton re-run when the model vendor ships a new version, so comparisons stay like-for-like? Because the headings are fixed, only the affected sections are re-run and re-dated, with prior figures retained alongside. Task accuracy and the failure-mode catalogue typically move on a point release while cost and sign-off often survive, so the committee reads a diff against the version it already approved rather than a new argument.
Where does the pack’s structure stop and benchmark methodology (LynxBenchAI) begin? The pack defends one buying decision to one committee on one date; it is a decision artefact, not a measurement methodology. Measurement construct, task taxonomy, sustained-load design and reproducibility sit with LynxBenchAI. We build the evidence pack; LynxBenchAI defines the methodology it applies.
Organize evidence to answer gatekeepers directly
Committees ask the same eight questions in every cycle, so structure your pack around their order of concern rather than your own testing sequence. If Structuring LLM Evaluation Pack is on your roadmap, the next step is to map it onto your own constraints rather than copy a reference architecture.