A cost comparison between the NVIDIA H100 80GB HBM3 and the NVIDIA H200 has two halves, and only one of them is a measurement. Price moves by region, contract, and availability — LynxBenchAI does not measure it at all. The published 26Q3 results do measure both devices, in one release, on a stated backend, driver, framework, and runtime. So the price half comes from your own quotes; the performance half comes from the catalogue. The comparison is only as honest as what you put in the denominator.
What does “H100 vs H200 price” actually mean in practice?
The naive version: take a street price or an hourly cloud rate for each card, divide by one headline benchmark figure, and call the smaller number the winner. That single figure collapses the Training, Inference, and Compute category scores into a blend, which is exactly the collapse the parent analysis of GPU specs versus real AI performance warns against. A blended denominator rewards whichever device looks stronger on the categories your workload does not resemble.
The expert version keeps the halves separate. Pick the category that matches the job — Training if you are fine-tuning, Inference if you are serving, Compute if you are running kernels or simulation — and divide your own quoted cost by that score. Record the release name (26Q3) and the runtime alongside the ratio, because the ratio expires when either changes.
Building the denominator honestly
| Step | Do | Do not |
|---|---|---|
| 1. Pick the category | Choose Training, Inference, or Compute by workload resemblance | Average the three into one “performance” number |
| 2. Pull both scores | Both from the same release name (26Q3), same architecture family | Mix a 26Q3 score with a score from another release |
| 3. Check the executor | Note backend, driver, framework, runtime for each score | Assume the quoted card ships the same software stack |
| 4. Scope the numerator | Match unit to unit: single accelerator cost against single-accelerator score | Divide a DGX-class node quote by a per-device score |
| 5. Stamp it | Write the release name next to the ratio | Reuse last quarter’s ratio after a release rolls |
Two consequences follow. First, cloud hourly rates and purchase prices are not interchangeable numerators: an hourly rate bundles a provider’s own backend, driver, and container image, which may not be the executor the published score was produced on. Second, a used or secondary-market quote tells you nothing the published result covers — the 26Q3 figure says nothing about the condition, firmware, or cooling of the specific unit being priced, so treat the discount as risk-adjusted, not free.
Regional price gaps are simpler than they look. The same card quoted differently in two markets changes only the numerator. The category score is unchanged, which is precisely why a per-category denominator is worth building once and reusing across quotes.
When the cheaper H100 genuinely wins
If your workload maps cleanly to one published category, a lower H100 price can beat an H200 uplift outright — that is a legitimate result, not a compromise. It stops being legitimate the moment the uplift you are discounting was measured on a category your workload does not resemble. And if the workload resembles none of the published categories, the honest outcome is that no catalogue-derived price ratio applies to it; you need your own run, on your own executor, before the arithmetic means anything. The same holds for next-generation parts: with no same-release result for a newer device, an H100/H200 ratio cannot be extended to cover it.
The methodology behind the category split, and how far a published score can be read, sits in the LynxBenchAI benchmarking work and in our GPU performance measurement material.
Before you divide anything: which category does your workload actually resemble, and can you name the release and runtime the score you are about to use came from?