“Fourteen times faster” is not a complete comparison. A speedup is a derived, relative measure: one measured time or throughput divided by another. To judge it, you need the baseline, the conditions, the statistic, and enough published numbers to repeat the division.
Three August 2026 claims make the problem concrete. FastVideo's 14.38× result compares one video model with its own dense baseline under named conditions. OpenAI's “up to 14×” preview did not publish the baseline workload. TypeScript 7 produces several defensible multiples on the same codebase when the operation or parallelism changes. The number is useful only with the comparison attached.
- 01A speedup is a relative measure, not a standalone score.It is derived from two measurements. Change the baseline or test conditions and the ratio changes even when the new system does not.
- 02Identical headline numbers can describe unrelated tests.FastVideo's 14.38× video result and OpenAI's “up to 14×” text-inference preview use different workloads, baselines, and disclosure levels.
- 03One codebase can support several honest multiples.Published TypeScript 7 results on vscode range from 8.0× to 16.7× depending on whether the test measures project load, a build, or checker count.
- 04Ask four questions before comparing.What is the baseline? Under which conditions? Which statistic over how many runs? Does the ratio recompute from the published numbers?
01 — The FortnightThree speedup claims, one fortnight, three different measurements.
Start with the coincidence, because it does the teaching for free. In a fourteen-day window in August 2026, three parties published speed claims about AI inference. Two of them used the same headline number. None of the three is comparable to either of the others, and the reasons differ in each pairing.
On August 13, 2026, OpenAI previewed an Ultrafast tier for GPT-5.6 Sol and described it as up to fourteen times standard speed. We covered that preview when it landed, and the relevant detail for this post is a negative one: our write-up of the Ultrafast preview records no disclosed baseline workload, no prompt set and no reasoning-effort setting. Even the Standard-tier figure that the multiple would have to be measured against is an inference we drew from two separate vendor numbers, and that post labels it as our inference rather than a vendor statement. We are not upgrading that hedge here.
Then, on August 27, two claims arrived on the same day about the same open-weights base model. FastVideo — the Hao AI Lab team at UC San Diego, working with Nuva Lab and the NVIDIA FastGen team — released FastH3 Preview v1, a four-step distilled derivative of MiniMax’s H3-Base weights, headlined “Up to 14x speedup on a single Nvidia Blackwell GPU”. The same day, the inference platform fal announced H3 Max, another post-trained derivative of the same base model, at roughly thirty-five times the throughput of its own hosted official MiniMax H3 endpoint.
Two of them used the same headline number. None of the three is comparable to either of the others, and the reasons differ in each pairing.Digital Applied analysis, August 30, 2026
FastH3 Preview v1 — fully disclosed
The denominator is named in the same table as the ratio: Base H3 with dense FA4 attention, run by the same team on the same hardware for the same clip duration. The conditions are named too — one B200, the 15-second shape, the VSA / Data-Free checkpoint, the median of three timed requests after one full warmup, model loading and compilation excluded, and an end-to-end figure that includes encoding, denoising, decoding, audio, muxing and file output. This is the good example in this post.
GPT-5.6 Sol Ultrafast — no baseline stated
The same headline number as the row above, attached to an entirely different kind of measurement: text inference rather than video generation. Our own coverage of the preview records no disclosed baseline workload, no prompt set and no reasoning-effort setting, which means the multiple cannot be located anywhere on a taxonomy of denominators. It is not that the number is wrong. It is that there is nothing published to divide by.
H3 Max — same base model
A different post-trained derivative of the same open-weights base model as the first row, announced the same day, and still not comparable to it. The denominator is named — fal's own hosted official MiniMax H3 endpoint, so fal controls both sides of the ratio — but the clip length, warmup, exclusions, statistic and sample count behind the 35× are not published. H3 Max was trained and served on GB200 NVL72, which is a serving fact rather than the baseline.
Read down that list and the disclosure spectrum is the finding, not the numbers. Rows one and three share a base model, a publication date and a modality, and are still not comparable, because the clip length, the hardware, the method and the baseline all differ. Rows one and two share a headline number and nothing else — one is a video-generation pipeline timed end to end, the other is a text inference tier. If you saw all three in a feed on August 28 you would have had no way to rank them, and ranking them was never a coherent operation in the first place.
One correction belongs here rather than in a footnote, because it is the most common error in circulation about the first row. FastH3 Preview v1 is FastVideo’s release, not MiniMax’s. The primary’s own citation block gives the author as the FastVideo Team; its acknowledgements thank MiniMax for releasing H3-Base, which is the posture of a team crediting the weights it post-trained rather than an author list. The Hugging Face card for the recommended checkpoint calls it the FastH3 Preview v1 checkpoint from FastVideo. MiniMax published the base weights and the community licence the derivative inherits; it did not publish this speedup.
The reason to spend a paragraph on attribution in a post about arithmetic is that the two errors are the same error. On August 28, 2026, the aggregator Digg published the story under the headline “MiniMax Announces Fast H3 v1 Video Model”, with a dek reading “MiniMax celebrates release of Fast H3 v1 video model with partner labs” and, per the page’s own counter, 253.4K combined views. Its body text is closer to right — it reports that MiniMax posted congratulations to the Hao AI Lab, Nuva Lab and NVIDIA FastGen teams — so this is a headline compressing a congratulation into an announcement, not a fabrication. Note that Digg is a rolling aggregator and the page may since have changed; the version quoted here is the one fetched on August 30, 2026. A headline that compresses “MiniMax congratulated the labs that built this” into “MiniMax announces” is doing exactly what a headline does when it compresses “14.38× at the 15-second shape, on one B200, against our own dense baseline” into “14x faster”. Same operation, same information loss.
Rows one and two share a headline number and nothing else — one is a video-generation pipeline timed end to end, the other is a text inference tier.Digital Applied analysis, August 30, 2026
02 — The ArithmeticA speedup is a ratio, and only the top half ever gets published.
The structural point is almost embarrassingly simple. A benchmark score is a measurement: one number, one system, one set of conditions. A speedup is a derived relative measure — a division. It requires two measurements, and the second one is a choice the claimant makes. Change the choice and the number changes without anything about the product changing at all. That is why the usual reading advice does not transfer: checking whether a benchmark was run honestly tells you nothing about whether a ratio’s baseline was well chosen, because a perfectly honest measurement divided by a favourable baseline still produces an incomparable multiple.
Our archive already covers the neighbouring questions, and the division of labour is worth stating so you can go to the right place. Reading a vendor benchmark table that shows a loss works through four levers inside a quality table — version, subset, metric shape, and who got a column. This post is about the fifth lever that table cannot show you: the denominator of a ratio. Our reproducibility audit of 42 vendor benchmark rows asked whether a buyer could re-run what a vendor published. And the voice-AI benchmark read happens to hold a multiple that moves with conditions in a completely different modality, which is a useful reminder that this is not a video problem or a compiler problem.
Before going further it is worth being precise about what this article can and cannot claim, because a sloppy absence claim in a post about rigour would be self-refuting. “Nobody publishes the denominator” is false, and it is not a sentence we will write. FastVideo publishes its denominator in the same table as its ratio, under a column labelled as the speedup over Base H3. Astral links its benchmark document from the headline sentence of the uv README. The honest version, scoped to what was actually read, is narrower and more useful.
Across the speedup claims examined for this article — FastVideo’s FastH3 Preview v1, fal’s H3 Max as recorded in our own August 27 coverage, Microsoft’s TypeScript 7 GA and 2025 native-port posts, Astral’s uv README and benchmark document, vLLM’s Blackwell and InferenceMAX post, and NVIDIA’s TensorRT-LLM deployment guide, each read on August 30, 2026 — every headline stated a multiple, and the matched run that produced it lived somewhere below the fold. Some headlines name a class of baseline (fal names its own hosted H3 endpoint; uv names pip; vLLM names Hopper GPUs). None of those eight put the matched conditions in the headline. The failure is in the compression, not in the disclosure. We did not survey the field, and this is what those eight documents show. Nothing in this post asserts a frequency beyond them.
That is a genuinely more interesting finding than a disclosure scandal would have been. It means the information you need mostly exists, and the work is retrieval rather than investigation — which in turn means a reader with four questions and five minutes can usually resolve a claim that looked unresolvable. Section 09 is those four questions.
03 — The Worked PairThe same base model, the same day, and nothing else in common.
This is the one section in this post whose subject is video generation. It earns the space because the pair is unusually clean: two derivatives of one open-weights base model, two speedup claims, one calendar day, and a complete disclosure asymmetry between them. Everything after this section is compilers, package managers, inference engines and standards documents, because the arithmetic does not care about the modality.
Side A is FastH3 Preview v1 from FastVideo. Side B is H3 Max from fal, which we covered on the day — our August 27 write-up of H3 Max has the launch detail, and this post takes only three multiples, the shape of the comparison set, and the serving hardware from it. The base model itself is MiniMax H3, which launched on July 31, 2026 with native audio; both sides post-trained it, and neither of them is MiniMax.
| Condition | Side A — FastH3 Preview v1 | Side B — H3 Max | Disclosed by |
|---|---|---|---|
| Publishing party | FastVideo / Hao AI Lab at UCSD, with Nuva Lab and the NVIDIA FastGen team | fal, an inference platform | Both |
| Announcement date | August 27, 2026 (the page’s own published-time tag) | August 27, 2026 | Both |
| Base model | MiniMax H3-Base, open weights | “the open-weights MiniMax H3 model”, per fal | Both |
| Headline ratio | “Up to 14x speedup on a single Nvidia Blackwell GPU” | “roughly 35x the throughput of the official MiniMax H3 endpoint” | Both |
| Precise ratio | 14.38× | ~35× (no decimal published) | Side A |
| The denominator | Base H3 with dense FA4 attention, run by FastVideo on matched hardware and matched clip duration | fal’s own hosted official MiniMax H3 endpoint — fal controls both sides | Both (named) |
| Hardware the ratio was taken on | 1× NVIDIA B200 (4× and 8× columns also published) | Not stated for the 35×; H3 Max was trained and served on GB200 NVL72 | Side A |
| Clip length the ratio applies to | The 15 s shape — 345 frames at 24 FPS, so 14.375 s of footage | Not stated | Side A |
| Warmup | “the median of three timed requests after one full warmup” | Not disclosed | Side A |
| Excluded from the timing | “Model loading and compilation are excluded.” | Not disclosed | Side A |
| Included in the timing | “encoding, denoising, decoding, audio, muxing, and file output” | Not disclosed | Side A |
| Statistic and sample count | Median, n = 3 | Not disclosed | Side A |
| How the ratio itself was computed | Stated: the unrounded timings, for the same duration and GPU count | Not disclosed | Side A |
| Comparison set | One matched baseline, run by the same team | 12 models, 6 of them named | Both |
| Weights released | Yes — checkpoint and pre-extracted LoRA, under the MiniMax H3 Community License | No — API and playground only | Both |
| Independent wall-clock re-run | None located in this review | None located in this review | Neither |
The asymmetry in the final column is the load-bearing observation, and it is a disclosure difference rather than a suggestion that either number is wrong. Side A publishes its warmup, its exclusions, its statistic and its sample count. Side B publishes none of those four, and publishes no absolute seconds for either half of its ratio. fal also published two further multiples in the same announcement — roughly fifteen times faster than anything with comparable quality, where comparability is set by an internal study that is not published, and a figure of more than fifty times that the announcement attributes to Design Arena. Three multiples, three different denominators, one press release.
Now the part that makes this section worth reading even if you never touch a video model. FastVideo published twelve speedup cells, not one, and the twelve cells disagree with each other in a way that demolishes the headline’s implied generality — honestly, legibly, and in the same table.
| GPUs | Clip shape | Base H3 dense FA4 (s) | FastH3 VSA (s) | Published × | Recomputed × |
|---|---|---|---|---|---|
| 1× B200 | 5 s | 132.5 | 16.2 | 8.16× | 8.179× — matches |
| 1× B200 | 10 s | 377.4 | 31.1 | 12.13× | 12.135× — matches |
| 1× B200 | 15 s — the headline cell | 678.7 | 47.2 | 14.38× | 14.379× — matches |
| 4× B200 | 5 s | 40.6 | 6.1 | 6.65× | 6.656× — matches |
| 4× B200 | 10 s | 108.7 | 12.0 | 9.03× | 9.058× — within rounding |
| 4× B200 | 15 s | 193.1 | 15.5 | 12.48× | 12.458× — within rounding |
Read across rather than down. The headline number, 14.38×, is the largest of the six cells. Change nothing but the clip length and it becomes 8.16× — a 1.76× swing from clip length alone. Across all six VSA cells the smallest is 6.65×, so the headline is 2.16 times the floor of its own table, and every one of those six numbers is equally true. Change nothing but the GPU count, from one B200 to four at the same fifteen-second shape, and it falls to 12.48×, because parallelism helps the baseline too and the ratio shrinks as you scale. Same weights, same kernel, same team, same measurement protocol, same day.
Nobody misled anyone. The table is right there, under the headline, and it is more complete than most vendors publish. That is precisely why it is the strongest exhibit available: it demonstrates that the compression from a six-dimensional result to a single multiple destroys comparability even when the publisher is scrupulous. A speedup claim does not need a bad actor to become unusable. It only needs a headline.
FastVideo also publishes the ablation that separates the two mechanisms behind the speed, and it is the honest analytical read of the length sensitivity above. The dense variant of the same four-step student sits at 7.24×, 7.52× and 7.43× at the five-, ten- and fifteen-second shapes on one B200 — essentially flat across clip length. The sparse variant climbs from 8.16× to 14.38× across the same three shapes. So step distillation, which cuts the transformer calls from 49 to 4, contributes a roughly constant factor, while the video sparse attention at 90% sparsity contributes a factor that grows with sequence length. At the fifteen-second shape, sparse attention is worth 1.94 times what dense is (14.38 ÷ 7.43). That is a mechanism, not a measurement trick — and it means the headline number is length-sensitive by physics, which is the opposite of an accident.
One last derived observation from the same primary, because it shows how far a label can drift from what it names. FastVideo states that its five-, ten- and fifteen-second shapes contain 124, 243 and 345 frames, and that the pipeline runs at 24 FPS. Divide and the shapes are 5.167 s, 10.125 s and 14.375 s of actual footage — so the shape called “15 s” is fourteen and three-eighths seconds long. Put those footage durations next to the published eight-GPU end-to-end times and the sub-realtime claim resolves cleanly: it is true, and it is true for exactly one of the three shapes.
| Shape label | Frames | Footage at 24 FPS (s) | 8× B200 end-to-end (s) | Generation ÷ playback |
|---|---|---|---|---|
| “5 s” | 124 | 5.167 | 6.84 | 1.32× — slower than playback |
| “10 s” | 243 | 10.125 | 11.66 | 1.15× — slower than playback |
| “15 s” | 345 | 14.375 | 12.88 | 0.90× — 1.12× faster than playback |
04 — The Non-Video MirrorOne compiler, one codebase, five published multiples.
If the video pair feels like a special case, here is the same structure in a domain almost every reader has on their laptop. TypeScript’s native compiler rewrite has been benchmarked repeatedly by its own team on the same public codebase — the vscode repository — across the March 2025 native-port announcement and the TypeScript 7.0 GA post, and has been published at five different multiples, every one of them defensible and every one of them measuring something slightly different.
The same compiler on the same repository, five ways
Microsoft TypeScript team, native-port (2025) and TypeScript 7.0 GA (2026) posts · bars scaled to the largest multipleMicrosoft’s headline across both announcements is “10x” — a number that appears nowhere in the five bars above, sitting instead between the third and second of them. Its own GA body text is more careful than the headline, describing optimisations that “typically yield speedups between 8x and 12x on full builds”, and the full GA table ranges from 7.7× to 11.9× across five codebases at the default checker count, then from 10.6× to 16.7× across the same five at --checkers 8. The vendor is not overstating; the headline is a rounding of a range, and the range is published two screens down.
Two honest notes on that chart. The 7.2× bar is a figure we found reported as the VS Code team’s own --noEmit measurement by Visual Studio Magazine; we could not locate a first-party page for it, so treat it as secondhand reporting rather than a vendor statement. And the machine matters here more than anywhere else in this post: the difference between 11.9× and 16.7× is purely the number of checker processes, and the GA post says only that both tables were run on the same machine — it never says what machine, which core count, or which operating system. To Microsoft’s credit the post publishes the caveat itself, noting that these codebases “get a better speedup from dedicating more cores, but results will differ across projects and underlying machines”. It publishes the lever and the warning and omits the hardware. That combination is the industry norm at its better end.
If you are actually deciding whether to move a codebase, our TypeScript 7 migration playbook is the post for that. Here the compiler is only a specimen, and the specimen’s value is arithmetic: this is what it looks like when one product, one repository and one honest team produce five different true multiples.
| Codebase | Current (s) | Native (s) | Published × | Recomputed × |
|---|---|---|---|---|
| VS Code | 77.8 | 7.5 | 10.4× | 10.373× — matches |
| Playwright | 11.1 | 1.1 | 10.1× | 10.091× — matches |
| TypeORM | 17.5 | 1.3 | 13.5× | 13.462× — matches |
| rxjs | 1.1 | 0.1 | 11.0× | 11.0× — matches |
| tRPC | 5.5 | 0.6 | 9.1× | 9.167× — rounds to 9.2×, precision artefact |
| date-fns | 6.5 | 0.7 | 9.5× | 9.286× — not reproducible at this precision |
That last row is teaching material, not an error, and it is worth being explicit about why. Divide the published times and date-fns comes out at 9.29×, not the published 9.5×. The explanation is almost certainly benign: 6.5 seconds and 0.7 seconds are one-decimal roundings of the real timings, and an unrounded 0.68 seconds would produce 9.5× exactly. The tRPC row shows the same artefact in the opposite direction — 9.167× recomputes to 9.2× and is published as 9.1×. This is not a fabricated number; it is what happens when a table publishes one-decimal seconds alongside two-significant-figure ratios.
The general rule that falls out of it is the useful part. When a table publishes rounded inputs and rounded ratios, the ratio column cannot be recomputed from the table. Every other ratio we checked does recompute: all twelve FastH3 cells and all ten TypeScript 7 GA cells land within ±0.03 of their published values. And FastVideo solved the whole problem with one sentence stating that its speedups use the unrounded timings. That sentence costs nothing and it is the single cheapest disclosure improvement available to anyone publishing a benchmark table.
Three more figures from the TypeScript GA post, because they show the same author being rigorous and loose in the same document. Slack’s CI type-check going from about 7.5 minutes to 1.25 minutes is a fully specified 6.0×. Canva’s first-error-in-editor time going from roughly 58 seconds to about 4.8 seconds recomputes to 12.08×. But the same post also reports that the new language server reduced failing language-server commands by over 80% and server crashes by over 60% compared with the previous version — two rates with no sample size for either, no time window, and no statement that they come from the same telemetry population. That is not dishonesty. It is the ordinary asymmetry: the numbers a team benchmarks get conditions, and the numbers a team telemeters get percentages.
05 — The MethodFour questions for the next speedup claim you see.
The practical payoff of everything above is short, and it is deliberately four questions rather than a rubric, because a rubric nobody runs is worth less than four questions somebody does. Each one is answerable from a primary source in about a minute, and the evidence in this post says the answers are usually there — below the fold, in the table, in the caption, or in a linked benchmark document.
What is the denominator?
Name the thing on the bottom of the ratio, and then name who ran it. An unoptimised version of the same product, run by the claimant, is a type-1 claim about engineering. A competitor is a type-4 claim about a market. A previous software version on the same hardware is a type-3 claim that reads like a hardware claim and is not. If the primary does not name a baseline at all — as with one of the three August claims here — stop. There is nothing to compare.
Under exactly what conditions?
The headline is one cell of a table. Find the table and find which cell the headline came from. Ask for the sequence or clip length, the device count, the concurrency, the cache state, the precision, the parallelism flags, and whether model loading and compilation were counted. Section 05 says how much each of those moves a number in practice, and the answer is frequently more than the difference between two competing products.
Which statistic, over how many runs?
A median of three is a different object from a 99th percentile over ten minutes of Poisson-distributed traffic, and the gap between them is where user-visible latency lives. Check whether the number is aggregate or per-user: the TensorRT-LLM table shows aggregate throughput up 19.0× while per-user throughput fell by a factor of 3.4 on identical hardware. Aggregate throughput is a claim about a fleet. Nobody experiences a fleet.
Does the ratio recompute?
Divide the published numbers yourself. It takes ten seconds and it catches two distinct things: an arithmetic slip, and a table whose rounding makes its own ratio column unreproducible — which is what the date-fns 9.5× turns out to be, benignly. If a publisher states that its ratios use unrounded timings, as FastVideo does, you have learned something about the publisher as well as about the number.
Two notes on using this in anger. First, the discipline is the same one that applies to any performance number you report internally, which is why we treat it as measurement work rather than as press-release literacy — it is the same habit we bring to analytics and measurement engagements, where a percentage lift with an unstated baseline period causes precisely this failure and costs considerably more than a misread benchmark. Second, if a claim survives all four questions, believe it — for the conditions it names, and not one condition further.
We reviewed the primary pages behind the three August claims, the FastVideo timing table, Microsoft's TypeScript 7 tables, and the relevant MLCommons rules. Ratios were recomputed from published figures where the source supplied both measurements. This is a worked method, not a frequency study of vendor claims.
06 — ConclusionThe denominator is the claim.
A speedup should be named after the thing it was divided by.
Three speed claims arrived inside a fortnight in August 2026. Two of them shared a headline number and measured unrelated things. Two of them concerned derivatives of one open-weights base model, arrived on the same day, and were not comparable to each other. And the best-documented of the three refuted its own headline’s implied generality in the table directly beneath it, by publishing 8.16× at a five-second clip and 14.38× at fifteen — the same checkpoint, the same GPU, the same protocol.
That is the finding, and it is a more useful one than a disclosure scandal would have been. Nobody in this post behaved badly. The information is there in the primaries; it dies in the compression to a headline. Which means the fix is available to both sides of the transaction: a publisher can add one sentence naming the baseline and the conditions in the same breath as the multiple, and a reader can spend four minutes on four questions — what was it divided by, under what conditions, with which statistic, and does the division check out.
If you want a standard to point at, MLPerf has already written one: thirty-six system-description fields, twenty-seven of them mandatory, a percentile fixed by the scenario rather than chosen by the submitter, adversarial review by other submitters, and a rule that the name of a result is part of its disclosure. No blog post will ever carry all of that, and none needs to. But the naming rule transposes for free, and it is the whole post in nine words: a speedup should be named after its denominator.