Skip to main content

What we got wrong, and how to check our work

An audit example we published was wrong. Here are the wrong numbers, the corrected sealed scan, and the register where signed corrections live.

Audit

Record what each named protocol observed.

Diagnose

Trace each production finding to evidence.

Verify

Re-run the protocol required by the remediation.

· Blog

On 2026-09-19 we published an audit example about www.usa.gov that was wrong: it claimed the site served almost no main-content text to a raw crawler. The defect was ours. This post puts the wrong numbers on the record one more time, explains what broke, shows the corrected scan, and describes the register where corrections like this one now live. Nothing here is a claim about how any AI system treats usa.gov — it is a record of what our tool measured, what it got wrong, and what it measures now.

The wrong numbers, verbatim

The example we published on 2026-09-19 reported raw 36 vs rendered 412 words for usa.gov's main content and called it a raw-content deficit. That finding was false. Independent re-measurement found 443 raw words in the page's main element — a deficit did not exist. We retracted the figures publicly on the homepage on 2026-09-29, and the numbers stayed wrong in public for ten days before that.

The sealed record of the failed scan is not hidden. The scan of record is 85be448f-8d8a-49fe-b25f-cb8d3030982c, and it now carries a signed, public correction in our errata register: paste that scan id into /errata and the RETRACTED entry appears, bound to the exact sealed bundle it corrects.

What actually broke

Our raw-content extraction arm used a whole-document heuristic to guess where a page's main content is. On usa.gov the heuristic picked the wrong region — effectively the page footer — and counted its handful of link labels as the page's main content. The rendered-browser arm counted the real content. The gap between those two numbers produced a confident, well-formed, and false finding: the audit compared its own bad guess against the rendered page and reported the difference as the site's defect. It was not the site's defect.

The fix is structural, not a patch for usa.gov. As of protocol set gm-ai-consumer-compat/1.8.0, the engine prefers a page's explicit semantic scope — the <main> element, the [role=main] landmark, or the <article> element — instead of a whole-document guess. When no semantic scope exists, the fallback heuristic now excludes landmark regions (navigation, banners, footers, aside content) rather than sweeping the whole document. And every extraction record now seals what it measured: which element was selected, why, and how much of the document it covered — so the next defect of this kind is inspectable in the evidence instead of invisible behind a number. The changelog documents the release.

The corrected scan

We re-audited www.usa.gov under protocol set 1.8.0 and published the real result on the sample-report page: a PASS / COMPLETE scan sealed on 2026-10-01, whose main-content arms measured raw 443 = rendered 443 = no-JavaScript 443 words on the same capture. The corrected numbers are not hand-edited: they are the main-content arm of the sealed bundle published on that page, and the bundle's SHA-256 digest (8f9da12a9fda3f0448b5a996da5e5f4990dccdd9b493fa697517344c44399fb4) can be recomputed by anyone from the bundle itself.

A PASS is an execution state — it means the required protocols completed without an open production blocker. It makes no statement about whether any provider crawls, indexes, ranks, cites, or trains on usa.gov, and nothing in this post changes that boundary.

The errata register

Sealed evidence bundles are never edited — that is the whole point of sealing them. So when one of them is wrong, the correction attaches to it instead: the public errata register holds signed, append-only entries that are looked up by scan id, one scan at a time. Each entry carries its category, its reason, whether it was reported to us by an external evaluator or found by our own pipeline, the date it was published, and the digest of the sealed bundle it corrects. Entries are signed with the same key discipline as the bundles themselves, and the register has no delete path: a correction, once published, cannot be quietly edited or removed.

Our reports and dashboards now point at the register too: a scan with a recorded correction says so on the result page and in the PDF, next to the verdict — the correction ships inside the artifact rather than living in a blog post.

What this does and does not establish

What it establishes is narrow and checkable: the example we published was wrong; the wrong numbers are quoted above, verbatim; the retraction and the corrected scan are public; the fix is a named protocol-set release; and the retracted scan carries a signed public correction you can look up by scan id. Every claim in this paragraph is a link to the record that backs it.

What it does not establish: that our scanner is now free of defects (this is the second scanner defect we have found in public, after the one we documented on September 19, and we treat that as a reason to keep publishing receipts, not as a claim of cleanliness); that the corrected sample generalizes to any other site; or anything about how AI systems treat usa.gov. An audit we later retract as our error entitles the purchaser to a free replacement scan — that remedy is written on the register itself.

Hold us to the same standard you would hold any auditor.

← Back to blog