External-testing docket
The public record of our 2026-10-05 external-testing round: the findings we confirmed, the releases that fixed them, the sealed before/after digests, and the claims we rejected with reasons.
Audit
Record what each named protocol observed.
Diagnose
Trace each production finding to evidence.
Verify
Re-run the protocol required by the remediation.
The round, in one paragraph
On 2026-10-05, 7 external auditors tested this product on Firm-tier beta accounts. All seven refused to subscribe; six of seven said they would pilot. We reconciled every claim against the code with an eleven-seat panel before building anything, fixed what was real, and publish this docket: what we confirmed, what fixed it, what we rejected with reasons, and the receipts — sealed digests you can re-verify yourself.
This page carries no account, customer, or auditor identifiers. Findings are published where they are reproducible from public artifacts; the one site-level case from the round is published as an anonymized aggregate with its identity withheld.
The demonstration: an operator self-audit arc
Operator self-audit — run by Generative Metrics on its own property, quota-exempt, protocol set 2.2.0.
Finding F18 was that no public fix-verified arc had ever been demonstrated. The engine's verification path existed — what was missing was a public receipt. This is one, end to end, on our own property. Both sides sealed under one protocol set, so the comparison is same-set by construction:
Our /.well-known/tdmrep.json carried the bare-object TDM-Rep form, which our own 2.2.0 validator honestly records as invalid. We moved it to the conforming per-location array form ([{"location":"/","tdm-reservation":1}]) and gave /ai.txt the section blocks the validator requires (directives unchanged). Deployed as commit 1352b98; the BEFORE bundle was sealed after the 2.2.0 deploy and before this fix, so both sides ran under one protocol set.
- Scan id
- 1344aaa1-e621-4495-98b4-f67ec6b73277
- Protocol set
- gm-ai-consumer-compat/2.2.0
sha256 8f46c7…d1f5d (8f46c773ed8709fdc6b16bb0248b89c758f7e8ae8a124769cecc89a4647d1f5d)
Records OPT_OUT_TDM_REP_INVALID (WARNING) — our own tdmrep.json, invalid under our own validator.
Verify this digest at /verify- Scan id
- 861e93dc-6578-49c0-9f2f-1a9a70593867
- Protocol set
- gm-ai-consumer-compat/2.2.0
sha256 29a494…3f145 (29a494bd7476dfafd26255b1cea313b4ae55adc086b8c85c902ab4615003f145)
The invalid-document finding is gone (the signal now records as OPT_OUT_POLICY_SIGNALS_OBSERVED, INFO).
Verify this digest at /verifyVerification receipt
Remediation REPAIR_TDM_REP_DOCUMENT · result VERIFIED · verified at 2026-10-06T07:43:47 · protocol set gm-ai-consumer-compat/2.2.0 · receipt b4f96f71-4e3f-4580-acfb-e5d21ff9fd7a
The site fix between the two bundles: commit 1352b98 — the own-site policy files that give the arc its before/after.
Rationale, verbatim from the receipt: “The candidate TDM-Rep reservation document parses as a TDM reservation.”
What the VERIFIED stamp means here: the rescan of the same target under the same protocol set sealed evidence that satisfied the registered criterion for this remediation (REPAIR_TDM_REP_DOCUMENT). It is a statement about the sealed evidence and the criterion — not a claim that any AI system now accesses or cites this site, and not a citation outcome. The engine resolves cross-set comparisons INDETERMINATE by design; this receipt is same-set.
Confirmed findings and their fixes
Every CONFIRMED-FIXED label on this page is computed from the row's evidence — a VERIFIED stamp plus both sealed digests reads “CONFIRMED-FIXED (verified)”; a fix without that triple reads “FIXED”. The labels are derived in the data module that builds this page, never hand-typed.
- F1FIXED
Entitlement split-brain: the dashboard displayed one limit while the API enforced another
Display and enforcement now read ONE resolver (gm_resolve_entitlements); display==enforcement unified. Verified live on a grant account: 625/625 URL audits in the calendar window, monitors limit 25, API-key create HTTP 200.
Fixed in Release A (3bbef28) · commit 3bbef28
- F2FIXED
Sold features answered 404 in production while the pricing page sold them
The bulk site-audit and comparison flags defaulted off in production, so the sold endpoints did not exist there. The flags are on by default and the endpoints answer 401 unauthenticated instead of 404.
Fixed in Release A (3bbef28) · commit 3bbef28
- F3FIXED
The pass certificate contradicted its own counts
The certificate printed a no-warning sentence beside a warning count and counted every execution instead of the required ones. Certificate copy is now derived from the bundle counts (count-derived copy).
Fixed in Release A (3bbef28) · commit 3bbef28
- F4FIXED
The PASS headline overrode its own probe evidence
The "path is open" headline fired while the same scan recorded required-actor probe denials and thin extraction coverage. The headline is now gated on that evidence, and protocol set 2.2.0 deepened the evidence with ACTOR_ACCESS_DIVERGENCE_SUSPECTED — a robots-corroborated differential-UA finding that never blocks the audit.
Fixed in Release A (3bbef28); deepened in Release A2 (0c9937b) · commit 3bbef28
- F6FIXED
Sitemap declarations in robots.txt were discarded
The sitemap protocol ignored the Sitemap: URLs robots.txt already declares, producing a false not-found finding plus remediation telling the site to do what it already does. Protocol set 2.2.0 threads the robots-declared sitemap facts into the finding, with a validate-the-declared-URLs remediation.
Fixed in Release A2 (0c9937b) · commit 0c9937b
- F13CONFIRMED-FIXED (verified)
The TDM-Rep validator rejected the conforming per-location array form
The validator accepted only the bare-object form, producing false findings on documents that use the W3C CG array shape. Fixed in protocol set 2.2.0 — and demonstrated end to end below in the operator self-audit arc, with the sealed before/after bundles and the verification receipt.
Fixed in Release A2 (0c9937b) — protocol set gm-ai-consumer-compat/2.2.0 · commit 0c9937b
before 8f46c7…d1f5d (8f46c773ed8709fdc6b16bb0248b89c758f7e8ae8a124769cecc89a4647d1f5d)
after 29a494…3f145 (29a494bd7476dfafd26255b1cea313b4ae55adc086b8c85c902ab4615003f145)
- F16FIXED-IN-PART
The actor registry published more actors than it measured
FIXED IN PART by measurement: protocol set 2.2.0 adds the tokens finding, which records robots-declared user-agent tokens outside the evaluated actor set as declared, unclassified facts. The one site-level case from the round publishes as an anonymized aggregate only (identity withheld): one audited public site’s robots.txt declared 16 user-agent tokens outside the 2.2.0 evaluated actor set (among them CCBot, cohere-ai, omgili, omgilibot), 3 documented AI-actor tokens disallowed by name, and the scanner itself was denied at the front door (recorded, not judged). Reproducible from vendors’ public robots.txt files. Protocol set 2.3.0 then expanded the evaluated set from vendor documentation: 14 documented actors added (CCBot, omgili and its omgilibot alias among them); 4 of the recorded outside tokens (petalbot, cohere-ai, tavilybot, youbot) remain unclassified because their vendors publish no live crawler documentation (checked 2026-10-06) — the tokens finding keeps recording them as declared.
Fixed in Release A2 (0c9937b) — the 2.2.0 tokens finding; registry wave 1 in 2.3.0 · commit 0c9937b
- F17FIXED
Silent failures: settings panels could hang with no error state
Settings fetches had no timeout, so a failed request rendered a permanent loading state. Fetches now time out at 10 seconds and render a visible error state with retry.
Fixed in Release A (3bbef28) · commit 3bbef28
- F18FIXED
No public fix-verified arc had ever been demonstrated
The verification path existed in the engine; no public receipt did, and the round’s testers could not complete one through the F1 entitlement defect. This arc is the proof — the operator self-audit below: a finding on our own site, the fix, the rescan under the same protocol set, and the VERIFIED receipt.
Fixed in B4 step 1 (1352b98), published on protocol set 2.2.0 (0c9937b) · commit 1352b98
Claims we rejected
The panel reconciled every claim from the round against the code before anything was built. These did not reproduce. We publish them with the one-line reason, quoted and refuted, because a docket that only lists confirmed work is not a record.
- R1REJECTED
Claim: the scanner is missing a byte-order-mark (BOM) strip
Reason: The BOM strip has been in the policy parsers since 2026-08-26; the claim did not reproduce against the code or its history. (Exists since 2026-08-26.)
- R2REJECTED
Claim: a 429 rate-limit response is judged a BLOCKER
Reason: 429 responses abstain from verdicts — they record as an unobserved state, not a finding about the site. That abstention has been live since 2026-09-15. (429 abstains from verdicts since 2026-09-15.)
- R3REJECTED
Claim: a policy file served as HTML is judged INVALID without a media-type check
Reason: Media-type gates on policy-file parsing are live (robots.txt since Sep 30; the ai.txt/TDM-Rep family with the 2.x releases). An HTML body is not parsed as a policy document. (Media-type gates live since September 30.)
- R4REJECTED
Claim: no signature exists on sealed evidence bundles
Reason: The server-side HMAC seal is real: our server re-verifies the seal under its key on every read, and the digest recipe is published at /diagnostics-proof for recomputation. (The server-side HMAC seal is real.)
- R5REJECTED
Claim: the terms of service forbid white-label reports
Reason: Terms of Service §78 explicitly permits white-label PDF branding for paid tiers (reselling the platform itself is a different, prohibited thing). (ToS §78 explicitly permits it.)
- R6REJECTED
Claim: exports exclude the evidence bundles
Reason: Exports include the sealed evidence bundles; the bundle download is part of the export surface. (Bundles are included.)
- R7REJECTED
Claim: there is no statement about machine-learning training on customer data
Reason: The privacy policy states it — line 29 of the published policy text declares that customer content is not used for model training. (Privacy policy line 29 states it.)
- R8REJECTED
Claim: a 429 on the analyze endpoint fails silently
Reason: The analyze endpoint renders an inline error state on a 429; the failure is visible where the request was made. (An inline error renders.)
- R9REJECTED
Claim: the rescan verification loop is impossible to complete
Reason: The loop is functional — the round’s attempts were blocked by the F1 entitlement defect, not by the loop. The operator self-audit arc below is a completed loop, published. (Functional; was quota-blocked.)
Forward-looking statements, dated
The docket subset of the claim-dating pass ships on this page now: every forward-looking statement here carries its date or its release label.
- Registry expansion — the outside-registry tokens recorded by the 2.2.0 tokens finding join the evaluated actor set from vendor documentation (robots-only wave; no new HTTP probes). Due: Protocol set gm-ai-consumer-compat/2.3.0 — the next release.
- The full claim-dating pass — every forward-looking claim on every public surface dated or owner-held — lands as a whole in release B3. This page carries the dated subset now. Due: Release B3.
How to check this page
- Every digest printed here re-verifies through the canonical recipe published at /diagnostics-proof using the in-browser checker at /verify: SHA-256 over the canonical JSON form of the sealed bundle, keys recursively sorted, no whitespace, UTF-8, with the
bundleSha256field removed. - The fixing commits named on this page are this repository's own release commits, recorded in the changelog release by release.
- The verdict labels are derived from the evidence triples in the data module that renders this page — the derivation itself is pinned by a test in our continuous integration.
Related surfaces: the month page shows the monitored-month formats these receipts feed into, and the errata register holds signed, append-only corrections attached to sealed records.