Skip to main content

External-testing docket

The public record of our 2026-10-05 external-testing round: the findings we confirmed, the releases that fixed them, the sealed before/after digests, and the claims we rejected with reasons.

Audit

Record what each named protocol observed.

Diagnose

Trace each production finding to evidence.

Verify

Re-run the protocol required by the remediation.

The round, in one paragraph

On 2026-10-05, 7 external auditors tested this product on Firm-tier beta accounts. All seven refused to subscribe; six of seven said they would pilot. We reconciled every claim against the code with an eleven-seat panel before building anything, fixed what was real, and publish this docket: what we confirmed, what fixed it, what we rejected with reasons, and the receipts — sealed digests you can re-verify yourself.

This page carries no account, customer, or auditor identifiers. Findings are published where they are reproducible from public artifacts; the one site-level case from the round is published as an anonymized aggregate with its identity withheld.

The demonstration: an operator self-audit arc

Operator self-audit — run by Generative Metrics on its own property, quota-exempt, protocol set 2.2.0.

Finding F18 was that no public fix-verified arc had ever been demonstrated. The engine's verification path existed — what was missing was a public receipt. This is one, end to end, on our own property. Both sides sealed under one protocol set, so the comparison is same-set by construction:

Our /.well-known/tdmrep.json carried the bare-object TDM-Rep form, which our own 2.2.0 validator honestly records as invalid. We moved it to the conforming per-location array form ([{"location":"/","tdm-reservation":1}]) and gave /ai.txt the section blocks the validator requires (directives unchanged). Deployed as commit 1352b98; the BEFORE bundle was sealed after the 2.2.0 deploy and before this fix, so both sides ran under one protocol set.

beforesealed 2026-10-06T07:30:54 · PASS / COMPLETE
Scan id
1344aaa1-e621-4495-98b4-f67ec6b73277
Protocol set
gm-ai-consumer-compat/2.2.0

sha256 8f46c7…d1f5d (8f46c773ed8709fdc6b16bb0248b89c758f7e8ae8a124769cecc89a4647d1f5d)

Records OPT_OUT_TDM_REP_INVALID (WARNING) — our own tdmrep.json, invalid under our own validator.

Verify this digest at /verify
aftersealed 2026-10-06T07:43:43 · PASS / COMPLETE
Scan id
861e93dc-6578-49c0-9f2f-1a9a70593867
Protocol set
gm-ai-consumer-compat/2.2.0

sha256 29a494…3f145 (29a494bd7476dfafd26255b1cea313b4ae55adc086b8c85c902ab4615003f145)

The invalid-document finding is gone (the signal now records as OPT_OUT_POLICY_SIGNALS_OBSERVED, INFO).

Verify this digest at /verify

Verification receipt

Remediation REPAIR_TDM_REP_DOCUMENT · result VERIFIED · verified at 2026-10-06T07:43:47 · protocol set gm-ai-consumer-compat/2.2.0 · receipt b4f96f71-4e3f-4580-acfb-e5d21ff9fd7a

The site fix between the two bundles: commit 1352b98 — the own-site policy files that give the arc its before/after.

Rationale, verbatim from the receipt: “The candidate TDM-Rep reservation document parses as a TDM reservation.”

What the VERIFIED stamp means here: the rescan of the same target under the same protocol set sealed evidence that satisfied the registered criterion for this remediation (REPAIR_TDM_REP_DOCUMENT). It is a statement about the sealed evidence and the criterion — not a claim that any AI system now accesses or cites this site, and not a citation outcome. The engine resolves cross-set comparisons INDETERMINATE by design; this receipt is same-set.

Confirmed findings and their fixes

Every CONFIRMED-FIXED label on this page is computed from the row's evidence — a VERIFIED stamp plus both sealed digests reads “CONFIRMED-FIXED (verified)”; a fix without that triple reads “FIXED”. The labels are derived in the data module that builds this page, never hand-typed.

  • F1FIXED

    Entitlement split-brain: the dashboard displayed one limit while the API enforced another

    Display and enforcement now read ONE resolver (gm_resolve_entitlements); display==enforcement unified. Verified live on a grant account: 625/625 URL audits in the calendar window, monitors limit 25, API-key create HTTP 200.

    Fixed in Release A (3bbef28) · commit 3bbef28

  • F2FIXED

    Sold features answered 404 in production while the pricing page sold them

    The bulk site-audit and comparison flags defaulted off in production, so the sold endpoints did not exist there. The flags are on by default and the endpoints answer 401 unauthenticated instead of 404.

    Fixed in Release A (3bbef28) · commit 3bbef28

  • F3FIXED

    The pass certificate contradicted its own counts

    The certificate printed a no-warning sentence beside a warning count and counted every execution instead of the required ones. Certificate copy is now derived from the bundle counts (count-derived copy).

    Fixed in Release A (3bbef28) · commit 3bbef28

  • F4FIXED

    The PASS headline overrode its own probe evidence

    The "path is open" headline fired while the same scan recorded required-actor probe denials and thin extraction coverage. The headline is now gated on that evidence, and protocol set 2.2.0 deepened the evidence with ACTOR_ACCESS_DIVERGENCE_SUSPECTED — a robots-corroborated differential-UA finding that never blocks the audit.

    Fixed in Release A (3bbef28); deepened in Release A2 (0c9937b) · commit 3bbef28

  • F6FIXED

    Sitemap declarations in robots.txt were discarded

    The sitemap protocol ignored the Sitemap: URLs robots.txt already declares, producing a false not-found finding plus remediation telling the site to do what it already does. Protocol set 2.2.0 threads the robots-declared sitemap facts into the finding, with a validate-the-declared-URLs remediation.

    Fixed in Release A2 (0c9937b) · commit 0c9937b

  • F13CONFIRMED-FIXED (verified)

    The TDM-Rep validator rejected the conforming per-location array form

    The validator accepted only the bare-object form, producing false findings on documents that use the W3C CG array shape. Fixed in protocol set 2.2.0 — and demonstrated end to end below in the operator self-audit arc, with the sealed before/after bundles and the verification receipt.

    Fixed in Release A2 (0c9937b) — protocol set gm-ai-consumer-compat/2.2.0 · commit 0c9937b

    before 8f46c7…d1f5d (8f46c773ed8709fdc6b16bb0248b89c758f7e8ae8a124769cecc89a4647d1f5d)

    after 29a494…3f145 (29a494bd7476dfafd26255b1cea313b4ae55adc086b8c85c902ab4615003f145)

  • F16FIXED-IN-PART

    The actor registry published more actors than it measured

    FIXED IN PART by measurement: protocol set 2.2.0 adds the tokens finding, which records robots-declared user-agent tokens outside the evaluated actor set as declared, unclassified facts. The one site-level case from the round publishes as an anonymized aggregate only (identity withheld): one audited public site’s robots.txt declared 16 user-agent tokens outside the 2.2.0 evaluated actor set (among them CCBot, cohere-ai, omgili, omgilibot), 3 documented AI-actor tokens disallowed by name, and the scanner itself was denied at the front door (recorded, not judged). Reproducible from vendors’ public robots.txt files. Protocol set 2.3.0 then expanded the evaluated set from vendor documentation: 14 documented actors added (CCBot, omgili and its omgilibot alias among them); 4 of the recorded outside tokens (petalbot, cohere-ai, tavilybot, youbot) remain unclassified because their vendors publish no live crawler documentation (checked 2026-10-06) — the tokens finding keeps recording them as declared.

    Fixed in Release A2 (0c9937b) — the 2.2.0 tokens finding; registry wave 1 in 2.3.0 · commit 0c9937b

  • F17FIXED

    Silent failures: settings panels could hang with no error state

    Settings fetches had no timeout, so a failed request rendered a permanent loading state. Fetches now time out at 10 seconds and render a visible error state with retry.

    Fixed in Release A (3bbef28) · commit 3bbef28

  • F18FIXED

    No public fix-verified arc had ever been demonstrated

    The verification path existed in the engine; no public receipt did, and the round’s testers could not complete one through the F1 entitlement defect. This arc is the proof — the operator self-audit below: a finding on our own site, the fix, the rescan under the same protocol set, and the VERIFIED receipt.

    Fixed in B4 step 1 (1352b98), published on protocol set 2.2.0 (0c9937b) · commit 1352b98

Claims we rejected

The panel reconciled every claim from the round against the code before anything was built. These did not reproduce. We publish them with the one-line reason, quoted and refuted, because a docket that only lists confirmed work is not a record.

  • R1REJECTED

    Claim: the scanner is missing a byte-order-mark (BOM) strip

    Reason: The BOM strip has been in the policy parsers since 2026-08-26; the claim did not reproduce against the code or its history. (Exists since 2026-08-26.)

  • R2REJECTED

    Claim: a 429 rate-limit response is judged a BLOCKER

    Reason: 429 responses abstain from verdicts — they record as an unobserved state, not a finding about the site. That abstention has been live since 2026-09-15. (429 abstains from verdicts since 2026-09-15.)

  • R3REJECTED

    Claim: a policy file served as HTML is judged INVALID without a media-type check

    Reason: Media-type gates on policy-file parsing are live (robots.txt since Sep 30; the ai.txt/TDM-Rep family with the 2.x releases). An HTML body is not parsed as a policy document. (Media-type gates live since September 30.)

  • R4REJECTED

    Claim: no signature exists on sealed evidence bundles

    Reason: The server-side HMAC seal is real: our server re-verifies the seal under its key on every read, and the digest recipe is published at /diagnostics-proof for recomputation. (The server-side HMAC seal is real.)

  • R5REJECTED

    Claim: the terms of service forbid white-label reports

    Reason: Terms of Service §78 explicitly permits white-label PDF branding for paid tiers (reselling the platform itself is a different, prohibited thing). (ToS §78 explicitly permits it.)

  • R6REJECTED

    Claim: exports exclude the evidence bundles

    Reason: Exports include the sealed evidence bundles; the bundle download is part of the export surface. (Bundles are included.)

  • R7REJECTED

    Claim: there is no statement about machine-learning training on customer data

    Reason: The privacy policy states it — line 29 of the published policy text declares that customer content is not used for model training. (Privacy policy line 29 states it.)

  • R8REJECTED

    Claim: a 429 on the analyze endpoint fails silently

    Reason: The analyze endpoint renders an inline error state on a 429; the failure is visible where the request was made. (An inline error renders.)

  • R9REJECTED

    Claim: the rescan verification loop is impossible to complete

    Reason: The loop is functional — the round’s attempts were blocked by the F1 entitlement defect, not by the loop. The operator self-audit arc below is a completed loop, published. (Functional; was quota-blocked.)

Forward-looking statements, dated

The docket subset of the claim-dating pass ships on this page now: every forward-looking statement here carries its date or its release label.

  • Registry expansion — the outside-registry tokens recorded by the 2.2.0 tokens finding join the evaluated actor set from vendor documentation (robots-only wave; no new HTTP probes). Due: Protocol set gm-ai-consumer-compat/2.3.0 — the next release.
  • The full claim-dating pass — every forward-looking claim on every public surface dated or owner-held — lands as a whole in release B3. This page carries the dated subset now. Due: Release B3.

How to check this page

  1. Every digest printed here re-verifies through the canonical recipe published at /diagnostics-proof using the in-browser checker at /verify: SHA-256 over the canonical JSON form of the sealed bundle, keys recursively sorted, no whitespace, UTF-8, with the bundleSha256 field removed.
  2. The fixing commits named on this page are this repository's own release commits, recorded in the changelog release by release.
  3. The verdict labels are derived from the evidence triples in the data module that renders this page — the derivation itself is pinned by a test in our continuous integration.

Related surfaces: the month page shows the monitored-month formats these receipts feed into, and the errata register holds signed, append-only corrections attached to sealed records.