Every vendor publishes an extraction accuracy figure; none publishes whether the value can be defended — that the citation points right, the system declined when it should, and an auditor could reconstruct the trail. So we wrote that measure and scored ourselves first.
Parsing benchmarks measure whether the characters came off the page correctly — the binding problem in 2023, not now.
Page references are rarely measured on whether they are right; a correct value with a wrong citation passes every extraction benchmark and fails the one test an examiner cares about.
Averaged accuracy hides the number a risk owner needs: of the values passed through without review, how many were wrong — the only population that reaches a client or regulator unchecked.
A confident error costs more than an abstention, yet headline accuracy punishes the system that abstains correctly — so vendors are rewarded for guessing.
A measure only means anything if the vendor who wrote it can lose on it. The methodology is open so that claim can be checked rather than believed — if somebody beats us, that is a working benchmark, not a failed launch.
Each is computed per field (or per sampled decision for case files) and aggregated per family, with deliberately no headline number — a composite would let a weak measure hide behind a strong one.
| Measure | What is computed | The question it answers |
|---|---|---|
| Citation correctness | Share of asserted values whose cited page and region actually contain the value, verified against the labelled source | “Show me where this number came from.” |
| Gate precision | Error rate within the population that cleared the confidence gate and went through without human review | “How many wrong values reached a client unchecked?” |
| Abstention correctness | Both directions: correct declines on genuinely ambiguous fields, and wrong declines on fields that were legible | “Does it know what it does not know?” |
| Reviewer agreement | Agreement between two independent qualified reviewers and the produced record, on a sampled subset | “Would an expert have recorded the same thing?” |
| Evidence completeness | Share of sampled fields for which a full trail — source, version, transform, gate decision, approver — can be reconstructed with no human involvement | “Reconstruct this record for the examiner.” |
| Case-file completeness | Share of sampled decisions whose case file has every required element — decision, source, how the output was used, human review, policy, a source page, and a verified signature | “Show me how this output shaped this decision.” |
A private corpus makes a benchmark a marketing asset; these are the parts we release, and the two limits we state rather than hide.
Synthetic and consented documents across the two shipping packs, built to reproduce real failure modes: superseded versions, handwritten amendments, scanned inserts, and values that appear twice with different meanings.
Field-level ground truth with the adjudication rules that produced it, so a disagreement about a score can be settled by reading the guide rather than by arguing about intent.
The scripts that compute all six measures from a vendor-neutral output format, plus the adapter we used for our own runs. If your system emits fields and citations, it can be scored.
A vendor-authored benchmark deserves scepticism — two families, one author, first publication. It becomes credible when others run it and argue with it.
No public corpus predicts your numbers — this benchmark makes vendors comparable, while the evaluation on your own documents tells you what you would actually get. Use both, in that order.
Each version is frozen on publication and every result carries its version; corpus changes ship as a new version, and scores are never restated retroactively.
Self-run, on the benchmark corpus, with the harness above — gothink measuring gothink, and the correct amount of trust is “enough to check”.
| Measure | Wealth onboarding | Insurance submissions | Where this is weakest |
|---|---|---|---|
| Citation correctness | — | — | Multi-page tables where a value is restated in a summary |
| Gate precision | — | — | Handwritten amendments that clear the gate on a clean scan |
| Abstention correctness | — | — | Over-abstains on legible but unusual clause layouts |
| Reviewer agreement | — | — | Fields where two reviewers disagree with each other first |
| Evidence completeness | — | — | Records that crossed a schema version boundary mid-run |
| Case-file completeness | — | — | New measure; decisions recorded from other vendors carry only what their logs hold |
The figures are withheld until the first run is complete and independently witnessed — publishing numbers nobody outside gothink has seen would repeat what this page exists to replace. The measures, weaknesses and method are published now because those are what a buyer can hold us to. The scores are published the day the witnessed run finishes, with raw outputs attached.
Ask every shortlisted vendor for their six numbers against v0.1 and their raw outputs — a fair test none can tune to a corpus they did not choose. Take the wording straight from this page.
Request the corpus and harnessThe corpus, labels, adjudication guide and harness are yours under a permissive licence, including for competitive marketing against us. The only condition: a published score carries the version and raw outputs.
Tell us you are running itMost teams that built extraction in-house have never measured gate precision or evidence completeness. Run the harness against your own pipeline before you buy — score well and you save a procurement cycle; score poorly and you learn which layers are missing.
The five layers, and what each one addsA public corpus is the right way to shortlist and the wrong way to choose; once it is down to two, run twenty-five of your own documents through each, scored on the measures above.