Home/The Attestation Benchmark
Open benchmark / v0.1, published Sep 2026

We published the benchmark. Including where we lose.

Every vendor publishes an extraction accuracy figure; none publishes whether the value can be defended — that the citation points right, the system declined when it should, and an auditor could reconstruct the trail. So we wrote that measure and scored ourselves first.

Open methodology Corpus available on request Any vendor may run it Our weakest measure named
01 / why a new measure

An extraction score tells you the value was read. It does not tell you it can be relied on.

Parsing benchmarks measure whether the characters came off the page correctly — the binding problem in 2023, not now.

The gap

A correct value with a wrong citation

Page references are rarely measured on whether they are right; a correct value with a wrong citation passes every extraction benchmark and fails the one test an examiner cares about.

The gap

Confidence with nothing behind it

Averaged accuracy hides the number a risk owner needs: of the values passed through without review, how many were wrong — the only population that reaches a client or regulator unchecked.

The gap

No credit for saying “I don’t know”

A confident error costs more than an abstention, yet headline accuracy punishes the system that abstains correctly — so vendors are rewarded for guessing.

Why we are publishing something we can be beaten on

A measure only means anything if the vendor who wrote it can lose on it. The methodology is open so that claim can be checked rather than believed — if somebody beats us, that is a working benchmark, not a failed launch.

02 / the six measures

Six measures, each tied to a question somebody will be asked under oath.

Each is computed per field (or per sampled decision for case files) and aggregated per family, with deliberately no headline number — a composite would let a weak measure hide behind a strong one.

MeasureWhat is computedThe question it answers
Citation correctnessShare of asserted values whose cited page and region actually contain the value, verified against the labelled source“Show me where this number came from.”
Gate precisionError rate within the population that cleared the confidence gate and went through without human review“How many wrong values reached a client unchecked?”
Abstention correctnessBoth directions: correct declines on genuinely ambiguous fields, and wrong declines on fields that were legible“Does it know what it does not know?”
Reviewer agreementAgreement between two independent qualified reviewers and the produced record, on a sampled subset“Would an expert have recorded the same thing?”
Evidence completenessShare of sampled fields for which a full trail — source, version, transform, gate decision, approver — can be reconstructed with no human involvement“Reconstruct this record for the examiner.”
Case-file completenessShare of sampled decisions whose case file has every required element — decision, source, how the output was used, human review, policy, a source page, and a verified signature“Show me how this output shaped this decision.”
Deliberately excluded: cost per page and pages per second. Both are real operational numbers, and neither belongs here — publishing them would invite a cheaper, faster run that clears the gate because the gate was lowered. We report throughput and cost to customers in the evaluation report.
03 / the corpus and the method

Published in full, so a result can be reproduced rather than trusted.

A private corpus makes a benchmark a marketing asset; these are the parts we release, and the two limits we state rather than hide.

On request

The document set

Synthetic and consented documents across the two shipping packs, built to reproduce real failure modes: superseded versions, handwritten amendments, scanned inserts, and values that appear twice with different meanings.

On request

The labels and the labelling guide

Field-level ground truth with the adjudication rules that produced it, so a disagreement about a score can be settled by reading the guide rather than by arguing about intent.

In the runtime

The scoring harness

The scripts that compute all six measures from a vendor-neutral output format, plus the adapter we used for our own runs. If your system emits fields and citations, it can be scored.

Stated limit

We wrote it, and it is v0.1

A vendor-authored benchmark deserves scepticism — two families, one author, first publication. It becomes credible when others run it and argue with it.

Stated limit

It is not your estate

No public corpus predicts your numbers — this benchmark makes vendors comparable, while the evaluation on your own documents tells you what you would actually get. Use both, in that order.

Versioning

Frozen and dated

Each version is frozen on publication and every result carries its version; corpus changes ship as a new version, and scores are never restated retroactively.

04 / our scores

Gothink AIR against Attestation Benchmark v0.1.

Self-run, on the benchmark corpus, with the harness above — gothink measuring gothink, and the correct amount of trust is “enough to check”.

MeasureWealth onboardingInsurance submissionsWhere this is weakest
Citation correctnessMulti-page tables where a value is restated in a summary
Gate precisionHandwritten amendments that clear the gate on a clean scan
Abstention correctnessOver-abstains on legible but unusual clause layouts
Reviewer agreementFields where two reviewers disagree with each other first
Evidence completenessRecords that crossed a schema version boundary mid-run
Case-file completenessNew measure; decisions recorded from other vendors carry only what their logs hold

Publish the run before you publish the number

The figures are withheld until the first run is complete and independently witnessed — publishing numbers nobody outside gothink has seen would repeat what this page exists to replace. The measures, weaknesses and method are published now because those are what a buyer can hold us to. The scores are published the day the witnessed run finishes, with raw outputs attached.

No competitor scores appear on this page, and none will be published by us. A vendor scoring its rivals is a marketing exercise, however careful the method; the public harness lets buyers, analysts and vendors produce those comparisons. We will link to any v0.1 run that publishes its outputs, including runs where Gothink AIR comes second.
05 / running it

Three ways to use this, depending on who you are.

If you are buying

Put it in the RFP

Ask every shortlisted vendor for their six numbers against v0.1 and their raw outputs — a fair test none can tune to a corpus they did not choose. Take the wording straight from this page.

Request the corpus and harness
If you are a vendor

Run it and publish

The corpus, labels, adjudication guide and harness are yours under a permissive licence, including for competitive marketing against us. The only condition: a published score carries the version and raw outputs.

Tell us you are running it
If it is your own build

Score what you already have

Most teams that built extraction in-house have never measured gate precision or evidence completeness. Run the harness against your own pipeline before you buy — score well and you save a procurement cycle; score poorly and you learn which layers are missing.

The five layers, and what each one adds
The comparison that actually decides it / your estate

The benchmark makes vendors comparable. Your documents decide it.

A public corpus is the right way to shortlist and the wrong way to choose; once it is down to two, run twenty-five of your own documents through each, scored on the measures above.