Four benchmark suites
These benchmarks exist to measure the distance between the number in a vendor deck and the number a process produces in production, and to state both in definitions a BPO can put in a contract.
The distance is usually made of four things: a clean test set standing in for messy inputs, a denominator that has quietly dropped the hard items, a success definition that counts an abandoned contact as a win, and a cost figure taken at a token count that real documents do not reach. None of that requires anyone to lie. It only requires the definition to go unstated.
Charter 1.0.0, written 2026-09-02
Every figure on this site is read at build time from the files in this repository. Nothing on any page is typed by hand. On 2026-09-02 the build read 34 source files.
Harness entail-bench 1.0.0. Data CC BY 4.0, code MIT, in nagu-io/benchmarks.
Status
The four suites
| Suite | What it measures | Unit scored | Headline metric | Runs |
|---|---|---|---|---|
| Messy Scan | What degradation does to extraction, and how much of the volume survives without a person | One document | Document-level straight-through-processing rate | 0 |
| Honest Containment | Whether a contact was resolved, or only ended | One contact | Containment under section 3.9 | 0 |
| Exception Economics | What automation costs when it is wrong, in reviewer minutes and money | One work item | Net cost per item at a stated confidence threshold | 0 |
| Day-60 | Whether the system is still trustworthy two months after go-live | One deployment | Day-60 score, 0 to 100 | 0 |
What exists, per suite
| Suite | Dataset | Labelled items | Published sample | Held privately | Results folder | Systems scored |
|---|---|---|---|---|---|---|
| Messy Scan | not a dataset | not applicable | not applicable | not applicable | does not exist | 0 |
| Honest Containment | not a dataset | not applicable | not applicable | not applicable | results/honest-containment-v1.0 | 0 |
| Exception Economics | not a dataset | not applicable | not applicable | not applicable | results/exception-economics-v1.0 | 0 |
| Day-60 | not a dataset | not applicable | not applicable | not applicable | does not exist | 0 |
The rules this runs under
5.1 Prompts are fixed and published in full. Every prompt used by any suite lives in harness/prompts/, is version-controlled, and is printed or hashed into every report that used it. There is no private prompt.
5.2 The same prompt goes to every model. The only permitted differences are the mechanical requirements of an interface: where a system instruction is placed, how an image is encoded, the maximum output tokens, and whether a structured-output mode is used where the interface has one. Every such difference is listed in the report, per model.
5.3 No vendor-specific tuning. No per-model prompt rewriting, no per-model few-shot selection, no per-model temperature search, no retry with a different prompt after a poor answer, and no post-processing that only one system receives. Normalisation is one shared function applied identically to every system's output, and it is published with the harness.
5.4 Three runs per model per suite at identical settings. Reported as the mean with the standard deviation and the minimum and maximum. A single run is never published as a figure. Where three runs cannot be completed, the row says how many ran and why.
5.5 Every result is reproducible from a commit. Every report records the dataset version and content hash, the harness version and commit hash, the prompt set hash, the model version string as the provider reported it, the run date and time, the price list date, and the exact command line. A figure that cannot be reproduced from those is withdrawn, not defended.
5.6 Our own system is scored by the same harness, from the same commit, with the same prompts, on the same data, and appears in the same table. It gets no highlight, no bold, no annotation, no footnote and no position advantage: tables sort on the metric, not on the vendor. Where our number is worse, it stays where the sort puts it.
5.7 No sponsorship, no paid placement, no pre-agreed outcome. If a vendor ever funds the interface cost of running its own system, the report header says so, and it changes nothing else: the vendor sees its rows only through the pre-publication notice in section 8.3, at the same time as every other vendor.
5.8 A system that cannot complete a suite is reported as "not run" or "incomplete" with the reason, whether the reason is a missing key, a rate limit, a refusal, a context limit or a crash. A partial figure is never promoted into a headline table, and an incomplete row is never compared with a complete one.
Charter 1.0.0, section 5
The neutrality rules in full, and what these benchmarks cannot tell a buyer.
Anyone may dispute any figure. The process is public and its outcomes are published whichever way they go.
The rest of this site
Methodology
The charter in full: every metric definition, the tier schemes, the neutrality rules, the data ethics, the versioning, the limitations and the current status.
Run it yourself
Install the harness, point it at your own folder of documents, and score any system through the same scorer.
Changelog
What has changed, at what version, and what is still at version one because nothing has been published.
Disputes
How to dispute a figure, what happens next, and the log of every dispute raised.