Entailment Labs benchmarks

Four benchmark suites

These benchmarks exist to measure the distance between the number in a vendor deck and the number a process produces in production, and to state both in definitions a BPO can put in a contract.

The distance is usually made of four things: a clean test set standing in for messy inputs, a denominator that has quietly dropped the hard items, a success definition that counts an abandoned contact as a win, and a cost figure taken at a token count that real documents do not reach. None of that requires anyone to lie. It only requires the definition to go unstated.

Charter 1.0.0, written 2026-09-02

Every figure on this site is read at build time from the files in this repository. Nothing on any page is typed by hand. On 2026-09-02 the build read 34 source files.

Harness entail-bench 1.0.0. Data CC BY 4.0, code MIT, in nagu-io/benchmarks.

Status

The four suites

One suite per line of work. Each scores a different unit, and the unit decides every denominator.
SuiteWhat it measuresUnit scoredHeadline metricRuns
Messy ScanWhat degradation does to extraction, and how much of the volume survives without a personOne documentDocument-level straight-through-processing rate0
Honest ContainmentWhether a contact was resolved, or only endedOne contactContainment under section 3.90
Exception EconomicsWhat automation costs when it is wrong, in reviewer minutes and moneyOne work itemNet cost per item at a stated confidence threshold0
Day-60Whether the system is still trustworthy two months after go-liveOne deploymentDay-60 score, 0 to 1000

What exists, per suite

Read from the dataset manifests, the harness registry and the results folders when this page was built. Honest Containment publishes its whole public split, and holds a separate private set beside it; the other two publish a sample drawn to the same tier and language mix as the whole set.
SuiteDatasetLabelled itemsPublished sampleHeld privatelyResults folderSystems scored
Messy Scannot a datasetnot applicablenot applicablenot applicabledoes not exist0
Honest Containmentnot a datasetnot applicablenot applicablenot applicableresults/honest-containment-v1.00
Exception Economicsnot a datasetnot applicablenot applicablenot applicableresults/exception-economics-v1.00
Day-60not a datasetnot applicablenot applicablenot applicabledoes not exist0

The rules this runs under

5.1 Prompts are fixed and published in full. Every prompt used by any suite lives in harness/prompts/, is version-controlled, and is printed or hashed into every report that used it. There is no private prompt.

5.2 The same prompt goes to every model. The only permitted differences are the mechanical requirements of an interface: where a system instruction is placed, how an image is encoded, the maximum output tokens, and whether a structured-output mode is used where the interface has one. Every such difference is listed in the report, per model.

5.3 No vendor-specific tuning. No per-model prompt rewriting, no per-model few-shot selection, no per-model temperature search, no retry with a different prompt after a poor answer, and no post-processing that only one system receives. Normalisation is one shared function applied identically to every system's output, and it is published with the harness.

5.4 Three runs per model per suite at identical settings. Reported as the mean with the standard deviation and the minimum and maximum. A single run is never published as a figure. Where three runs cannot be completed, the row says how many ran and why.

5.5 Every result is reproducible from a commit. Every report records the dataset version and content hash, the harness version and commit hash, the prompt set hash, the model version string as the provider reported it, the run date and time, the price list date, and the exact command line. A figure that cannot be reproduced from those is withdrawn, not defended.

5.6 Our own system is scored by the same harness, from the same commit, with the same prompts, on the same data, and appears in the same table. It gets no highlight, no bold, no annotation, no footnote and no position advantage: tables sort on the metric, not on the vendor. Where our number is worse, it stays where the sort puts it.

5.7 No sponsorship, no paid placement, no pre-agreed outcome. If a vendor ever funds the interface cost of running its own system, the report header says so, and it changes nothing else: the vendor sees its rows only through the pre-publication notice in section 8.3, at the same time as every other vendor.

5.8 A system that cannot complete a suite is reported as "not run" or "incomplete" with the reason, whether the reason is a missing key, a rate limit, a refusal, a context limit or a crash. A partial figure is never promoted into a headline table, and an incomplete row is never compared with a complete one.

Charter 1.0.0, section 5

The neutrality rules in full, and what these benchmarks cannot tell a buyer.

Anyone may dispute any figure. The process is public and its outcomes are published whichever way they go.