Entailment Labs benchmarks

Suite · Back-office automation

Exception Economics

What automation costs when it is wrong, in reviewer minutes and money.

Charter 1.0.0, section 2

Unit scored: one work item. Headline metric: net cost per item at a stated confidence threshold.

Built from datasets/exception-economics/.

On this page: leaderboard, metric definitions, difficulty tiers, reproduce, versions and hashes.

Status

Leaderboard

The tier here describes the exercise, not an item in a dataset: this suite has no labelled set and no language dimension. Selected slice: every tier. No count is published for this slice. The filter changes the denominator a row would be measured over. Every measured cell reads not run in every slice.

No sort applied. Rows are in source order.

Exception Economics — every system scored by the same harness, from the same commit, on the same data
GPT (latest)not runnot runnot runnot run0not run — no model interface key and no reachable model interface in the build environment
Claude (latest)not runnot runnot runnot run0not run — no model interface key and no reachable model interface in the build environment
Gemini (latest)not runnot runnot runnot run0not run — no model interface key and no reachable model interface in the build environment
Mistral (latest)not runnot runnot runnot run0not run — no model interface key and no reachable model interface in the build environment
Open model Anot runnot runnot runnot run0not run — no model interface key and no reachable model interface in the build environment
Open model Bnot runnot runnot runnot run0not run — no model interface key and no reachable model interface in the build environment
Entailment Labs pipelinenot runnot runnot runnot run0not run — no model interface key and no reachable model interface in the build environment

Column headings sort the table. A figure that was never produced sorts last in both directions; it is not a low score. Charter 5.6: our own system is scored by the same harness, from the same commit, with the same prompts, and sits wherever the sort puts it, with no highlight and no position advantage.

Net cost per item at a stated confidence threshold, by system

Nothing to plot. No system has been run, so there is no net cost per item at a stated confidence threshold to draw. Charter 10.4: a chart with no run renders an empty state, not example bars. The reason for each row is in the status column above.

Metric definitions

Read from charter 1.0.0 section 3 when this page was built. Each metric carries its formula, its numerator, its denominator and its exclusions, and each has clause language in charter/contract-clauses.md so it can be written into a statement of work. The charter's arithmetic examples are left out here: they are invented numbers that demonstrate a formula, and beside a table of "not run" they would read as results.

Automation rate

Charter 1.0.0 section 3.14 · read it in the charter · SOW clause 13

Definition

The share of back-office work items the system carried to a final state with no human action. The unit is a work item, which may span several documents, several lookups and several writes, which is why it is defined separately from the document-level rate in section 3.4.

Formula

Automation rate = items completed with no human action ÷ items admitted to processing.

Numerator

Items that reached a final state, posted, matched, closed or rejected by rule, with no human action of any kind, including no reviewer opening the item. The sampled-audit convention in section 3.4.2 applies unchanged: an audited item counts as automated if the audit changed nothing.

Denominator

Items admitted to processing in the window: items received, less items rejected before processing by a rule named in the statement of work. Both counts are reported.

Excluded

Items still open at the end of the window, counted and reported as in flight. Items abandoned because an upstream system was unavailable, counted and reported as upstream failures. Items the partner withdrew.

The rule that governs this metric

Automation rate is never published, quoted or contracted alone. It is meaningless without section 3.15 beside it, because a system that automates every item and is wrong on a tenth of them can cost more than one that automates half and is right. Every table carrying an automation rate carries the wrong-automation figure on the same row.

Contract sentence

"Automation rate means the number of Items carried to a final state with no human action, divided by the number of Items admitted to processing in the measurement window. The Provider shall achieve an automation rate of not less than placeholder percent while keeping wrong-automation rework at or below the ceiling in clause [X]. The automation rate is not satisfied in any month in which the rework ceiling is exceeded."

Wrong-automation cost in rework minutes

Charter 1.0.0 section 3.15 · read it in the charter · SOW clause 14

Definition

The labour needed to find and put right the items the system completed automatically and wrongly, expressed in minutes per one thousand automated items, and then in money at the reviewer cost the labour model states.

Formula

Rework minutes per 1,000 automated items = (the sum over wrong automated items of detection minutes plus correction minutes plus downstream correction minutes) ÷ automated items × 1,000.

Numerator

Minutes taken from the labour model published with the dataset, which states, per error class, the expected minutes to detect the error by its detection route and the minutes to correct it and to correct anything downstream that consumed it. In production the minutes come from the time recorded against the correction. Where an error class has no detection route inside the window, it is not given a minute figure. It is reported as an open exposure count, with the class named. Inventing a detection time for an error nobody would find is exactly the kind of number this charter exists to prevent.

Denominator

Items completed automatically, that is the numerator of section 3.14, scaled to one thousand.

Excluded

Rework on items that went to human review in the first place, which is reviewer time under section 3.16. Rework caused by a source-data error the system could not have detected, reported separately as an input-quality class. Rework caused by a partner-side change outside the statement of work.

Money

Cost = rework minutes × the fully loaded reviewer cost per minute stated in the dataset labour model, reported in INR and USD, both marked placeholder until a partner supplies the rate. Net cost per item, which is the figure a chief financial officer reads, combines this with section 3.16 and with cost per document from section 3.7, and is reported at three confidence thresholds so that the trade between automation and rework is visible rather than argued.

Contract sentence

"Wrong-automation rework means the sum of detection, correction and downstream correction minutes attributable to Items completed automatically and incorrectly, divided by the number of Items completed automatically and expressed per one thousand such Items, using the labour model in Schedule [X]. The Provider shall keep wrong-automation rework at or below placeholder minutes per one thousand automated Items. Error classes with no detection route within the measurement window shall be reported as an open exposure count with the class named, and shall not be assigned an estimated minute figure."

Reviewer minutes per exception

Charter 1.0.0 section 3.16 · read it in the charter · SOW clause 15

Definition

The reviewer time taken to bring an exception to a final state. Reported as the mean and the median, and always broken down by queue entry code, because a low-confidence check and a processing failure are different pieces of work.

Formula

Mean reviewer minutes per exception = total reviewer minutes recorded on exceptions closed in the window ÷ exceptions closed in the window.

Numerator

Reviewer active time from the reviewer opening the item to its final state, summed across every reviewer who touched it, including second review and escalation to a senior reviewer. In production, taken from queue timestamps with an idle cut-off, placeholder seconds, after which a reviewer's session is closed and further time is not counted. In the benchmark, taken from the labour model published with the dataset.

Denominator

Exceptions closed in the window, not exceptions opened. Counting the ones that opened would let the slowest items sit outside the figure forever.

Excluded

Waiting time in the queue, which is queue age. Training and calibration time. Time on items rejected as out of scope before review, reported separately. Sampled-audit reviews, which are a measurement device and are reported separately as audit minutes.

Whose number this is

Reviewer time is usually the partner's staff cost, not ours, so this is not a metric a supplier can promise on its own. It is contracted as a reporting duty with a redesign trigger, and Clause 15 in contract-clauses.md is written that way.

Contract sentence

"Reviewer minutes per exception means the total reviewer active time recorded against exceptions closed in the measurement window, divided by the number of exceptions closed, reported as a mean and a median and broken down by queue entry code. The Provider shall report this measure in every Monthly Report. Where the mean exceeds placeholder minutes for two consecutive months, the Provider shall, at no charge, analyse the causes and propose a change to thresholds, validation rules, queue design or the model, and shall implement the agreed change within placeholder Business Days."

90-day drift

Charter 1.0.0 section 3.17 · read it in the charter · SOW clause 16

Definition

The change in a tracked metric between the acceptance baseline and the same metric measured ninety days later. Two figures are required, and neither is complete without the other.

FigureMeasured onWhat it isolates
Frozen-set driftThe same frozen labelled set version used at acceptanceChange in the system: a provider model version, a prompt, a threshold, a library
Live-distribution driftA fresh labelled sample drawn from the last thirty days of live inputChange in the input: new formats, new senders, new categories, new languages
Formula

Drift = the metric at day 90 minus the metric at acceptance, in percentage points for rates and in percent of baseline for cost and latency. The report always shows both endpoints. A difference published without its endpoints hides which one moved.

Numerator and denominator

Those of the underlying metric, unchanged. Drift is a difference between two measurements of the same metric, not a metric of its own, and the underlying sample size is printed at both endpoints.

In the benchmark

The ninety-day drift is simulated by re-scoring under a shifted input distribution defined in the dataset manifest, with new vendor formats and new item categories at a share the manifest states. It is labelled a simulation everywhere it appears, in the table, the chart and the prose. A simulated drift is evidence about a system's sensitivity, not a record of what happened to a deployment.

Excluded

Change caused by a scope change agreed under change control, which is reported with its change reference and then excluded from the drift figure. The hypercare period stated in the statement of work, because thresholds are still being set in it; the baseline is the acceptance measurement, not the first day of production.

Contract sentence

"Ninety-day drift means the change in a tracked measure between the acceptance measurement and the measurement taken ninety days later, reported both on the frozen labelled set version used at acceptance and on a fresh labelled sample drawn from the preceding thirty days of live input, with both endpoints and both sample sizes stated. The Provider shall report both figures in the Monthly Report. A fall of more than placeholder percentage points on the frozen labelled set, or any fall that takes a tracked measure below its floor on the live sample, is a Severity 2 Incident under the Retainer Schedule clause 4."

Difficulty tiers

Charter 1.0.0 section 4.4. A tier is a property of the item, assigned when it is generated, and it never changes because a system found the item hard.

TierSources to reconcileKey availabilityMatch cardinalityTolerance and rulesUnseen categoriesGround truth
T1OneExact key present on both sidesOne to oneExact match onlyNoneSingle correct answer, no judgement
T2TwoExact key present but formatted differently on each sideOne to oneOne tolerance rule, such as rounding or a date windowNoneSingle correct answer
T3Two or threeNo exact key; matching on a combination of name, amount and date windowOne to manyTwo or more tolerance rules that can interactNoneSingle correct answer, reached by a judgement rule written in the labelling guide
T4Three or moreNo exact keyMany to many, including split and partial settlementsAs T3, plus a rule whose outcome depends on the order of applicationA share stated in the manifest, drawn from categories absent from the baseline distributionSingle correct answer; the cost of the wrong answer exceeds the cost of review, stated in the labour model
T5Three or moreNo exact keyMany to manyAs T4As T4, plus at least one category with no labelled example anywhere in the setTwo sources disagree and the rule deciding which prevails sits in a policy document, not in the data; the distribution shift of section 3.17.4 is applied part-way through the window

4.4.1 The labour model states reviewer minutes and rework minutes per error class per tier, so that the cost of being wrong rises with the tier in the way it does in a real back office.

4.4.2 T5 is the only tier in which the ninety-day drift simulation is applied. It is labelled a simulation in every table it appears in.

The data

This suite has no dataset. It is a rubric and a set of scripted exercises run against a live deployment, ours or anyone's, sixty days after go-live.

The public sample

Nothing to download. The rubric, the scripted incidents and the self-assessment are documents, listed under the reproduce section below.

Reproduce

These commands are read at build time from results/exception-economics-v1.0/reproduce.md. They are the commands that produced the files behind this page, and the commands that replace the "not run" rows with a measurement.

Prerequisites

pip install pyyaml --break-system-packages

The four commands

python3 generate.py --seed 20260902
python3 validate.py
python3 score.py --out ../../results/exception-economics-v1.0/scores-baseline.json
python3 drift.py  --out ../../results/exception-economics-v1.0/drift.json
python3 report.py --results ../../results/exception-economics-v1.0

Scoring a real system

{"item_id": "EE-0001", "proposed_outcome": "recon:matched:PO-446576", "confidence": 0.91}

Scoring a real system

python3 score.py --predictions runs/vendor-a.jsonl --label "Vendor A" \
    --out ../../results/exception-economics-v1.0/scores-vendor-a.json

Reproducing a single figure

python3 -c "import json;d=json.load(open('scores-baseline.json'));print(d['thresholds'][0]['rates']['automation_rate'])"

The documents this suite is run from

  • Exception Economics dataset v1.0.0datasets/exception-economics/README.md · It runs with no model API, and here is why that is honest · Regenerate everything from a seed · What is in the set · The labour model · The ninety-day drift simulation · Files
  • Datasheet — Exception Economics dataset v1.0.0datasets/exception-economics/datasheet.md · Motivation · Composition · Collection and generation process · Difficulty parameters · Uses · Distribution and licence

Charter 5.5

Every result is reproducible from a commit. A figure that cannot be reproduced from the dataset version and hash, the harness version and commit, the prompt set hash, the model version string, the run date, the price list date and the exact command line is withdrawn, not defended.

Run it on your own data.

Versions and hashes

Charter 5.5: a figure that cannot be reproduced from the items below is withdrawn rather than defended.
Datasetdatasets/exception-economics/ — not a dataset
Dataset seednot applicable
Dataset version in the resultsnot run
Harness version1.0.0
Harness commitnot run
Scorer versionnot run
Charter version1.0.0
Ground-truth hashnot published for this suite
Dataset manifest hashnot published for this suite
Results folderresults/exception-economics-v1.0
Run datenot run
Price list datenot run — no provider charge has been incurred

Charter 7.3

No table mixes versions. A row produced under a different dataset, harness or charter version sits in a different table, and a superseded table stays published, marked superseded, with a link to the one that replaced it.

Changelog · Dispute a figure on this page