Suite · Back-office automation
Exception Economics
What automation costs when it is wrong, in reviewer minutes and money.
Charter 1.0.0, section 2
Unit scored: one work item. Headline metric: net cost per item at a stated confidence threshold.
Built from datasets/exception-economics/.
On this page: leaderboard, metric definitions, difficulty tiers, reproduce, versions and hashes.
Status
Leaderboard
The tier here describes the exercise, not an item in a dataset: this suite has no labelled set and no language dimension. Selected slice: every tier. No count is published for this slice. The filter changes the denominator a row would be measured over. Every measured cell reads not run in every slice.
No sort applied. Rows are in source order.
| GPT (latest) | not run | not run | not run | not run | 0 | not run — no model interface key and no reachable model interface in the build environment |
|---|---|---|---|---|---|---|
| Claude (latest) | not run | not run | not run | not run | 0 | not run — no model interface key and no reachable model interface in the build environment |
| Gemini (latest) | not run | not run | not run | not run | 0 | not run — no model interface key and no reachable model interface in the build environment |
| Mistral (latest) | not run | not run | not run | not run | 0 | not run — no model interface key and no reachable model interface in the build environment |
| Open model A | not run | not run | not run | not run | 0 | not run — no model interface key and no reachable model interface in the build environment |
| Open model B | not run | not run | not run | not run | 0 | not run — no model interface key and no reachable model interface in the build environment |
| Entailment Labs pipeline | not run | not run | not run | not run | 0 | not run — no model interface key and no reachable model interface in the build environment |
Column headings sort the table. A figure that was never produced sorts last in both directions; it is not a low score. Charter 5.6: our own system is scored by the same harness, from the same commit, with the same prompts, and sits wherever the sort puts it, with no highlight and no position advantage.
Nothing to plot. No system has been run, so there is no net cost per item at a stated confidence threshold to draw. Charter 10.4: a chart with no run renders an empty state, not example bars. The reason for each row is in the status column above.
Metric definitions
Read from charter 1.0.0 section 3 when this page was built. Each metric carries its formula, its numerator, its denominator and its exclusions, and each has clause language in charter/contract-clauses.md so it can be written into a statement of work. The charter's arithmetic examples are left out here: they are invented numbers that demonstrate a formula, and beside a table of "not run" they would read as results.
Automation rate
Charter 1.0.0 section 3.14 · read it in the charter · SOW clause 13
- Definition
The share of back-office work items the system carried to a final state with no human action. The unit is a work item, which may span several documents, several lookups and several writes, which is why it is defined separately from the document-level rate in section 3.4.
- Formula
Automation rate = items completed with no human action ÷ items admitted to processing.
- Numerator
Items that reached a final state, posted, matched, closed or rejected by rule, with no human action of any kind, including no reviewer opening the item. The sampled-audit convention in section 3.4.2 applies unchanged: an audited item counts as automated if the audit changed nothing.
- Denominator
Items admitted to processing in the window: items received, less items rejected before processing by a rule named in the statement of work. Both counts are reported.
- Excluded
Items still open at the end of the window, counted and reported as in flight. Items abandoned because an upstream system was unavailable, counted and reported as upstream failures. Items the partner withdrew.
- The rule that governs this metric
Automation rate is never published, quoted or contracted alone. It is meaningless without section 3.15 beside it, because a system that automates every item and is wrong on a tenth of them can cost more than one that automates half and is right. Every table carrying an automation rate carries the wrong-automation figure on the same row.
- Contract sentence
"Automation rate means the number of Items carried to a final state with no human action, divided by the number of Items admitted to processing in the measurement window. The Provider shall achieve an automation rate of not less than placeholder percent while keeping wrong-automation rework at or below the ceiling in clause [X]. The automation rate is not satisfied in any month in which the rework ceiling is exceeded."
Wrong-automation cost in rework minutes
Charter 1.0.0 section 3.15 · read it in the charter · SOW clause 14
- Definition
The labour needed to find and put right the items the system completed automatically and wrongly, expressed in minutes per one thousand automated items, and then in money at the reviewer cost the labour model states.
- Formula
Rework minutes per 1,000 automated items = (the sum over wrong automated items of detection minutes plus correction minutes plus downstream correction minutes) ÷ automated items × 1,000.
- Numerator
Minutes taken from the labour model published with the dataset, which states, per error class, the expected minutes to detect the error by its detection route and the minutes to correct it and to correct anything downstream that consumed it. In production the minutes come from the time recorded against the correction. Where an error class has no detection route inside the window, it is not given a minute figure. It is reported as an open exposure count, with the class named. Inventing a detection time for an error nobody would find is exactly the kind of number this charter exists to prevent.
- Denominator
Items completed automatically, that is the numerator of section 3.14, scaled to one thousand.
- Excluded
Rework on items that went to human review in the first place, which is reviewer time under section 3.16. Rework caused by a source-data error the system could not have detected, reported separately as an input-quality class. Rework caused by a partner-side change outside the statement of work.
- Money
Cost = rework minutes × the fully loaded reviewer cost per minute stated in the dataset labour model, reported in INR and USD, both marked placeholder until a partner supplies the rate. Net cost per item, which is the figure a chief financial officer reads, combines this with section 3.16 and with cost per document from section 3.7, and is reported at three confidence thresholds so that the trade between automation and rework is visible rather than argued.
- Contract sentence
"Wrong-automation rework means the sum of detection, correction and downstream correction minutes attributable to Items completed automatically and incorrectly, divided by the number of Items completed automatically and expressed per one thousand such Items, using the labour model in Schedule [X]. The Provider shall keep wrong-automation rework at or below placeholder minutes per one thousand automated Items. Error classes with no detection route within the measurement window shall be reported as an open exposure count with the class named, and shall not be assigned an estimated minute figure."
Reviewer minutes per exception
Charter 1.0.0 section 3.16 · read it in the charter · SOW clause 15
- Definition
The reviewer time taken to bring an exception to a final state. Reported as the mean and the median, and always broken down by queue entry code, because a low-confidence check and a processing failure are different pieces of work.
- Formula
Mean reviewer minutes per exception = total reviewer minutes recorded on exceptions closed in the window ÷ exceptions closed in the window.
- Numerator
Reviewer active time from the reviewer opening the item to its final state, summed across every reviewer who touched it, including second review and escalation to a senior reviewer. In production, taken from queue timestamps with an idle cut-off, placeholder seconds, after which a reviewer's session is closed and further time is not counted. In the benchmark, taken from the labour model published with the dataset.
- Denominator
Exceptions closed in the window, not exceptions opened. Counting the ones that opened would let the slowest items sit outside the figure forever.
- Excluded
Waiting time in the queue, which is queue age. Training and calibration time. Time on items rejected as out of scope before review, reported separately. Sampled-audit reviews, which are a measurement device and are reported separately as audit minutes.
- Whose number this is
Reviewer time is usually the partner's staff cost, not ours, so this is not a metric a supplier can promise on its own. It is contracted as a reporting duty with a redesign trigger, and Clause 15 in
contract-clauses.mdis written that way.- Contract sentence
"Reviewer minutes per exception means the total reviewer active time recorded against exceptions closed in the measurement window, divided by the number of exceptions closed, reported as a mean and a median and broken down by queue entry code. The Provider shall report this measure in every Monthly Report. Where the mean exceeds placeholder minutes for two consecutive months, the Provider shall, at no charge, analyse the causes and propose a change to thresholds, validation rules, queue design or the model, and shall implement the agreed change within placeholder Business Days."
90-day drift
Charter 1.0.0 section 3.17 · read it in the charter · SOW clause 16
- Definition
The change in a tracked metric between the acceptance baseline and the same metric measured ninety days later. Two figures are required, and neither is complete without the other.
Figure Measured on What it isolates Frozen-set drift The same frozen labelled set version used at acceptance Change in the system: a provider model version, a prompt, a threshold, a library Live-distribution drift A fresh labelled sample drawn from the last thirty days of live input Change in the input: new formats, new senders, new categories, new languages - Formula
Drift = the metric at day 90 minus the metric at acceptance, in percentage points for rates and in percent of baseline for cost and latency. The report always shows both endpoints. A difference published without its endpoints hides which one moved.
- Numerator and denominator
Those of the underlying metric, unchanged. Drift is a difference between two measurements of the same metric, not a metric of its own, and the underlying sample size is printed at both endpoints.
- In the benchmark
The ninety-day drift is simulated by re-scoring under a shifted input distribution defined in the dataset manifest, with new vendor formats and new item categories at a share the manifest states. It is labelled a simulation everywhere it appears, in the table, the chart and the prose. A simulated drift is evidence about a system's sensitivity, not a record of what happened to a deployment.
- Excluded
Change caused by a scope change agreed under change control, which is reported with its change reference and then excluded from the drift figure. The hypercare period stated in the statement of work, because thresholds are still being set in it; the baseline is the acceptance measurement, not the first day of production.
- Contract sentence
"Ninety-day drift means the change in a tracked measure between the acceptance measurement and the measurement taken ninety days later, reported both on the frozen labelled set version used at acceptance and on a fresh labelled sample drawn from the preceding thirty days of live input, with both endpoints and both sample sizes stated. The Provider shall report both figures in the Monthly Report. A fall of more than placeholder percentage points on the frozen labelled set, or any fall that takes a tracked measure below its floor on the live sample, is a Severity 2 Incident under the Retainer Schedule clause 4."
Difficulty tiers
Charter 1.0.0 section 4.4. A tier is a property of the item, assigned when it is generated, and it never changes because a system found the item hard.
| Tier | Sources to reconcile | Key availability | Match cardinality | Tolerance and rules | Unseen categories | Ground truth |
|---|---|---|---|---|---|---|
| T1 | One | Exact key present on both sides | One to one | Exact match only | None | Single correct answer, no judgement |
| T2 | Two | Exact key present but formatted differently on each side | One to one | One tolerance rule, such as rounding or a date window | None | Single correct answer |
| T3 | Two or three | No exact key; matching on a combination of name, amount and date window | One to many | Two or more tolerance rules that can interact | None | Single correct answer, reached by a judgement rule written in the labelling guide |
| T4 | Three or more | No exact key | Many to many, including split and partial settlements | As T3, plus a rule whose outcome depends on the order of application | A share stated in the manifest, drawn from categories absent from the baseline distribution | Single correct answer; the cost of the wrong answer exceeds the cost of review, stated in the labour model |
| T5 | Three or more | No exact key | Many to many | As T4 | As T4, plus at least one category with no labelled example anywhere in the set | Two sources disagree and the rule deciding which prevails sits in a policy document, not in the data; the distribution shift of section 3.17.4 is applied part-way through the window |
4.4.1 The labour model states reviewer minutes and rework minutes per error class per tier, so that the cost of being wrong rises with the tier in the way it does in a real back office.
4.4.2 T5 is the only tier in which the ninety-day drift simulation is applied. It is labelled a simulation in every table it appears in.
The data
This suite has no dataset. It is a rubric and a set of scripted exercises run against a live deployment, ours or anyone's, sixty days after go-live.
The public sample
Nothing to download. The rubric, the scripted incidents and the self-assessment are documents, listed under the reproduce section below.
Reproduce
These commands are read at build time from results/exception-economics-v1.0/reproduce.md. They are the commands that produced the files behind this page, and the commands that replace the "not run" rows with a measurement.
Prerequisites
pip install pyyaml --break-system-packagesThe four commands
python3 generate.py --seed 20260902
python3 validate.py
python3 score.py --out ../../results/exception-economics-v1.0/scores-baseline.json
python3 drift.py --out ../../results/exception-economics-v1.0/drift.json
python3 report.py --results ../../results/exception-economics-v1.0Scoring a real system
{"item_id": "EE-0001", "proposed_outcome": "recon:matched:PO-446576", "confidence": 0.91}Scoring a real system
python3 score.py --predictions runs/vendor-a.jsonl --label "Vendor A" \
--out ../../results/exception-economics-v1.0/scores-vendor-a.jsonReproducing a single figure
python3 -c "import json;d=json.load(open('scores-baseline.json'));print(d['thresholds'][0]['rates']['automation_rate'])"The documents this suite is run from
- Exception Economics dataset v1.0.0 —
datasets/exception-economics/README.md· It runs with no model API, and here is why that is honest · Regenerate everything from a seed · What is in the set · The labour model · The ninety-day drift simulation · Files - Datasheet — Exception Economics dataset v1.0.0 —
datasets/exception-economics/datasheet.md· Motivation · Composition · Collection and generation process · Difficulty parameters · Uses · Distribution and licence
Charter 5.5
Every result is reproducible from a commit. A figure that cannot be reproduced from the dataset version and hash, the harness version and commit, the prompt set hash, the model version string, the run date, the price list date and the exact command line is withdrawn, not defended.
Versions and hashes
| Dataset | datasets/exception-economics/ — not a dataset |
|---|---|
| Dataset seed | not applicable |
| Dataset version in the results | not run |
| Harness version | 1.0.0 |
| Harness commit | not run |
| Scorer version | not run |
| Charter version | 1.0.0 |
| Ground-truth hash | not published for this suite |
| Dataset manifest hash | not published for this suite |
| Results folder | results/exception-economics-v1.0 |
| Run date | not run |
| Price list date | not run — no provider charge has been incurred |
Charter 7.3
No table mixes versions. A row produced under a different dataset, harness or charter version sits in a different table, and a superseded table stays published, marked superseded, with a link to the one that replaced it.