Entailment Labs benchmarks

Suite · AI operations

Day-60

Whether the system is still trustworthy two months after go-live.

Charter 1.0.0, section 2

Unit scored: one deployment. Headline metric: day-60 score, 0 to 100.

Built from day-60/ rubric and scripted exercises.

On this page: leaderboard, metric definitions, difficulty tiers, reproduce, versions and hashes.

Status

Leaderboard

Selected slice: every tier, every language. No deployment has been exercised at any tier, so the slice carries no count. A Day-60 score is only comparable with another Day-60 score run at the same tier. The filter changes the denominator a row would be measured over. Every measured cell reads not run in every slice.

No sort applied. Rows are in source order.

Day-60 — deployments exercised against the published rubric, scored 0 to 100
No deployment has been exercised. A Day-60 score needs a live deployment, a partner and an exercise window agreed in writing under charter 4.5.1.

Column headings sort the table. A figure that was never produced sorts last in both directions; it is not a low score. Charter 5.6: our own system is scored by the same harness, from the same commit, with the same prompts, and sits wherever the sort puts it, with no highlight and no position advantage.

Day-60 score, 0 to 100, by deployment

Nothing to plot. No deployment has been exercised, so there is no Day-60 score to draw. Charter 10.4: a chart with no run renders an empty state, not example bars.

The rubric

A Day-60 score is 100 points against the rubric below, run against a live deployment sixty days after go-live. Every observed value is empty because no deployment has been exercised. Charter 4.5.1: an exercise that cannot meet the safety conditions is not run, and the affected rubric lines are scored "not exercised", never assumed.

17 scoring lines shown, 100 points of the 100 the rubric allocates.

No sort applied. Rows are in source order.

Day-60 rubric lines, read from day-60/scoresheet.csv
A1Detection lead timeratio8not exercisednot exercised
A2Notice lead timeratio8not exercisednot exercised
A3Notice contentchecklist5not exercisednot exercised
A4False-alert accountingcomputed4not exercisednot exercised
B1Acknowledgementratio5not exercisednot exercised
B2Time to restoreratio10not exercisednot exercised
B3Communication during the incidentchecklist5not exercisednot exercised
B4Post-incident recordchecklist5not exercisednot exercised
C1Decision timeratio4not exercisednot exercised
C2Rollback timeratio8not exercisednot exercised
C3Re-queue completionratio5not exercisednot exercised
C4Restored configuration provenchecklist3not exercisednot exercised
D1Completenesscomputed12not exercisednot exercised
D2Changes stated with their dateschecklist4not exercisednot exercised
D3Timelinessbanded4not exercisednot exercised
E1Timestamps come from system recordschecklist5not exercisednot exercised
E2A reported figure can be reproducedchecklist5not exercisednot exercised

A ratio line scores full points where the observed value meets the contractual target, half where it is up to twice the target, and zero beyond that or where the contract carries no target. The bands are in day-60/rubric.md section 3.3.

Recorded before any line is scored. A Day-60 score is only comparable with another Day-60 score run at the same tier, so the tier is printed beside the score.
FieldWhat goes in itRecorded
Tier exercisedOne of T1 to T5, from rubric.md section 2. Print it beside the score everywhere the score appears (charter 4.5 and 9.9).not exercised
Exercise windowStart and end dates of the exercise. Fixes the meaning of 'not achieved within the window' on every ratio line.not exercised
AssessorsTwo, named. Each scores every line independently from the same evidence pack. Disagreements resolve under rubric.md 3.7.not exercised
Rules that apply to every row1. A line with no evidence_reference scores 0, whatever the assessors believe happened (rubric.md 3.8). 2. A line that could not be exercised is written 'not exercised' in band, scores no points, and is removed from points_achievable; it is never assumed. 3. A line scored 0 because the contract sets no target is written 'no target in the contract' in note and listed separately in the exercise report. 4. This file is shipped with no scores in it and stays that way; no Day-60 exercise has been run.not exercised
Point allocation, and what has been awarded.
SectionPoints availablePoints awarded
Section A subtotal, drift detection and notice25not exercised
Section B subtotal, incident response and communication25not exercised
Section C subtotal, rollback20not exercised
Section D subtotal, monthly report completeness20not exercised
Section E subtotal, evidence and reproducibility10not exercised
Points awarded and points achievable100not exercised
Day-60 score100not exercised

Metric definitions

Read from charter 1.0.0 section 3 when this page was built. Each metric carries its formula, its numerator, its denominator and its exclusions, and each has clause language in charter/contract-clauses.md so it can be written into a statement of work. The charter's arithmetic examples are left out here: they are invented numbers that demonstrate a formula, and beside a table of "not run" they would read as results.

Drift detection lead time

Charter 1.0.0 section 3.18 · read it in the charter · SOW clause 17

Definition

The elapsed time from the first item affected by a drift to the supplier telling the partner about it in writing. Two clocks are reported, because they answer different questions.

ClockFromToThe question it answers
DetectionFirst affected itemThe internal alert or the ticket that records the driftDid the monitoring see it
NoticeFirst affected itemThe written notice to the partner naming the class and the evidenceDid anyone act on what the monitoring saw
Formula

Lead time = the notice timestamp minus the onset timestamp, reported in hours and in business hours, with the detection clock reported beside it. Numerator and denominator: a single lead time is an elapsed time and has neither. Where a report gives a mean over several drift observations, the numerator is the sum of the lead times and the denominator is the count of drift observations in the window, and both are printed, together with the maximum.

Onset

The timestamp of the first input item belonging to the shifted population, established from the record after the fact. In a Day-60 exercise the injection timestamp is known in advance and held by the person running the exercise, not by the team being measured.

Excluded

Drift the partner announced in advance, such as a scheduled format change, which is a planned change and not a detection. Injections the supplier was told about. Signals raised and withdrawn as false, which are counted separately as a false-alert count, because a monitoring system that alerts constantly also has a lead time near zero and is worth nothing.

Where the drift was not found

If the exercise window closes with no notice, the figure is written "not detected within the window", with the window length stated. It is never extrapolated to a longer window.

Contract sentence

"Drift detection lead time means the elapsed time from the first Item belonging to a shifted input population to the Provider's written notice to the Partner naming the affected class and the supporting evidence. The Provider shall give that notice within placeholder business hours of onset, and shall report both the detection and the notice timestamps for every drift observation in the Monthly Report. A drift not notified within placeholder business hours of onset is a Severity 2 Incident under the Retainer Schedule clause 4."

Incident mean time to restore

Charter 1.0.0 section 3.19 · read it in the charter · SOW clause 18

Definition

The mean elapsed time to restore service, computed separately for each severity in the Retainer Schedule clause 4. Reported with the count, the median and the maximum, because a mean over three incidents conceals the worst of them.

Formula

Mean time to restore for a severity = the sum of restoration durations for incidents of that severity closed in the window ÷ the count of those incidents. The numerator is that sum of durations, measured as 3.19.3 and 3.19.4 define. The denominator is the count of incidents of that severity closed in the window, and it is printed with the figure, because a mean over two incidents is not a rate.

Start of the clock

The earlier of the supplier's own detection timestamp and the partner's notice timestamp. Not the time the ticket was opened, and not the time an engineer picked it up.

End of the clock

The time the covered system returned to the state where work submitted at intake reaches an output or the review queue within the service level, confirmed by a monitoring check. Not the time the ticket was closed, and not the time the root cause was found.

Excluded

Incidents excluded under the Retainer Schedule clause 10. Incidents raised and withdrawn as not reproducible, counted separately. Time inside an agreed change window. Time waiting on a partner dependency, which stops the clock and is reported as a separate "waiting on partner" total, so that neither party can hide inside the other's delay.

Contract sentence

"Mean time to restore means the sum of restoration durations for Incidents of a given Severity closed in the measurement window, divided by the number of such Incidents, where the clock starts at the earlier of the Provider's detection and the Partner's notice and stops when a monitoring check confirms that work submitted at intake reaches an output or the human review queue within the Service Level. The Provider shall meet the resolution targets in the Retainer Schedule clause 3 and shall report, for each Severity, the count, the mean, the median and the maximum, together with any time excluded as waiting on a Partner dependency."

Rollback time

Charter 1.0.0 section 3.20 · read it in the charter · SOW clause 19

Definition

The elapsed time from the decision to roll back to the previous release serving production traffic and passing the smoke set with its model pin, prompt version and threshold set restored. Two further times are reported beside it: decision time, and re-queue completion.

TimeFromTo
Decision timeThe rollback criterion being met, per 06-delivery/build-standards.md section 9The on-call engineer recording the decision
Rollback timeThe decision being recordedThe smoke set passing on the restored release
Re-queue completionThe decision being recordedEvery item processed by the rolled-back release identified and re-queued for review
Formula

Rollback time = the smoke-set pass timestamp minus the decision timestamp, in minutes. Numerator and denominator: a single rollback is an elapsed time and has neither. Where a report covers more than one rollback, the numerator is the sum of the rollback times and the denominator is the count of rollbacks in the period, reported separately for rehearsals and for rollbacks performed on production, with the maximum of each.

Why three times

A rollback that restores the service quickly but leaves the outputs of a bad release sitting in the partner's downstream systems is not finished. A rollback that takes four minutes after a two-hour argument is not fast. Publishing one number and not the other three is the usual way this measure is made to look good.

Excluded and reported separately

A rollback that also requires a data restore reports the restore time separately under 02-security/security-policy-set/07-backup-and-recovery.md. Rehearsals on a non-production copy are labelled as rehearsals and are never mixed in a table with rollbacks performed on production.

Contract sentence

"Rollback time means the elapsed time from the recording of a rollback decision to the previous signed release serving production traffic and passing the smoke set with its model pin, prompt version and threshold set restored. The Provider shall complete a rollback within placeholder minutes, shall rehearse a rollback at least once per quarter and record the rehearsal result in the Monthly Report, and shall report decision time and re-queue completion time alongside rollback time for every rollback performed."

<!-- Benchmark charter, part 4. Indexed in ../methodology.md. Sections 3.21 to 7. -->

Report completeness

Charter 1.0.0 section 3.21 · read it in the charter · SOW clause 20

Definition

The share of required elements of the monthly report that are present, cover the whole period, and carry the basis or the sample size where the element is a figure.

Formula

Report completeness = elements present and passing ÷ elements required.

Numerator

Elements that are present, cover the whole reporting period, and are supported. An element that is present but carries a figure without its sample size or its basis fails, because an unsupported figure in a partner report is the problem this whole charter exists to address.

Denominator

The ten elements listed in 01-legal/ops-retainer-schedule.md clause 5.2, plus any element the statement of work adds. The denominator is printed with the figure, since it differs per statement of work.

Excluded

Elements that do not apply in the period, such as incidents in a month with none. These pass if the report states "none" explicitly, and fail if they are simply absent. The count of not-applicable elements is reported.

Scored separately

Timeliness, being delivery by the fifth business day of the following month. A complete report delivered late and an incomplete report delivered on time are different failures and are not averaged together.

Contract sentence

"Report completeness means the number of required Monthly Report elements that are present, cover the whole reporting period and carry the sample size or the basis for each figure, divided by the number of elements required by the Retainer Schedule clause 5.2 and this SOW. The Provider shall achieve report completeness of 100 percent, an element that does not apply in the period being satisfied by an explicit statement to that effect, and shall deliver the Monthly Report by the fifth Business Day of the following month. A figure reported without its sample size or its basis does not satisfy the element it belongs to."

Difficulty tiers

Charter 1.0.0 section 4.5. A tier is a property of the item, assigned when it is generated, and it never changes because a system found the item hard.

The tier describes the exercise run against a deployment, not an item in a dataset. A Day-60 score is only comparable with another Day-60 score run at the same tier, and the tier is printed beside the score.

TierEnvironmentDrift injectedIncidentRollbackReport audit
T1Non-production copyOne class, announced windowOne Severity 3, scriptedRehearsed on the copyOne month
T2Live-like environment with production-shaped volumeOne class, unannounced within an agreed exercise periodOne Severity 2, scriptedPerformed during an agreed change windowOne month
T3Production, within the change-window rules of the statement of workTwo concurrent signals, for example input drift and a confidence shiftOne Severity 2 that requires a partner dependency to resolvePerformed, with the re-queue of affected items measuredThree months
T4ProductionA class with no labelled examplesA Severity 1 exercised as a tabletop onlyPerformed, including a data restore from backupThree months, spanning a provider model version change
T5ProductionAs T4, and the exercise is initiated by the partner without prior notice to the delivery team, inside the rules the statement of work allowsAs T4As T4Three months, spanning a provider model version change and a threshold change

4.5.1 Safety rules that hold at every tier, without exception. No exercise ever injects a real data exposure, exfiltration or a simulated breach involving live personal data. Every injection is agreed in writing in advance, has a named owner who can stop it, has a defined stop condition, and has a rollback owner on call for its duration. An exercise that cannot meet these conditions is not run, and the affected rubric lines are scored "not exercised", never assumed.

4.5.2 The self-assessment version a BPO runs against its current supplier is capped at T2, because the higher tiers require the ability to change a production configuration.

The data

This suite has no dataset. It is a rubric and a set of scripted exercises run against a live deployment, ours or anyone's, sixty days after go-live.

The public sample

Nothing to download. The rubric, the scripted incidents and the self-assessment are documents, listed under the reproduce section below.

Reproduce

This suite is not run from a command line. It is a set of scripted exercises run against a live deployment, scored against a published rubric by a person. The documents below hold the procedure.

The documents this suite is run from

  • Day-60 rubric v1.0.0day-60/rubric.md · 1. What this measures, and what it does not · 2. Tiers · 3. How the score is built · 4. Section A — drift detection and notice, 25 points · 5. Section B — incident response and communication, 25 points · 6. Section C — rollback, 20 points
  • Day-60 scripted exercises v1.0.0day-60/scripted-incidents.md · 1. Roles · 2. Safety rules, which hold at every tier without exception · 3. Exercise D-1 — drift injection · 4. Exercise I-1 — incident · 5. Exercise R-1 — rollback · 6. Exercise A-1 — monthly report audit
  • Day-60 self-assessmentday-60/self-assessment.md · What this is · What this is worth · Before the afternoon · How to score · Section A — do they notice when the work changes, 25 points · Section B — what happens when something breaks, 25 points

Charter 5.5

Every result is reproducible from a commit. A figure that cannot be reproduced from the dataset version and hash, the harness version and commit, the prompt set hash, the model version string, the run date, the price list date and the exact command line is withdrawn, not defended.

Run it on your own data.

Versions and hashes

Charter 5.5: a figure that cannot be reproduced from the items below is withdrawn rather than defended.
Datasetday-60/ rubric and scripted exercises — not a dataset
Dataset seednot applicable
Dataset version in the resultsnot run
Harness version1.0.0
Harness commitnot run
Scorer versionnot run
Charter version1.0.0
Ground-truth hashnot published for this suite
Dataset manifest hashnot published for this suite
Results folderdoes not exist — the suite has not been run
Run datenot run
Price list datenot run — no provider charge has been incurred

Charter 7.3

No table mixes versions. A row produced under a different dataset, harness or charter version sits in a different table, and a superseded table stays published, marked superseded, with a link to the one that replaced it.

Changelog · Dispute a figure on this page