Suite · AI operations
Day-60
Whether the system is still trustworthy two months after go-live.
Charter 1.0.0, section 2
Unit scored: one deployment. Headline metric: day-60 score, 0 to 100.
Built from day-60/ rubric and scripted exercises.
On this page: leaderboard, metric definitions, difficulty tiers, reproduce, versions and hashes.
Status
Leaderboard
Selected slice: every tier, every language. No deployment has been exercised at any tier, so the slice carries no count. A Day-60 score is only comparable with another Day-60 score run at the same tier. The filter changes the denominator a row would be measured over. Every measured cell reads not run in every slice.
No sort applied. Rows are in source order.
| No deployment has been exercised. A Day-60 score needs a live deployment, a partner and an exercise window agreed in writing under charter 4.5.1. | ||||
Column headings sort the table. A figure that was never produced sorts last in both directions; it is not a low score. Charter 5.6: our own system is scored by the same harness, from the same commit, with the same prompts, and sits wherever the sort puts it, with no highlight and no position advantage.
Nothing to plot. No deployment has been exercised, so there is no Day-60 score to draw. Charter 10.4: a chart with no run renders an empty state, not example bars.
The rubric
A Day-60 score is 100 points against the rubric below, run against a live deployment sixty days after go-live. Every observed value is empty because no deployment has been exercised. Charter 4.5.1: an exercise that cannot meet the safety conditions is not run, and the affected rubric lines are scored "not exercised", never assumed.
17 scoring lines shown, 100 points of the 100 the rubric allocates.
No sort applied. Rows are in source order.
| A1 | Detection lead time | ratio | 8 | not exercised | not exercised |
|---|---|---|---|---|---|
| A2 | Notice lead time | ratio | 8 | not exercised | not exercised |
| A3 | Notice content | checklist | 5 | not exercised | not exercised |
| A4 | False-alert accounting | computed | 4 | not exercised | not exercised |
| B1 | Acknowledgement | ratio | 5 | not exercised | not exercised |
| B2 | Time to restore | ratio | 10 | not exercised | not exercised |
| B3 | Communication during the incident | checklist | 5 | not exercised | not exercised |
| B4 | Post-incident record | checklist | 5 | not exercised | not exercised |
| C1 | Decision time | ratio | 4 | not exercised | not exercised |
| C2 | Rollback time | ratio | 8 | not exercised | not exercised |
| C3 | Re-queue completion | ratio | 5 | not exercised | not exercised |
| C4 | Restored configuration proven | checklist | 3 | not exercised | not exercised |
| D1 | Completeness | computed | 12 | not exercised | not exercised |
| D2 | Changes stated with their dates | checklist | 4 | not exercised | not exercised |
| D3 | Timeliness | banded | 4 | not exercised | not exercised |
| E1 | Timestamps come from system records | checklist | 5 | not exercised | not exercised |
| E2 | A reported figure can be reproduced | checklist | 5 | not exercised | not exercised |
A ratio line scores full points where the observed value meets the contractual target, half where it is up to twice the target, and zero beyond that or where the contract carries no target. The bands are in day-60/rubric.md section 3.3.
| Field | What goes in it | Recorded |
|---|---|---|
| Tier exercised | One of T1 to T5, from rubric.md section 2. Print it beside the score everywhere the score appears (charter 4.5 and 9.9). | not exercised |
| Exercise window | Start and end dates of the exercise. Fixes the meaning of 'not achieved within the window' on every ratio line. | not exercised |
| Assessors | Two, named. Each scores every line independently from the same evidence pack. Disagreements resolve under rubric.md 3.7. | not exercised |
| Rules that apply to every row | 1. A line with no evidence_reference scores 0, whatever the assessors believe happened (rubric.md 3.8). 2. A line that could not be exercised is written 'not exercised' in band, scores no points, and is removed from points_achievable; it is never assumed. 3. A line scored 0 because the contract sets no target is written 'no target in the contract' in note and listed separately in the exercise report. 4. This file is shipped with no scores in it and stays that way; no Day-60 exercise has been run. | not exercised |
| Section | Points available | Points awarded |
|---|---|---|
| Section A subtotal, drift detection and notice | 25 | not exercised |
| Section B subtotal, incident response and communication | 25 | not exercised |
| Section C subtotal, rollback | 20 | not exercised |
| Section D subtotal, monthly report completeness | 20 | not exercised |
| Section E subtotal, evidence and reproducibility | 10 | not exercised |
| Points awarded and points achievable | 100 | not exercised |
| Day-60 score | 100 | not exercised |
Metric definitions
Read from charter 1.0.0 section 3 when this page was built. Each metric carries its formula, its numerator, its denominator and its exclusions, and each has clause language in charter/contract-clauses.md so it can be written into a statement of work. The charter's arithmetic examples are left out here: they are invented numbers that demonstrate a formula, and beside a table of "not run" they would read as results.
Drift detection lead time
Charter 1.0.0 section 3.18 · read it in the charter · SOW clause 17
- Definition
The elapsed time from the first item affected by a drift to the supplier telling the partner about it in writing. Two clocks are reported, because they answer different questions.
Clock From To The question it answers Detection First affected item The internal alert or the ticket that records the drift Did the monitoring see it Notice First affected item The written notice to the partner naming the class and the evidence Did anyone act on what the monitoring saw - Formula
Lead time = the notice timestamp minus the onset timestamp, reported in hours and in business hours, with the detection clock reported beside it. Numerator and denominator: a single lead time is an elapsed time and has neither. Where a report gives a mean over several drift observations, the numerator is the sum of the lead times and the denominator is the count of drift observations in the window, and both are printed, together with the maximum.
- Onset
The timestamp of the first input item belonging to the shifted population, established from the record after the fact. In a Day-60 exercise the injection timestamp is known in advance and held by the person running the exercise, not by the team being measured.
- Excluded
Drift the partner announced in advance, such as a scheduled format change, which is a planned change and not a detection. Injections the supplier was told about. Signals raised and withdrawn as false, which are counted separately as a false-alert count, because a monitoring system that alerts constantly also has a lead time near zero and is worth nothing.
- Where the drift was not found
If the exercise window closes with no notice, the figure is written "not detected within the window", with the window length stated. It is never extrapolated to a longer window.
- Contract sentence
"Drift detection lead time means the elapsed time from the first Item belonging to a shifted input population to the Provider's written notice to the Partner naming the affected class and the supporting evidence. The Provider shall give that notice within placeholder business hours of onset, and shall report both the detection and the notice timestamps for every drift observation in the Monthly Report. A drift not notified within placeholder business hours of onset is a Severity 2 Incident under the Retainer Schedule clause 4."
Incident mean time to restore
Charter 1.0.0 section 3.19 · read it in the charter · SOW clause 18
- Definition
The mean elapsed time to restore service, computed separately for each severity in the Retainer Schedule clause 4. Reported with the count, the median and the maximum, because a mean over three incidents conceals the worst of them.
- Formula
Mean time to restore for a severity = the sum of restoration durations for incidents of that severity closed in the window ÷ the count of those incidents. The numerator is that sum of durations, measured as 3.19.3 and 3.19.4 define. The denominator is the count of incidents of that severity closed in the window, and it is printed with the figure, because a mean over two incidents is not a rate.
- Start of the clock
The earlier of the supplier's own detection timestamp and the partner's notice timestamp. Not the time the ticket was opened, and not the time an engineer picked it up.
- End of the clock
The time the covered system returned to the state where work submitted at intake reaches an output or the review queue within the service level, confirmed by a monitoring check. Not the time the ticket was closed, and not the time the root cause was found.
- Excluded
Incidents excluded under the Retainer Schedule clause 10. Incidents raised and withdrawn as not reproducible, counted separately. Time inside an agreed change window. Time waiting on a partner dependency, which stops the clock and is reported as a separate "waiting on partner" total, so that neither party can hide inside the other's delay.
- Contract sentence
"Mean time to restore means the sum of restoration durations for Incidents of a given Severity closed in the measurement window, divided by the number of such Incidents, where the clock starts at the earlier of the Provider's detection and the Partner's notice and stops when a monitoring check confirms that work submitted at intake reaches an output or the human review queue within the Service Level. The Provider shall meet the resolution targets in the Retainer Schedule clause 3 and shall report, for each Severity, the count, the mean, the median and the maximum, together with any time excluded as waiting on a Partner dependency."
Rollback time
Charter 1.0.0 section 3.20 · read it in the charter · SOW clause 19
- Definition
The elapsed time from the decision to roll back to the previous release serving production traffic and passing the smoke set with its model pin, prompt version and threshold set restored. Two further times are reported beside it: decision time, and re-queue completion.
Time From To Decision time The rollback criterion being met, per 06-delivery/build-standards.mdsection 9The on-call engineer recording the decision Rollback time The decision being recorded The smoke set passing on the restored release Re-queue completion The decision being recorded Every item processed by the rolled-back release identified and re-queued for review - Formula
Rollback time = the smoke-set pass timestamp minus the decision timestamp, in minutes. Numerator and denominator: a single rollback is an elapsed time and has neither. Where a report covers more than one rollback, the numerator is the sum of the rollback times and the denominator is the count of rollbacks in the period, reported separately for rehearsals and for rollbacks performed on production, with the maximum of each.
- Why three times
A rollback that restores the service quickly but leaves the outputs of a bad release sitting in the partner's downstream systems is not finished. A rollback that takes four minutes after a two-hour argument is not fast. Publishing one number and not the other three is the usual way this measure is made to look good.
- Excluded and reported separately
A rollback that also requires a data restore reports the restore time separately under
02-security/security-policy-set/07-backup-and-recovery.md. Rehearsals on a non-production copy are labelled as rehearsals and are never mixed in a table with rollbacks performed on production.- Contract sentence
"Rollback time means the elapsed time from the recording of a rollback decision to the previous signed release serving production traffic and passing the smoke set with its model pin, prompt version and threshold set restored. The Provider shall complete a rollback within placeholder minutes, shall rehearse a rollback at least once per quarter and record the rehearsal result in the Monthly Report, and shall report decision time and re-queue completion time alongside rollback time for every rollback performed."
<!-- Benchmark charter, part 4. Indexed in ../methodology.md. Sections 3.21 to 7. -->
Report completeness
Charter 1.0.0 section 3.21 · read it in the charter · SOW clause 20
- Definition
The share of required elements of the monthly report that are present, cover the whole period, and carry the basis or the sample size where the element is a figure.
- Formula
Report completeness = elements present and passing ÷ elements required.
- Numerator
Elements that are present, cover the whole reporting period, and are supported. An element that is present but carries a figure without its sample size or its basis fails, because an unsupported figure in a partner report is the problem this whole charter exists to address.
- Denominator
The ten elements listed in
01-legal/ops-retainer-schedule.mdclause 5.2, plus any element the statement of work adds. The denominator is printed with the figure, since it differs per statement of work.- Excluded
Elements that do not apply in the period, such as incidents in a month with none. These pass if the report states "none" explicitly, and fail if they are simply absent. The count of not-applicable elements is reported.
- Scored separately
Timeliness, being delivery by the fifth business day of the following month. A complete report delivered late and an incomplete report delivered on time are different failures and are not averaged together.
- Contract sentence
"Report completeness means the number of required Monthly Report elements that are present, cover the whole reporting period and carry the sample size or the basis for each figure, divided by the number of elements required by the Retainer Schedule clause 5.2 and this SOW. The Provider shall achieve report completeness of 100 percent, an element that does not apply in the period being satisfied by an explicit statement to that effect, and shall deliver the Monthly Report by the fifth Business Day of the following month. A figure reported without its sample size or its basis does not satisfy the element it belongs to."
Difficulty tiers
Charter 1.0.0 section 4.5. A tier is a property of the item, assigned when it is generated, and it never changes because a system found the item hard.
The tier describes the exercise run against a deployment, not an item in a dataset. A Day-60 score is only comparable with another Day-60 score run at the same tier, and the tier is printed beside the score.
| Tier | Environment | Drift injected | Incident | Rollback | Report audit |
|---|---|---|---|---|---|
| T1 | Non-production copy | One class, announced window | One Severity 3, scripted | Rehearsed on the copy | One month |
| T2 | Live-like environment with production-shaped volume | One class, unannounced within an agreed exercise period | One Severity 2, scripted | Performed during an agreed change window | One month |
| T3 | Production, within the change-window rules of the statement of work | Two concurrent signals, for example input drift and a confidence shift | One Severity 2 that requires a partner dependency to resolve | Performed, with the re-queue of affected items measured | Three months |
| T4 | Production | A class with no labelled examples | A Severity 1 exercised as a tabletop only | Performed, including a data restore from backup | Three months, spanning a provider model version change |
| T5 | Production | As T4, and the exercise is initiated by the partner without prior notice to the delivery team, inside the rules the statement of work allows | As T4 | As T4 | Three months, spanning a provider model version change and a threshold change |
4.5.1 Safety rules that hold at every tier, without exception. No exercise ever injects a real data exposure, exfiltration or a simulated breach involving live personal data. Every injection is agreed in writing in advance, has a named owner who can stop it, has a defined stop condition, and has a rollback owner on call for its duration. An exercise that cannot meet these conditions is not run, and the affected rubric lines are scored "not exercised", never assumed.
4.5.2 The self-assessment version a BPO runs against its current supplier is capped at T2, because the higher tiers require the ability to change a production configuration.
The data
This suite has no dataset. It is a rubric and a set of scripted exercises run against a live deployment, ours or anyone's, sixty days after go-live.
The public sample
Nothing to download. The rubric, the scripted incidents and the self-assessment are documents, listed under the reproduce section below.
Reproduce
This suite is not run from a command line. It is a set of scripted exercises run against a live deployment, scored against a published rubric by a person. The documents below hold the procedure.
The documents this suite is run from
- Day-60 rubric v1.0.0 —
day-60/rubric.md· 1. What this measures, and what it does not · 2. Tiers · 3. How the score is built · 4. Section A — drift detection and notice, 25 points · 5. Section B — incident response and communication, 25 points · 6. Section C — rollback, 20 points - Day-60 scripted exercises v1.0.0 —
day-60/scripted-incidents.md· 1. Roles · 2. Safety rules, which hold at every tier without exception · 3. Exercise D-1 — drift injection · 4. Exercise I-1 — incident · 5. Exercise R-1 — rollback · 6. Exercise A-1 — monthly report audit - Day-60 self-assessment —
day-60/self-assessment.md· What this is · What this is worth · Before the afternoon · How to score · Section A — do they notice when the work changes, 25 points · Section B — what happens when something breaks, 25 points
Charter 5.5
Every result is reproducible from a commit. A figure that cannot be reproduced from the dataset version and hash, the harness version and commit, the prompt set hash, the model version string, the run date, the price list date and the exact command line is withdrawn, not defended.
Versions and hashes
| Dataset | day-60/ rubric and scripted exercises — not a dataset |
|---|---|
| Dataset seed | not applicable |
| Dataset version in the results | not run |
| Harness version | 1.0.0 |
| Harness commit | not run |
| Scorer version | not run |
| Charter version | 1.0.0 |
| Ground-truth hash | not published for this suite |
| Dataset manifest hash | not published for this suite |
| Results folder | does not exist — the suite has not been run |
| Run date | not run |
| Price list date | not run — no provider charge has been incurred |
Charter 7.3
No table mixes versions. A row produced under a different dataset, harness or charter version sits in a different table, and a superseded table stays published, marked superseded, with a link to the one that replaced it.