Readiness Scoring

How an exercise becomes a 0 to 100 readiness score: the eight capability areas, how gaps and decisions move them, the limits on single-participant runs, and what the score is compared against.

A readiness score is computed when an exercise is completed and stored with that exercise. The exercise results page recomputes it from the exercise's current gaps, so remediating a gap shows there straight away; the Readiness page, board report and certificates use the stored score.

The score measures what an exercise recorded: the gaps found and the decisions made. It does not inspect your environment, so a high score means the exercise surfaced few weaknesses, not that your controls have been tested.

The eight capability areas

| Capability area | Weight | What lowers it | |---|---|---| | Detection & Monitoring | 15% | Detection gaps | | Crisis Communications | 15% | Communications gaps | | Containment & Response | 15% | Containment and technology gaps | | Operational Recovery | 15% | Recovery gaps | | Decision Speed | 10% | Low stated confidence, overrunning the planned duration | | Executive Alignment | 10% | Process gaps | | Incident Command | 10% | Escalation and personnel gaps | | Regulatory Readiness | 10% | Legal / regulatory and documentation gaps |

Weights sum to 100%. Each area is scored 0 to 100. The weighted average of the eight is the starting point for the overall score; the adjustments below can only lower it.

How gaps lower a score

Every gap recorded on the exercise costs points in the area its category maps to (table above). A gap whose category maps to no area costs the same total, spread evenly across all eight.

| Gap severity | Points | |---|---| | Critical | 25 | | High | 15 | | Medium | 8 | | Low | 3 | | Info | 0 |

A remediated gap still costs half its points: it was found, and fixing it is recorded separately from never having had it. An area cannot go below 0.

Operational Recovery starts 3 points lower than the other areas. The engine can adjust it from a business impact analysis (recovery point and time objectives per function), but no exercise score supplies one today, so every score carries the 3-point deduction.

Decisions, confidence and timing

Each decision carries the participant's stated confidence, 1 to 5. A decision submitted without one counts as 3.

  • Decision Speed starts from the average stated confidence on a 0 to 100 scale (1 = 0, 3 = 50, 5 = 100), or 50 when there are no decisions. It gains 10 points when the session finished within its planned duration and loses 10 when it ran more than 1.5 times longer. An exercise with no planned duration is treated as 90 minutes.
  • Executive Alignment gains 5 points when the average stated confidence is 70 or higher on that scale.

Confidence is self-reported. It describes how sure participants said they were, not whether they were right.

Single-participant runs

A run with one participant (a self-service exercise) grades one person's stated decisions. It cannot show whether the controls those decisions rely on work, or whether a team would carry them out. Its overall score is adjusted in three ways that a team exercise is not:

  1. Capped at 79. The cap applies to the score the run would have had with no gaps; the cost of its gaps is then taken off that. A run that records harmful decisions therefore scores below a clean one rather than both sitting at the cap.
  2. Recorded gaps. When the coach recognises a harmful call, such as paying a ransom before sanctions screening or giving an all-clear before forensics, it is recorded as a gap titled "Exercise decision: …" with the severity of that call. Every inject category without such a gap gets a "Verify:" gap instead, naming the control the decisions assumed (for example, that a backup restore has actually been tested). Every single-participant run also gets a low-severity gap to validate the decisions with a team exercise. These are real gaps and lower the score like any other.
  3. Answer coverage, up to −15. Each answer is compared with the inject's expected response, split into its separate actions. An action counts as covered when at least half of its key words appear in the decision or its rationale. Injects whose expected response lists fewer than two actions are not counted. An average coverage of 60% or more costs nothing; below that, each percentage point costs a quarter of a point, to a maximum of 15.

The results page shows each adjustment beside the capability bars, so the overall and the bars can be read together. Remediating gaps does not recover the answer-coverage deduction; re-running the exercise does.

Because a single-participant run is capped at 79, the Optimized band requires a team exercise.

Open critical gaps

While any critical gap is open, the overall score is capped at 64, whatever the other areas score. This applies to every exercise. It lifts when the critical gaps are remediated.

Bands

| Band | Overall score | |---|---| | Optimized | 80 and above | | Strong | 70 to 79 | | Developing | 65 to 69 | | At Risk | 40 to 64 | | Critical | below 40 |

The same table is used on the results page, the Readiness page, the dashboard and the board report.

What a score is compared against

Reference baselines. Each sector has a fixed reference baseline per area (financial services, healthcare, energy, government, technology, and a general baseline for any other sector). These are illustrative starting points, not measured peer data, and are labelled "est." wherever they appear. The exercise results page always uses the reference baseline for your sector.

Live peer benchmark. On the Readiness page and in the board report, once at least five other organizations in your sector have readiness scores, the reference baseline is replaced by the average of each of those organizations' latest score. Your own organization is excluded. A percentile rank is shown only in that case; with fewer than five peers it is withheld. The board report states which of the two it used.

Benchmarks are not segmented by organization size or by exercise difficulty.

Prescriptions

Each area is rated against its score and its baseline:

  • Critical below 40
  • High below 60, or more than 15 below the baseline
  • Moderate below 80, or below the baseline
  • On track otherwise

Each rating carries a fixed set of recommended actions for that area, with the framework references they derive from, and names the gaps behind it. Recommendations are ranked by how many points remediating the related gaps would add. "Improvement available" is the overall gain if every open gap were remediated, computed by re-scoring the exercise, not by adding up the areas.

Score history and comparisons

Every stored score is kept, one per exercise, and the Readiness page plots them over time. A change is shown only between scores produced by the same scoring method. Single-participant scores stored before 4 October 2026 were computed under the previous method (a flat cap at 79, without recorded decision gaps or the answer-coverage adjustment). They stay in the history, but no delta or trend line is drawn between them and later scores; they are marked "Scored under the previous method; not compared".

Scores do not decay with time. An old score stays as it was stored; the date beside it says how old it is.