Blind cross-event validation of the asset-dynamics cascade
This document reports, in full and without rounding toward a favourable answer, whether the per-asset dynamics framework's RoCoF-trip cascade has predictive power beyond calibration. It was written in response to the direct question: "we no longer diverge from the official report — is that because we're biased toward it, or because it's genuinely correct?" The honest answer is below, with the numbers.
TL;DR verdict
- The cascade mechanism is genuine (embedded units trip on their own RoCoF protection, unprompted).
- The cascade magnitude/depth is not predicted — it is set by the unmeasured embedded-trip volume, and the nadir swings ~2.7 Hz across plausible volumes.
- The cascade cannot be blind-validated on GB data at all, because 9 Aug 2019 is the only major GB embedded-generation cascade and the ALoMCP programme deliberately eliminated the mechanism afterwards. There is no held-out cascade event, by construction. This is structural, not a data-gathering gap.
- Therefore the cascade is a mechanism / envelope tool, permanently
UnvalidatedAtScaleon GB data — not a validated predictor. Use it for counterfactual/scenario/sensitivity work, not to assert a specific predicted cascade nadir as fact. - The base swing+response model (which predates this framework) can be cross-checked and lands in the right region for both recorded events with shared unfitted constants — but that is not the asset framework's contribution.
The protocol
A real blind test freezes the unmeasurable model parameters on one event and predicts a held-out event without touching them. Only recorded/measured inputs may differ between events (inertia, demand, loss magnitudes, procured response volumes, pre-event frequency — these are facts, not knobs).
Shared, unfitted constants (identical for every event): DroopPu 0.05, DampingPctPerHz 1.0,
ResponsiveFraction 0.20, GovernorTimeConstantS 8. Per-event inputs come entirely from
EventCatalog (Ofgem / settlement-metered / NESO-recorded).
Why the obvious held-out event (2026-05-31) is disqualified
reports/infeed_loss-20260531T180031-31869 flags a 992 MW infeed loss, nadir 49.806 Hz. But its recorded
peak RoCoF is +0.007 Hz/s — for a 992 MW step at the recorded 183 GVA·s, step physics demands
50·992/(2·183000) = 0.135 Hz/s, ~19× larger. The report's own physics cross-check says the swing-equation
inertia is 19.32× the NESO value — i.e. this was a gradual imbalance, not a step loss. You cannot
step-predict an event that was not a step. Disqualified. (For contrast, 2019-08-09's closure is 1.25×,
and 2023-12-22's is consistent — both obey step physics.)
Results — base-model cross-event replay (numbers as measured)
| Event | Type | Inputs (recorded) | Recorded nadir | Model nadir | Residual |
|---|---|---|---|---|---|
| 2019-08-09 | replay (all losses + recorded LFDD) | 215 GVA·s, 21.1 GW, governor arrest | 48.80 Hz | 48.82 Hz | 0.02 Hz |
| 2023-12-22 | forward (measured losses + measured DC/DM/DR, no LFDD, no fitting) | 161 GVA·s, 22.4 GW, 1226 MW response | 49.275 Hz | 48.98 Hz | 0.30 Hz too deep |
Parameter-free physics (RoCoF = f0·ΔP/(2E), zero free parameters), first-1s loss: 2019 → 0.132 Hz/s;
2023 → 0.155 Hz/s. Both are the right order and consistent with the recorded window-averaged figures
(2019 −0.151/−0.17; 2023 −0.102 on a 1 s window; 1 s averaging lowers a sharp step's peak).
Reading these honestly:
- 2019 (0.02 Hz) is replay, not prediction. We fed it the recorded losses and the recorded LFDD arrest; reproducing recorded facts is not evidence of predictive power. It confirms the integrator only.
- 2023 (0.30 Hz too deep) is the genuinely forward number. No LFDD occurred, the losses and response
volumes are measured, nothing was fitted to the nadir — and the model under-arrests by ~0.3 Hz. That is
the right region (a ~1.3 GW loss on a light-inertia system correctly produces a shallow ~0.3–0.7 Hz dip),
but it is not tight, and it exercises the base arrest model (damping + response time constants), not
the asset framework. The residual is pinned in
BlindPredictionTestsso a regression would surface.
Why the cascade itself cannot be blind-tested (the structural argument)
The asset framework's novel capability is the RoCoF-trip cascade. Validating it needs an event where embedded generation actually tripped on RoCoF. In the GB record:
- 9 Aug 2019 is the only major GB embedded-generation RoCoF cascade. ~350–500 MW of distributed generation tripped as frequency fell, deepening the nadir (Ofgem 2.4.x).
- The Accelerated Loss-of-Mains Change Programme (ALoMCP) then raised embedded relays from 0.125 Hz/s to 1.0 Hz/s with a 500 ms delay, and EREC G99 phased out vector-shift — for the express purpose of removing this mechanism. It was substantially complete by ~2022.
- Consequently every later GB event (2023-12-22, 2026-05-31, …) has no cascade: modern relays ride
through.
BlindPredictionTests.The_cascade_has_no_second_gb_event_to_blind_predictpins this — 2019 carries "distributed generation" RoCoF-trip losses; 2023 carries none.
So there is no held-out GB cascade event, and there never will be one unless the mechanism recurs
(which policy is designed to prevent). The cascade model can be calibrated to 2019 (fleet size → nadir),
but its predictive accuracy for a cascade cannot be independently checked on GB data. Any single-event
"agreement" is calibration; the framework's own honesty tags (UnvalidatedAtScale, [VERIFY]) are the
correct and permanent labels.
What the cascade capability actually is (and isn't)
Is: a mechanistic, physically-grounded model of how an embedded-generation RoCoF cascade unfolds — the trips emerge from local-RoCoF protection (uplifted ~1.16× per the sub-cycle grounding), feed back into the deficit, and deepen the fall. Correct for: counterfactuals ("what if X GW of DG were still on 0.125 Hz/s relays"), sensitivity/envelope studies, screening, and reproducing the 2019 mechanism for explanation.
Isn't: a validated predictor of a specific cascade nadir/RoCoF. The embedded-trip volume is unmeasured, the depth is strongly sensitive to it, and there is no second event to test transfer. Presenting a point-predicted cascade outcome as validated fact would be exactly the pseudo-precision the framework is built to avoid.
What would change this verdict
Only new measured evidence: (a) per-feeder embedded-DG trip telemetry for a real event (not in any
public GB dataset), or (b) a future recorded GB cascade (which ALoMCP is designed to prevent). Absent
those, the honest ceiling is: reproduce 2019's mechanism, bracket its outcome in an envelope, and label it
UnvalidatedAtScale. See also docs/technical/13-asset-dynamics-data-grounding.md.