Expert Detection of Decision Reconstruction Risk
Decision Reconstruction Risk (DRR) is the condition in which a record cannot explain, on its own terms, why a consequential decision was made. This page summarises what the programme has measured about that construct, in the order a reviewing peer would want it: the boundaries of the claim first, the result second, the statistics that produced it third. Every figure below comes from a study whose thresholds were fixed before any data were examined, and one of those thresholds was not met. That failure is reported here in the same type size as the result that succeeded.
- The corpus is deliberately bimodal. Twelve clearly grounded and twelve clearly unsupported records, constructed at the two ends of the severity range so the detection question could be asked cleanly. Real documentation is continuous. A test evaluated on clear-cut cases overstates field performance, and this one is no exception.
- There is no criterion validity against real records. The stimuli are constructed and AI-generated. Nothing here shows that the same accuracy holds on real workplace files, on human-authored records, or on ambiguous ones. That study is separate and open.
- The five review conditions are not psychometrically validated. They were derived from investigative casework, not from a factor-analytic procedure. Whether they are distinguishable rather than redundant, whether all five are necessary, and whether they form a coherent scale are all open questions. Read them as a checklist with face validity, not as a validated multidimensional scale.
- One pre-registered criterion was not met. The reliability criterion had two parts: a point estimate of at least 0.61 and a lower confidence bound of at least 0.41. The point estimates clear the first part. Neither clears the second on the analytic interval that the plan specified. Details below.
- Group-level detectability is not individual-level dependability. Accuracy at the level of a single reviewer ranged from 37.5 to 100 percent with a standard deviation of 21.0. An organisation routing one record to one reviewer cannot treat the panel mean as that reviewer's expected performance.
Nothing on this page is presented as validated. The programme is in its operational validation phase, and the claim control that governs it is published in full in the Methodology Archive.
95% confidence interval 72.7 to 95.1 at participant level. Sensitivity 87.0 percent for unsupported records, specificity 80.7 percent for grounded ones. Produced by 16 independent experts across 11 countries and 5 continents, reading cold and blind, over 384 graded reads.
Three estimators carry the weight of this result. Each is stated here with the choice behind it, because a summary that reports a coefficient without its construction cannot be checked by the reader it is trying to persuade.
The 83.9 percent figure is the mean of the 16 individual reviewer accuracy scores, each scored out of 24 records. It is not a pooled proportion over the 384 reads, and it is not an estimate of field performance. The reviewer is the unit of observation, which avoids treating 384 correlated reads as 384 independent trials.
83.9 ± t(15, .975) × 21.0 / √16 = 72.7 to 95.1The interval is a Student t interval on the 16 reviewer scores, with 15 degrees of freedom. It accounts for variation between reviewers. It does not account for the possibility that this particular draw of 24 records was easier than another draw from the same construct, which is the limitation the crossed model below was fitted to size.
Agreement is reported as Gwet's AC1 rather than Cohen's or Fleiss' kappa. The reason is the prevalence problem: when raters agree that most records fall in the same category, kappa collapses toward zero even as raw agreement stays high, and the resulting coefficient describes the marginal distribution more than it describes the raters. AC1 uses a chance-agreement term that does not behave that way under skewed prevalence.
On a shared 10-record set carrying 113 submitted determinations, reduced to 104 after keeping one label per rater per record: 0.739 for invited participants, 0.623 for open-enrolment participants. Both are reported as interim. They rest on 10 records against a pooled target of about 26, the intervals are wide, and the pre-registered lower bound is not cleared. Earlier versions of this work placed 0.739 in the substantial band of Landis and Koch. That characterisation has been dropped: the band boundaries are acknowledged as arbitrary by their own authors, and the adjective added a suggestion of adequacy that the failed criterion contradicts.
Per-condition association is reported in the study appendix with Wilson score intervals (Wilson, 1927) rather than the normal approximation. The reason is cell size: Ready determinations number 14 in this corpus, so every rate in that column rests on 14 observations. The normal approximation misbehaves badly at those counts and can produce bounds outside the unit interval; the Wilson construction stays inside it and keeps nominal coverage at small n. The intervals are correspondingly wide, and that width is the honest report of what 14 observations can support.
The primary analysis treats the reviewer as the unit of observation and does not model record difficulty. To size that omission rather than argue about it, a mixed-effects logistic model was fitted over all 384 graded reads, crossing both random factors:
correct ~ 1 + (1 | reviewer) + (1 | record)Fitted by Laplace-approximated maximum likelihood, with no participant excluded by the pre-registered rule. The intercept places an average reviewer on an average record at 89.2 percent. Reviewer SD is 1.769 (profile interval 1.292 to 3.000); record SD is 0.011 and sits at the boundary, with a profile interval of 0.001 to 0.556. The latent-scale intraclass correlation is 0.488 for reviewers and 0.0000 for records.
Two things follow, and only two. Modelling item difficulty does not materially move the panel estimate, because the participant-level mean sits inside the reviewer-level spread the model estimates. And the variation that matters on this corpus is between reviewers, not between records, which is the same conclusion the raw dispersion reaches by a different route. This analysis was added after pre-registration and is labelled exploratory throughout. The record component is weakly identified, and its profile interval permits a materially larger record effect than the point estimate suggests.
This research concerns a construct and the artifacts that carry it. It does not concern a product, and the separation is kept deliberately sharp so a procurement reviewer never has to guess which side of the line a sentence sits on.
What this research addresses
- Whether an operationalised documentation-risk construct is detectable by independent experts
- Artifact-level risk: what a finalised record does and does not carry on its own terms
- Whether independent reviewers apply the five conditions consistently with one another
- How consistently independent AI models from different vendors apply the same conditions; this is measured as raw agreement and does not by itself establish the pre-registered reproducibility criterion
- How much of the observed variation belongs to reviewers rather than to records
What it does not address
- Whether any tool prevents Decision Reconstruction Risk from arising
- Whether automated review substitutes for a qualified human reviewer
- Whether flagged records produce different real-world outcomes when contested
- Compliance status under any external framework, including the EU AI Act, the NIST AI RMF and ISO/IEC 42001
- Any measure of return on investment, cost avoidance or litigation risk reduction
The research above supports one operational conclusion and licenses one design response. The conclusion: a record can fail to explain its own decision in ways a trained reader can identify, reliably enough to be worth looking for, and unevenly enough that a single reader is not a control. The design response follows directly from the dispersion, which is why every asset below assumes calibration and second review rather than a single pass.