Evidence Development Program · Research Summary

Expert Detection of Decision Reconstruction Risk

Decision Reconstruction Risk (DRR) is the condition in which a record cannot explain, on its own terms, why a consequential decision was made. This page summarises what the programme has measured about that construct, in the order a reviewing peer would want it: the boundaries of the claim first, the result second, the statistics that produced it third. Every figure below comes from a study whose thresholds were fixed before any data were examined, and one of those thresholds was not met. That failure is reported here in the same type size as the result that succeeded.

Five boundaries on every figure on this page
  • The corpus is deliberately bimodal. Twelve clearly grounded and twelve clearly unsupported records, constructed at the two ends of the severity range so the detection question could be asked cleanly. Real documentation is continuous. A test evaluated on clear-cut cases overstates field performance, and this one is no exception.
  • There is no criterion validity against real records. The stimuli are constructed and AI-generated. Nothing here shows that the same accuracy holds on real workplace files, on human-authored records, or on ambiguous ones. That study is separate and open.
  • The five review conditions are not psychometrically validated. They were derived from investigative casework, not from a factor-analytic procedure. Whether they are distinguishable rather than redundant, whether all five are necessary, and whether they form a coherent scale are all open questions. Read them as a checklist with face validity, not as a validated multidimensional scale.
  • One pre-registered criterion was not met. The reliability criterion had two parts: a point estimate of at least 0.61 and a lower confidence bound of at least 0.41. The point estimates clear the first part. Neither clears the second on the analytic interval that the plan specified. Details below.
  • Group-level detectability is not individual-level dependability. Accuracy at the level of a single reviewer ranged from 37.5 to 100 percent with a standard deviation of 21.0. An organisation routing one record to one reviewer cannot treat the panel mean as that reviewer's expected performance.

Nothing on this page is presented as validated. The programme is in its operational validation phase, and the claim control that governs it is published in full in the Methodology Archive.

83.9%
Panel detection accuracy against a reference classification fixed before scoring

95% confidence interval 72.7 to 95.1 at participant level. Sensitivity 87.0 percent for unsupported records, specificity 80.7 percent for grounded ones. Produced by 16 independent experts across 11 countries and 5 continents, reading cold and blind, over 384 graded reads.

Pre-registered detection threshold
Met on both parts
A point estimate of at least 70 percent, and a lower confidence bound above chance. Both fixed before any data were seen, and both cleared.
Pre-registered reliability criterion
Not met
Measured on a separate, smaller sample than the detection panel: 22 raters over the 10 records that carried two or more raters, 104 labels. Gwet's AC1 0.739 for invited participants and 0.623 for open-enrolment participants both exceed the point floor of 0.61. The required lower bound of 0.41 is not cleared on the analytic interval: 0.402 for invited participants, 0.252 for open enrolment. The 83.9 percent detection figure is not a reliability measure.
Reviewer dispersion
37.5 to 100%
Accuracy across individual reviewers, SD 21.0. Six of the panel scored 100 percent. The spread is reported as a finding, not as noise around the mean.
Programme reviewer base
58 reviewers
Have graded at least one record across the three studies, unpaid and in a personal capacity. 36 of them (all completers, both arms) completed a full 24-record set across 16 countries. The detection panel is a subset of that base, not the whole of it.

Three estimators carry the weight of this result. Each is stated here with the choice behind it, because a summary that reports a coefficient without its construction cannot be checked by the reader it is trying to persuade.

1. The participant-level mean and its interval

The 83.9 percent figure is the mean of the 16 individual reviewer accuracy scores, each scored out of 24 records. It is not a pooled proportion over the 384 reads, and it is not an estimate of field performance. The reviewer is the unit of observation, which avoids treating 384 correlated reads as 384 independent trials.

83.9 ± t(15, .975) × 21.0 / √16 = 72.7 to 95.1

The interval is a Student t interval on the 16 reviewer scores, with 15 degrees of freedom. It accounts for variation between reviewers. It does not account for the possibility that this particular draw of 24 records was easier than another draw from the same construct, which is the limitation the crossed model below was fitted to size.

2. Gwet's AC1 for inter-rater agreement

Agreement is reported as Gwet's AC1 rather than Cohen's or Fleiss' kappa. The reason is the prevalence problem: when raters agree that most records fall in the same category, kappa collapses toward zero even as raw agreement stays high, and the resulting coefficient describes the marginal distribution more than it describes the raters. AC1 uses a chance-agreement term that does not behave that way under skewed prevalence.

On a shared 10-record set carrying 113 submitted determinations, reduced to 104 after keeping one label per rater per record: 0.739 for invited participants, 0.623 for open-enrolment participants. Both are reported as interim. They rest on 10 records against a pooled target of about 26, the intervals are wide, and the pre-registered lower bound is not cleared. Earlier versions of this work placed 0.739 in the substantial band of Landis and Koch. That characterisation has been dropped: the band boundaries are acknowledged as arbitrary by their own authors, and the adjective added a suggestion of adequacy that the failed criterion contradicts.

3. Wilson score intervals on the per-condition rates

Per-condition association is reported in the study appendix with Wilson score intervals (Wilson, 1927) rather than the normal approximation. The reason is cell size: Ready determinations number 14 in this corpus, so every rate in that column rests on 14 observations. The normal approximation misbehaves badly at those counts and can produce bounds outside the unit interval; the Wilson construction stays inside it and keeps nominal coverage at small n. The intervals are correspondingly wide, and that width is the honest report of what 14 observations can support.

4. The crossed mixed-effects model Exploratory

The primary analysis treats the reviewer as the unit of observation and does not model record difficulty. To size that omission rather than argue about it, a mixed-effects logistic model was fitted over all 384 graded reads, crossing both random factors:

correct ~ 1 + (1 | reviewer) + (1 | record)

Fitted by Laplace-approximated maximum likelihood, with no participant excluded by the pre-registered rule. The intercept places an average reviewer on an average record at 89.2 percent. Reviewer SD is 1.769 (profile interval 1.292 to 3.000); record SD is 0.011 and sits at the boundary, with a profile interval of 0.001 to 0.556. The latent-scale intraclass correlation is 0.488 for reviewers and 0.0000 for records.

Two things follow, and only two. Modelling item difficulty does not materially move the panel estimate, because the participant-level mean sits inside the reviewer-level spread the model estimates. And the variation that matters on this corpus is between reviewers, not between records, which is the same conclusion the raw dispersion reaches by a different route. This analysis was added after pre-registration and is labelled exploratory throughout. The record component is weakly identified, and its profile interval permits a materially larger record effect than the point estimate suggests.

This research concerns a construct and the artifacts that carry it. It does not concern a product, and the separation is kept deliberately sharp so a procurement reviewer never has to guess which side of the line a sentence sits on.

What this research addresses

  • Whether an operationalised documentation-risk construct is detectable by independent experts
  • Artifact-level risk: what a finalised record does and does not carry on its own terms
  • Whether independent reviewers apply the five conditions consistently with one another
  • How consistently independent AI models from different vendors apply the same conditions; this is measured as raw agreement and does not by itself establish the pre-registered reproducibility criterion
  • How much of the observed variation belongs to reviewers rather than to records

What it does not address

  • Whether any tool prevents Decision Reconstruction Risk from arising
  • Whether automated review substitutes for a qualified human reviewer
  • Whether flagged records produce different real-world outcomes when contested
  • Compliance status under any external framework, including the EU AI Act, the NIST AI RMF and ISO/IEC 42001
  • Any measure of return on investment, cost avoidance or litigation risk reduction
JRS does not establish compliance with the EU AI Act, the NIST AI RMF, ISO/IEC 42001 or any other framework, and no page on this site claims otherwise. JRS measures the record, not the tool that produced it, and not the organisation that holds it.

The research above supports one operational conclusion and licenses one design response. The conclusion: a record can fail to explain its own decision in ways a trained reader can identify, reliably enough to be worth looking for, and unevenly enough that a single reader is not a control. The design response follows directly from the dispersion, which is why every asset below assumes calibration and second review rather than a single pass.

Practitioner · JRS Investigator Guides
Free · 3 editions
The five conditions written as investigative questions, in Employment and EEO, Fair Housing, and International editions. This is the construct from the study, expressed as something a practitioner can apply to a closed file the same afternoon. A supplement to existing procedure, not legal advice. Open the guides →
Self-assessment · The Seven-Point Record Defensibility Check
No registration
Seven named ways an AI-assisted record fails when it is challenged, each with the question that detects it. Runs entirely in the browser: nothing is uploaded and nothing is sent anywhere. Check a closed file →
Instrument · JRS Codebook, version 1.0
Versioned
The measurement instrument itself: the five conditions, their decision rules, and the routing bands. Every figure on this page was produced by human reviewers and independent models applying this document, unchanged, to the same records. Read the codebook →
Calibration · Reviewer Training, six modules
Free · open
The dispersion finding is the argument for calibration, and this is the calibration. Six modules with a role-gated path for HR, compliance, investigations, employee relations and administration. No code and no registration to begin, and a certificate at the end for anyone who wants one. Start Module 1 →
Organisation · Structured Pilot
Scoped engagement
A bounded pilot on an organisation's own closed records, run under the same conditions and the same claim control as the research. Sampled double review and adjudication of disagreements are part of the design rather than an upgrade, for the reason the dispersion figure gives. Scope a pilot →
Platform · Enterprise and Integration
Licensing
Licensing of the standard and the review engine for organisations embedding record-level review into an existing workflow. The engine holds no record text at rest, which narrows what a vendor security review has to examine. Integration pathways →
Participation and independence. Every reviewer took part unpaid and in a personal capacity. Panel figures on this page are read from the study database at page load rather than transcribed, so a sentence here cannot drift from the data behind it. A figure that could not be read live is marked as a last known value rather than shown as current.
Claim control. Cross-vendor consistency, reproducibility, reliability, accuracy and validation are distinct properties and are never treated as equivalent in this programme. Raw cross-vendor agreement is reported; the pre-registered reproducibility criterion is not established because the required chance-corrected AC1 was not computed for that study. Reliability and accuracy are reported only where separately measured. Real-world validation has not been established.