Evidence Development Program
Research & Validation
JRS is in a staged validation program that tests whether structured pre-finalization review can identify Decision Reconstruction Risk, the condition in which a record cannot explain why a consequential decision was made. Three stages have now produced data. As of 5 August 2026, an international panel of 16 independent experts across 11 countries and 5 continents detected unreconstructable records at 83.9% accuracy against a verified key, clearing the threshold set before any data were examined; these figures are provisional until the study closes and are recomputed at that point. Independent AI models apply the standard reproducibly (86.7% cross-vendor agreement on the latest nightly run, range 82.2 to 93.3 percent across 37 runs at the full 15-record set), and independent human reviewers apply it reliably (Gwet's AC1 0.739 experts, 0.624 trained reviewers, interim on 10 records). Detection, reproducibility, and reliability are distinct from real-world validation, which is a later stage, and nothing here is presented as validated until pre-registered thresholds are met. This page registers the program's studies and reports their honest current status.
Empirical Validation: Headline Metrics
Expert Detection Accuracy
83.9%
Against a verified answer key, read blind. 16 independent experts, 11 countries, 5 continents, 384 determinations. 95% CI 72.7 to 95.1 at participant level. Sensitivity 87.0%, specificity 80.7%. Clears the pre-registered threshold of 70% fixed before any data were seen.
Inter-Rater Reliability
0.739 / 0.624
Gwet's AC1, experts and trained reviewers, on a shared 10-record set carrying 99 retained labels. Both clear the pre-registered floor of 0.61 and sit in the substantial band under Landis and Koch. Intervals are wide and the plan's secondary lower-bound criterion is not met, so both are reported as interim.
Cross-Vendor Model Consistency
86.7%
Mean pairwise agreement across three vendors, Anthropic, OpenAI and Google, applying the same five conditions to the same 15 records. Latest nightly run 9 August 2026; range 82.2 to 93.3 percent across 37 runs at the full record set. Recomputed every night, not transcribed.
Reviewer Base
57 reviewers
Have graded at least one record across the three studies, unpaid and in a personal capacity. 36 of them completed a full 24-record set, spanning 16 countries and 5 continents. Counted from the study database on 11 August 2026, not from a mailing list.
Detection, reproducibility and reliability are distinct from real-world validation, which is a later stage. Nothing above is presented as validated until its pre-registered threshold is met, and the two figures that have not met theirs say so.
The Five Review Conditions
The standard is these five conditions and the rules for applying them. Every figure above was produced by human reviewers and independent AI models applying this same set, unchanged, to the same records. The conditions are versioned, documented in the Codebook, and implemented identically in the review engine, which is what makes a result from one reviewer comparable to a result from another.
Condition 1 · Reconstructability
Can the conclusion be reconstructed from the record? A future reviewer must be able to trace the path from documented evidence to the conclusion reached, without relying on the author's recollection or added explanation.
Condition 2 · Basis Identification
Is the basis for the conclusion identifiable? The source of each characterization must be visible and traceable, not implied or summarized without attribution.
Condition 3 · Chronology
Is the chronology understandable? The sequence of events must be followable from the record alone, including the timing of prior interventions, escalation steps, and the period under review.
Condition 4 · Decision-Process Traceability
Can a future reviewer determine how the conclusion was reached? Who reviewed the matter, what criteria or threshold triggered the conclusion, and whether responsive or mitigating information was considered before finalization.
Condition 5 · Evidentiary Sufficiency
Could a reviewer with no prior knowledge evaluate the evidentiary sufficiency of this record? The record must stand on its own so an independent reviewer can assess whether the conclusion is supported by the documented evidence.
Each condition returns Pass, Needs Attention or Fail with a written note, and the five combine into a routing band. Section 5.4 of the detection study reports that all five carry information separating a reconstructable record from an unreconstructable one, at p between 1.0e-08 and 1.5e-11 on 108 labels from 21 raters.
This Week's Observation
Scope: the latest observation appears below. Observations to date include a nightly cross-vendor reproducibility check (independent AI models judging the same constructed records) and human inter-rater reliability among independent reviewers. Reproducibility and reliability mean the standard is applied consistently; they are not accuracy, not validation, and not evidence about real workplace records.
No observation published yet. One appears here automatically once a study run completes.
Questions Emerging From the Data
We publish questions, not conclusions. Each is open and under investigation. Follow a question rather than a report.
Validation Maturity
Current Stage
Operational Validation
Early results in (reproducibility and reliability); no claim validated yet.
Evidence Base
Detection Reported
Programme-wide, 57 international reviewers have graded records across the three studies. Panel detection 83.9% against a verified key (16 independent experts, 11 countries); reproducibility 86.7% on the latest nightly cross-vendor run; reviewer reliability Gwet's AC1 0.74 experts / 0.62 trained, interim. Not real-world validation.
Condition Maturity
Experimental
Readiness Scores
Not Established
Pending accuracy-benchmark data and larger samples.
Current Findings
Study Registry
Study 001 · AI Reproducibility (cross-vendor, synthetic)
ActiveThree independent AI models from different vendors each judge the same set of constructed (synthetic) records, and the nightly run reports how often the models agree. It involves no human reviewers and no ground-truth labels, so it cannot speak to accuracy or to real records. Independent vendors agreeing is a stronger signal than one model repeating itself, but agreement is still not accuracy and not validation.
Study 003 · Condition Performance
CollectingPer-condition data now collected: 108 scored determinations across 10 records, distributed Gap 69%, Needs work 18%, Ready 13%. Per-condition agreement analysis continues as volume grows.
Study 004 · Reviewer Reliability
Result reportedIndependent reviewers have scored a shared record set; inter-rater reliability is substantial (Gwet's AC1 0.739 experts, 0.624 trained reviewers, 10 records), clearing the pre-registered point floor of 0.61. These figures are interim: they rest on 10 records against a pooled target of about 26, the confidence intervals are wide, and the plan's secondary lower-bound criterion is not yet met. Reliability, not accuracy. See the Reliability & Accuracy methods paper below.
Study 008 · Professional Reviewer
CollectingRole and profession are captured at participation; patterns by reviewer type are reported as numbers grow.
Study 009 · Organizational-Psychology Readiness
Dataset readyPer-condition reliability and agreement data (10 records, multiple raters, all five conditions scored) have been assembled and exported for independent organizational-psychology review of construct validity. Awaiting an organizational-psychology reviewer.
Study 010 · Criterion Validity (real-outcome)
CollectingDe-identified public determinations are paired with their documented real-world outcomes (upheld, overturned, challenged) to test whether JRS reads correspond to results when a record is contested. Cases are accruing across HR, public-records, and related domains in small batches. No results are reported until the sample is adequate.
Study 012 · Randomized Comparison: structured review versus unaided judgment
CollectingA separate study, with its own participants and its own recruitment, testing whether applying the five conditions improves detection relative to unaided professional judgment on the same 24-record corpus. Independent experts with no prior exposure to the method are randomly assigned by a deterministic hash of their participant code, before they judge any record, either to the five conditions or to a single general question about adequacy of support. Participants are blind to the two-condition design and are debriefed when the study closes. Data collection is still open and no result is reported. It will be reported separately and in full when it closes, whatever the outcome. This study asks a different question from Study 011 and neither result substitutes for the other.
Study 011 · AI-Assisted Records Detection
Result reportedWhether JRS distinguishes AI-generated records whose conclusions are grounded in their source from records that read convincingly but are not, judged blind by independent reviewers against a held-out key. Constructed stimuli with known ground truth. Sixteen reviewers across 11 countries and 5 continents have completed the full 24-record set, producing 384 graded reads. Panel detection accuracy against the verified key is 83.9% (95% confidence interval 72.7 to 95.1 at participant level; sensitivity 87.0%, specificity 80.7%), clearing the pre-registered threshold on both criteria. Detection of a known AI documentation risk. Reviewers took part in a personal capacity, unpaid. Figures are current as of 5 August 2026 and remain provisional until the study closes, expected 14 August 2026. Invited reviewers may still complete, which would change the counts and the point estimate. Every figure is recomputed at close. A manuscript reporting this result in full is in preparation.
Status legend. Active, Collecting, Result reported, Dataset ready = real data accruing or reported from live participation. This is a validation-phase program; no claim is presented as validated.
What Would Count as Evidence (and What Would Falsify a Claim)
A JRS claim of usefulness would be supported only if independent reviewers, applying the five conditions to records they did not author, identify deficiencies that standard review misses, and agree with one another above chance. It would be weakened or falsified if reviewers cannot apply the conditions consistently, if flagged records are no less reviewable than unflagged ones under expert assessment, or if agreement is no better than chance. Two early stages have now produced data: cross-vendor model reproducibility on synthetic records, and human inter-rater reliability among independent expert and trained reviewers at substantial, chance-corrected agreement (see the Reliability and Accuracy methods paper below). These support reproducible and reliable application; they are not accuracy and not validation. Accuracy against a held-out key is in progress. Nothing here is presented as validated until pre-registered thresholds are met.
Origin & Approach
JRS originated from civil rights investigative and documentation-review experience, with a cognitive-behavioral and AI-governance lens. Read the origin and what JRS is and is not →
Perspectives
Concept · Decision Reconstruction Risk (DRR)
Definition
The condition in which a record cannot explain, on its own terms, why a consequential decision was made. The named failure mode behind indefensible AI-assisted records. Read →
Perspective · Why Good Decisions Fail on Paper
Essay
A practitioner essay on Decision Reconstruction Risk: how sound decisions leave indefensible records, and why AI accelerates it. By Phillip Wikes. Read →
Working Paper · The Justification Review Standard (JRS)
PDF
Full working research paper: origins, the five conditions, the proportionality principle, the evidence-development program and pilots to date, and the enterprise plan. Validation phase; preliminary and observational. Download →
Methods Paper · Reliability and Accuracy of JRS (Rungs 1 & 2)
PDF
Cross-vendor AI reproducibility (86.7% latest nightly, 15 records) and human inter-rater reliability (Gwet's AC1 0.739 experts, 0.624 trained reviewers, 10 records), with accuracy reported as preliminary. Pre-registered thresholds; preliminary and observational. Download →
Perspective · When the Record Sounds Right but Says Nothing
AI Governance
AI-assisted documentation and the next layer of AI governance: why governing the model is no longer enough, and what comes after. By Phillip Wikes and Jake McDonough. Read →
Program Layers
Claim control: reproducibility, accuracy, and validation are distinct and are not treated as equivalent anywhere in this program. Figures shown are observational, not statistical findings or validated data.