Evidence Development Program
Research & Validation
JRS is in a staged validation program that tests whether structured pre-finalization review can identify Decision Reconstruction Risk, the condition in which a record cannot explain why a consequential decision was made. Three stages have now produced data. The detection panel, 16 independent reviewers across 11 countries and 5 continents, detected unreconstructable records at 83.9% accuracy against a verified key, clearing the threshold set before any data were examined; these figures carry the provisional and methodological limitations recorded against each study; the operational validation study closed on 4 September 2026 and analysis continues. Independent AI models apply the standard with measurable consistency (66.7 to 93.3 percent across 61 recorded runs, mean 85.3 percent, 12 June to 21 August 2026; study closed 21 August 2026) (range 82.2 to 93.3 percent across 37 runs at the full 15-record set), and a separate human-reviewer reliability analysis measured Gwet's AC1 at 0.739 for invited participants and 0.623 for open enrolment on 10 analysed records; both point estimates exceeded 0.61, but the pre-registered two-part reliability criterion was not met because neither lower confidence bound reached 0.41. Detection, reproducibility, and reliability are distinct from real-world validation, which is a later stage, and nothing here is presented as validated until pre-registered thresholds are met. This page registers the program's studies and reports their honest current status.
Empirical Validation: Headline Metrics
Expert Detection Accuracy
83.9%
Against a verified answer key, read blind. 16 independent experts on the detection panel, 11 countries, 5 continents, 384 determinations. 95% CI 72.7 to 95.1 at participant level. Sensitivity 87.0%, specificity 80.7%. Clears the pre-registered threshold of 70% fixed before any data were seen.
Inter-Rater Reliability
0.739 / 0.623 criterion not met
Gwet's AC1, invited and open-enrolment participants, on a shared 10-record analysed set carrying 113 submitted determinations reduced to 104 after keeping one label per rater per record. Point estimates of 0.739 and 0.623 both exceed the pre-registered 0.61 point floor, but neither confidence-interval lower bound reaches the required 0.41. The pre-registered two-part reliability criterion was therefore not met. No verbal adequacy band is attached, and both results are reported as interim.
Cross-Vendor Model Consistency
66.7–93.3%
Mean pairwise agreement across three vendors, Anthropic, OpenAI and Google, applying the same five conditions to the same constructed record set. Across 61 recorded runs between 12 June and 21 August 2026 agreement ranged from 66.7 to 93.3 percent, mean 85.3 percent. Restricted to the runs that returned every record, the range is 82.2 to 93.3 percent across 37 runs at the full 15-record set. The range is shown rather than a single run because a single figure from a closed 61-run study conceals the dispersion, which is the finding. The study closed on 21 August 2026 and is no longer recomputed.
Reviewer Base
58 reviewers
Have graded at least one record across the three studies, unpaid and in a personal capacity. 36 of them (all completers, both arms) completed a full 24-record set, spanning 16 countries and 5 continents. Counted from the study database at 11 August 2026, not from a mailing list.
Detection, cross-vendor consistency, reliability, and real-world validation are distinct. The cross-vendor study reports raw agreement only; because its pre-registered reproducibility criterion requires chance-corrected AC1 and no AC1 was computed for that study, the reproducibility criterion is not established. Nothing above is presented as real-world validation.
The Five Review Conditions
The standard is these five conditions and the rules for applying them. Every figure above was produced by human reviewers and independent AI models applying the same published condition set to the same records. The conditions are versioned and documented in the Codebook. The Review Engine operationalizes them through a documented technical mapping. That mapping remains under methodological review and must not be described as exact or identical until the correspondence record supports that conclusion.
Condition 1 · Reconstructability
Can the conclusion be reconstructed from the record? A future reviewer must be able to trace the path from documented evidence to the conclusion reached, without relying on the author's recollection or added explanation.
Condition 2 · Basis Identification
Is the basis for the conclusion identifiable? The source of each characterization must be visible and traceable, not implied or summarized without attribution.
Condition 3 · Chronology
Is the chronology understandable? The sequence of events must be followable from the record alone, including the timing of prior interventions, escalation steps, and the period under review.
Condition 4 · Decision-Process Traceability
Can a future reviewer determine how the conclusion was reached? Who reviewed the matter, what criteria or threshold triggered the conclusion, and whether responsive or mitigating information was considered before finalization.
Condition 5 · Evidentiary Sufficiency
Could a reviewer with no prior knowledge evaluate the evidentiary sufficiency of this record? The record must stand on its own so an independent reviewer can assess whether the conclusion is supported by the documented evidence.
Each condition returns Pass, Needs Attention or Fail with a written note, and the five combine into a routing band. Section 5.4 of the detection study reports that all five carry information separating a reconstructable record from an unreconstructable one, at p between 1.0e-08 and 1.5e-11 on 108 labels from 21 raters.
This Week's Observation
Scope: the latest observation appears below. Observations to date include a nightly cross-vendor consistency study (independent AI models judging the same constructed records) and a separate human inter-rater reliability analysis among independent reviewers. The cross-vendor series is raw agreement on constructed records; it is not accuracy, reliability, or established reproducibility. Its pre-registered reproducibility criterion requires chance-corrected AC1, which was not computed for that study. Neither study is real-world validation or evidence about real workplace records.
No observation published yet. One appears here automatically once a study run completes.
Questions Emerging From the Data
We publish questions, not conclusions. Each is open and under investigation. Follow a question rather than a report.
Validation Maturity
Current Stage
Operational Validation
Early results in (cross-vendor consistency and a separate reliability analysis); the cross-vendor reproducibility criterion is not established and no claim is real-world validated.
Evidence Base
Detection Reported
Programme-wide, 58 international reviewers have graded records across the three studies. Panel detection 83.9% against a verified key (16 independent reviewers on the detection panel, 11 countries); cross-vendor raw agreement ranged 66.7 to 93.3 percent, mean 85.3 percent, across 61 recorded mixed-denominator runs, and 82.2 to 93.3 percent across 37 runs restricted to the full 15-record set, study closed 21 August 2026; no chance-corrected AC1 was computed for that cross-vendor study, so its pre-registered reproducibility criterion is not established; reviewer reliability Gwet's AC1 0.739 invited / 0.623 open enrolment, interim; the pre-registered two-part reliability criterion was not met because neither lower confidence bound reached 0.41. Not real-world validation.
Condition Maturity
Experimental
Readiness Scores
Not Established
Pending accuracy-benchmark data and larger samples.
Current Findings
Study Registry
Study 001 · AI Cross-Vendor Consistency (synthetic)
Closed 21 Aug 2026Three independent AI models from different vendors each judge the same set of constructed (synthetic) records, and the nightly run reported how often the models agreed. The study closed on 21 August 2026 and is no longer recomputed. It involves no human reviewers and no ground-truth labels, so it cannot speak to accuracy or to real records. Independent vendors agreeing is a stronger signal than one model repeating itself, but agreement is still not accuracy and not validation.
Study 003 · Condition Performance
CollectingPer-condition data now collected: 108 scored determinations across 10 records, distributed Gap 69%, Needs work 18%, Ready 13%. Per-condition agreement analysis continues as volume grows.
Study 004 · Reviewer Reliability
Result reportedIndependent reviewers have scored a shared record set; inter-rater reliability is reported as chance-corrected agreement (Gwet's AC1 0.739 invited, 0.623 open enrolment, 10 analysed records). Both point estimates exceed the pre-registered 0.61 point floor, but the pre-registered two-part criterion was not met because neither confidence-interval lower bound reached 0.41. These figures are interim: they rest on 10 records against a pooled target of about 26 and the intervals are wide. Reliability, not accuracy. See the Reliability & Accuracy methods paper below.
Study 008 · Professional Reviewer
CollectingRole and profession are captured at participation; patterns by reviewer type are reported as numbers grow.
Study 009 · Organizational-Psychology Readiness
Dataset readyPer-condition reliability and agreement data (10 records, multiple raters, all five conditions scored) have been assembled and exported for independent organizational-psychology review of construct validity. Awaiting an organizational-psychology reviewer.
Study 010 · Criterion Validity (real-outcome)
CollectingDe-identified public determinations are paired with their documented real-world outcomes (upheld, overturned, challenged) to test whether JRS reads correspond to results when a record is contested. Cases are accruing across HR, public-records, and related domains in small batches. No results are reported until the sample is adequate.
Study 012 · Randomized Comparison: structured review versus unaided judgment
CollectingA separate study, with its own participants and its own recruitment, testing whether applying the five conditions improves detection relative to unaided professional judgment on the same 24-record corpus. Independent experts with no prior exposure to the method are randomly assigned by a deterministic hash of their participant code, before they judge any record, either to the five conditions or to a single general question about adequacy of support. Participants are blind to the two-condition design and are debriefed when the study closes. Data collection is still open and no result is reported. It will be reported separately and in full when it closes, whatever the outcome. This study asks a different question from Study 011 and neither result substitutes for the other.
Study 011 · AI-Assisted Records Detection
Result reportedWhether JRS distinguishes AI-generated records whose conclusions are grounded in their source from records that read convincingly but are not, judged blind by independent reviewers against a held-out key. Constructed stimuli with known ground truth. The detection panel, 16 reviewers across 11 countries and 5 continents, has completed the full 24-record set, producing 384 graded reads. Panel detection accuracy against the verified key is 83.9% (95% confidence interval 72.7 to 95.1 at participant level; sensitivity 87.0%, specificity 80.7%), clearing the pre-registered threshold on both criteria. Detection of a known AI documentation risk. Reviewers took part in a personal capacity, unpaid. Study status: the operational validation study closed on 4 September 2026. Figures are current as of 5 August 2026 and carry the methodological and provisional limitations stated above. Analysis and reporting continue: a manuscript reporting this result in full is in preparation.
Status legend. Active, Collecting, Result reported, Dataset ready = real data accruing or reported from live participation. This is a validation-phase program; no claim is presented as validated.
What Would Count as Evidence (and What Would Falsify a Claim)
A JRS claim of usefulness would be supported only if independent reviewers, applying the five conditions to records they did not author, identify deficiencies that standard review misses, and agree with one another above chance. It would be weakened or falsified if reviewers cannot apply the conditions consistently, if flagged records are no less reviewable than unflagged ones under expert assessment, or if agreement is no better than chance. Two early stages have now produced data: raw cross-vendor model agreement on synthetic records, and a separate human inter-rater reliability analysis among invited and open-enrolment participants (see the Reliability and Accuracy methods paper below). The cross-vendor study does not establish the pre-registered reproducibility criterion because the required chance-corrected AC1 was not computed. These results are not real-world validation. Accuracy against a held-out key is reported separately where measured.
Origin & Approach
JRS originated from civil rights investigative and documentation-review experience, with a cognitive-behavioral and AI-governance lens. Read the origin and what JRS is and is not →
Perspectives
Concept · Decision Reconstruction Risk (DRR)
Definition
The condition in which a record cannot explain, on its own terms, why a consequential decision was made. The named failure mode behind indefensible AI-assisted records. Read →
Perspective · Why Good Decisions Fail on Paper
Essay
A practitioner essay on Decision Reconstruction Risk: how sound decisions leave indefensible records, and why AI accelerates it. By Phillip Wikes. Read →
Working Paper · The Justification Review Standard (JRS)
PDF
Full working research paper: origins, the five conditions, the proportionality principle, the evidence-development program and pilots to date, and the enterprise plan. Validation phase; preliminary and observational. Download →
Methods Paper · Reliability and Accuracy of JRS (Rungs 1 & 2)
PDF
Cross-vendor AI consistency (66.7 to 93.3 percent across 61 recorded mixed-denominator runs, mean 85.3 percent; 82.2 to 93.3 percent across 37 runs restricted to the full 15-record set; closed 21 August 2026; no chance-corrected AC1 computed, so the pre-registered reproducibility criterion is not established) and human inter-rater reliability (Gwet's AC1 0.739 invited, 0.623 open enrolment, 10 analysed records; the pre-registered two-part criterion was not met), with accuracy reported as preliminary. Pre-registered thresholds; preliminary and observational. Download →
Perspective · When the Record Sounds Right but Says Nothing
AI Governance
AI-assisted documentation and the next layer of AI governance: why governing the model is no longer enough, and what comes after. By Phillip Wikes and Jake McDonough. Read →
Program Layers
Claim control: cross-vendor consistency, reproducibility, reliability, accuracy, and validation are distinct. The cross-vendor study reports raw agreement; its pre-registered reproducibility criterion is not established because the required chance-corrected AC1 was not computed. Figures shown do not establish real-world validation.