Methodology Archive
How the JRS Evidence Development Program collects, governs, and reports evidence on whether structured pre-finalization review can identify Decision Reconstruction Risk and support decision defensibility. This work is preliminary and observational. This archive is versioned: methodology changes are recorded in the Evidence Ledger.
Measurement instrument
All review and analysis reference the JRS Codebook, which defines the five review conditions, detection criteria, examples, risks, and severity guidance. No review occurs outside the codebook's governed definitions. Conditions carry a maturity level (Experimental → Emerging → Stable → Validated); all are currently Experimental.
Data sources
- AI-Assisted Records Detection: 24 constructed records; reviewer reads against a held-out, independently verified key.
- Bench reliability: a shared record set scored on the five conditions by independent experts and trained reviewers.
- Real-case criterion: public determinations paired with their documented outcomes.
- Pilot contact: practitioner inquiries (not a study input; operational).
Records used in exercises are constructed and fictional. Participation is voluntary and may be anonymous.
What is measured
- Agreement: whether reviewers converge on the same assessment of a record.
- Accuracy: agreement with an expert benchmark mapping (separate from reviewer agreement; requires benchmark datasets not yet established).
- Condition performance: per-condition agreement and dispute patterns.
Cross-vendor consistency, reproducibility, reliability, accuracy, and validation are distinct and are never used interchangeably. Raw cross-vendor agreement does not by itself establish reproducibility; for the cross-vendor study, the pre-registered reproducibility criterion requires chance-corrected AC1 and that coefficient was not computed. Reliability is reported only for the separate human-reviewer sample. Accuracy requires correctness against a benchmark, and none of these properties alone constitutes real-world validation of the framework.
- Records are constructed, not sampled from real organizational populations.
- Self-selected participants; not a representative sample.
- Early sample sizes; public aggregates are withheld until a minimum threshold.
- Author-applied reviews are illustrative, not independent validation.