Measuring Point-in-Time Correctness & Response Divergence using legal-rag-audit tool
Memon Systems
0:00 / 0:00
Measuring Point-in-Time Correctness & Response Divergence using legal-rag-audit tool
13 просмотров · 9 дней назад
Memon Systems
3 подписчика
13 просмотров · 9 дней назад
An operational demonstration of legal-rag-audit in existing-corpus mode: sealing a statutory battery before execution, validating JSONPaths pre-flight, running 3-pass multi-query generation, and scoring point-in-time correctness offline with zero remote sockets.
Tested against 3 live commercial UK legal AI products across a fixed 14-probe battery (12 point-in-time anchors across 6 statutory provisions + 2 licensed-content checks):
• 3 of 3 targets exhibited material response divergence on identical repeated queries.
• 1 of 3 answered a dated question using superseded statutory text while correctly citing the section.
• 2 of 3 answered every dated question correctly on a single pass — repetition removed one, which gave the following month's statutory cap on passes 2 and 3.
---
TIMESTAMPS:
00:00 - The Problem: Why Single-Pass Evaluations Measure Nothing
00:45 - Architecture & Invariants: Ingest, Plant, Hash, and Pre-Commitment
02:24 - Configuration & Pre-Flight Validation (Preventing False Positives)
03:29 - Multi-Pass Generation & Air-Gapped Deterministic Scoring
05:34 - Live Run: Ingesting UK Statutes & Anchors (legislation.gov.uk)
06:02 - Live Run: Generating Existing-Corpus Battery & Answer Key
07:10 - Live Run: Sealing Ground Truth with SHA-256 Hashes
07:47 - Live Run: Pre-Flight Validation & 3-Pass Response Generation
08:50 - Live Run: Offline Scoring, Tier 2 Bypass & Evidence Reports
10:55 - Empirical Findings: Testing 3 Commercial UK Legal AI Systems
11:23 - Target 1: Substantive Value Drift across Passes (£68,400 vs £72,300)
12:15 - Target 2: Agentic Routing Dropout & Empty Stream Thrashing
12:56 - Target 3: Temporal Citation Illusion (Section 108 Modern vs 2011 Law)
13:50 - Public Harness, GitHub Repo & Extending Domain Anchors
14:07 - TPRM Audits: Why Re-runnability Beats SOC 2
14:41 - The Self-Attestation Problem & Independent Deployment Audit Close
---
RESOURCES & REPRODUCIBILITY:
• Open Source Harness: https://github.com/memonsystems/legal...
• Diagnostic Engagement (£500 fixed / 10-day turnaround): https://memonsystems.com/engagements
• Companion Research & Transcripts: https://memonsystems.com/tools
METHODOLOGICAL DISCLOSURE:
Scores are evaluated deterministically (Tier 1 exact match) against pre-committed SHA-256 digests authored prior to query dispatch. Transport dropouts are classified as uncaptured records rather than pipeline failures. Findings characterise reproducibility and temporal grounding across a fixed 14-probe battery, not an overall system accuracy rate. All targets received private raw transcripts prior to publication.