EvidenceGate
Documents
/
DOC-006
Saved, v0.8
Export
Save version
Quality
Draft
In review
Approved
Hemnaath
Sai
Both
Ground truth, query reduction, reviewer time, extraction quality, evidence coverage, safety metrics, and rollout gates.
Markdown
H2
List
Use # headings, - lists, and plain text. Every save creates a revision.
# Pilot Evaluation and Scorecard ## Pilot question Can EvidenceGate reduce avoidable pre-authorization queries and reviewer handling time while preserving complete evidence lineage and human control? ## Design - One hospital partner and one TPA or insurer review partner where possible. - One hundred permissioned historical or shadow-mode planned pre-authorization packets. - Two procedure families. - Five document types. - One published rule bundle. - Blind comparison against the existing workflow baseline. ## Ground truth Two trained annotators label document type, critical fields, page regions, required-document state, contradictions, and expected query reasons. Disagreements go to an adjudicator. Policy outcomes come only from the approved rule bundle and partner review. ## Primary metrics - Avoidable query cycles per case. - Time from packet start to submission readiness. - Reviewer handling time to verify a packet. - Critical missing-document precision and recall. - Critical-field exact-match and normalized accuracy. - Cross-document contradiction precision. - Evidence-reference coverage for material findings. - Human correction and override rate. ## Safety metrics - False pass on mandatory deterministic checks. - Unsupported finding rate. - Citation mismatch rate. - Cases where the packet obscures an unknown or failed worker. - Policy result produced from an inactive or wrong-version bundle. ## Pilot targets Targets are hypotheses until the partner baseline is measured. - At least 30 percent fewer avoidable query cycles. - At least 40 percent lower median preparation or first-review handling time. - At least 95 percent evidence coverage. - At least 98 percent precision for critical missing-document findings. - Zero false passes caused by treating unknown as present. - Zero autonomous final decisions. ## Rollout gates ### Gate A: Offline evaluation Critical safety metrics pass on a held-out packet set. ### Gate B: Shadow mode EvidenceGate runs without changing the operational submission. Reviewers compare packets and record corrections. ### Gate C: Assisted mode Hospital users may correct packets before submission. All findings remain advisory. ### Gate D: Expansion decision Proceed only if measured query reduction, time savings, evidence coverage, and reviewer trust meet the agreed thresholds. ## Reporting Report results by document type, procedure family, insurer bundle, confidence band, and failure category. Do not publish blended averages that hide critical-field failures.