Context
Compliance work built on dense regulation, strict documentation requirements, and real consequences for error — the kind of environment where a plausible-sounding wrong answer is worse than no answer.
The problem
Files arrive as scanned, inconsistent document sets that must be classified, extracted, and checked against program rules.
Some determinations are pure rule evaluation; others require judgment over incomplete information. The system has to know which is which.
Every determination needs an explanation an auditor can follow.
The system
Critical decisions
Human-in-the-loop by consequence, not by AI category
Review effort concentrates where an error costs the most, not uniformly across every AI-touched step.
Deterministic rules stay deterministic
Income limits and eligibility math are computed, never generated. AI extracts, classifies, and drafts — rules decide.
Evidence-linked outputs
Every extracted value points back to its source page, so verification is a glance rather than a re-read.
Execution
Workflow-first design: the system mirrors how compliance teams actually process a file, with AI stages embedded at the steps where they remove real hours. Evaluation sets built from historical files before automation was trusted.
Outcome
- Compliance processing where AI does the reading and people do the deciding.
- Audit-ready determinations with document-level evidence.
- [Quantified outcomes to confirm before publication]
Lessons
- Regulated AI is a systems and accountability problem before it is a model problem.
- The boundary between computed and generated must be explicit in the architecture.
- Trust is built with evaluation sets, not demos.
Have a similar problem?
Discuss an AI initiative