Guide
AI utilization-review compliance: the human-in-the-loop mandate
What "human-in-the-loop" means in utilization review
Utilization review (UR) and utilization management (UM) are how payers decide whether a requested service is covered and medically necessary. AI is attractive here because it can read charts and flag routine cases fast. The risk is that a model's suggested denial becomes the actual denial with no meaningful human judgment in between. "Human-in-the-loop" means the reverse: AI assists and prioritizes, but a qualified person makes and stands behind the medical-necessity call — and there is a record proving it.
Why the pressure is building
- California SB 1120 (Physicians Make Decisions Act). Effective January 1, 2025, it requires that when a California health plan or disability insurer uses AI in utilization review, a licensed physician or qualified health professional — not the tool — makes the medical-necessity determination, and the tool is audited and disclosed to regulators. See our SB 1120 explainer.
- Federal prior-authorization scrutiny. CMS's prior-authorization and interoperability rule (CMS-0057-F) tightens timelines and transparency for affected payers, and CMS guidance has stressed that coverage decisions must rest on an individual's circumstances rather than an algorithm applied in isolation. The direction of travel is more accountability for automated denials, not less.
- Litigation and defensibility. High-profile suits have alleged that insurers used algorithms to deny claims at scale with inadequate human review. Whatever their outcome, they set the expectation that a payer must be able to show a qualified human actually evaluated each denial.
None of this bans AI. All of it raises the bar on being able to demonstrate qualified human review — which is where most programs are weakest.
Why policy alone isn't enough
Plenty of payers already require physician sign-off on denials as policy. The gap is evidentiary. When a decision is questioned months later, the payer has to answer concrete questions: which named, credentialed person reviewed this claim, what did they see, and can that be shown without relying on the payer's own uncorroborated word? If the trail is scattered across systems and editable logs, "we had a human review it" is an assertion, not proof. Defensible compliance means a per-decision record that is specific, durable, and tamper-evident.
What a verified, signed decision record provides
The reusable pattern that satisfies the mandate has four properties:
- No automated denial. A proposed deny or pend can never finalize on the model's say-so; it is always routed to a human before the decision issues.
- Named, credentialed attestation. The reviewer who clears a denial is recorded with their qualification — name, credential type, license number, jurisdiction — attached to that specific decision.
- Cryptographic integrity. The decision and its attestation are sealed in an Ed25519-signed Proof Object, so the record can't be silently altered and can be verified offline by anyone holding the public key — no need to trust the vendor or the payer.
- A complete audit trail. The record shows what the AI proposed, that a qualified human reviewed it, and who that was — reproducible on demand for a regulator, an appeal, or a court.
Where HandInLoop fits
HandInLoop supplies exactly this layer, and is careful about its boundary: it verifies decisions; it does not make medical-necessity determinations. It forces every proposed denial to a human, captures the reviewer's credential as a required attestation, and returns a signed, independently verifiable record that a licensed human owned the call. The credential is currently self-declared by the reviewer rather than checked against a licensing board, but it produces a durable, tamper-evident chain of accountability — the artifact these rules effectively demand. It's demonstrated today on synthetic and de-identified claims, so teams can evaluate the mechanism before any protected data is involved.