FREE TO READ

Everything here is free to read, and nothing here is a claim about your case. Built from California’s published records and modified by us — not official government data. How these numbers were made.

Methodology · data as of 2026-09-04

How these numbers were made

This page exists so that any figure on this site can be traced to the record behind it, and so that the things that did not work are as visible as the things that did.

42,749
rows in the published file
42,710
unique after de-duplication
42,621
analysable — self-contradicting excluded
250
hand-read and coded by a person

The source

verbatim

California Department of Managed Health Care, Independent Medical Review (IMR) Determinations, Trend, published on the CalHHS Open Data Portal. The file carries every IMR decision the Department has administered since 1 January 2001. Data as of 2026-09-04: 42,749 rows, of which 42,710 are unique.

Publication is not discretionary. Health & Safety Code § 1374.33(g) requires the Department to make the de-identified decisions available in a searchable database, and § 1374.33(h)(1) prescribes the fields — which map one-to-one onto the columns of this file, including:

“a detailed case summary that includes the specific standards, criteria, and medical and scientific evidence, if any, that led to the case decision”
Health & Safety Code § 1374.33(h)(1)(K)

The reasoning this site describes is content the Legislature required to be published.

Licence position. The dataset page displays a notice that commercial uses must be approved. A written commercial-use request to the Department is outstanding. Until it is answered or counsel confirms the Public Records Act route, this deployment sells nothing, carries no price and is not indexed. The notice’s status is itself unsettled — identical text appears on datasets published by an unrelated department — which is why it is being resolved in writing rather than assumed either way.

The rules every figure obeys

enforced in code
  1. No rate below n = 30. A group smaller than that is published as a count, never as a percentage. A precedent list is a legitimate product; a rate computed on six decisions is not. The explorer runs into this constantly, on purpose.
  2. Every rate carries its era. Headline figures are 2021-2025. The all-years rate is 52.5% and it describes no era: the annual rate has risen from 35.1% in 2002 to 72.4% in 2025.
  3. Duplicates are dropped before counting. The published file carries 39 exact duplicate rows.
  4. Self-contradicting decisions are excluded. 89 decisions state a result in their own text that disagrees with the published determination field. They are counted in what the source contains and left out of every rate, rather than resolved in favour of one side.
  5. Partial outcomes are flagged, not flattened. 753 decisions record a partial overturn that the published binary field collapses into a full win or a full loss.
  6. No decision text is published. Not on any page, not in any data file, not as a per-decision URL. The build database holds the summaries; the deployment holds counts and codes. The build fails if a case summary reaches the output.

Accuracy testing, and what failed it

3 published · 11 held back · 2 withdrawn

Classification codes are produced by explicit pattern rules, not by a model, so any classification can be answered with the sentence that triggered it. The rules were scored against 250 decisions read in full and coded by hand — 100 of them before the harness existed, and 150 drawn disjoint and coded blind, with the sampler physically withholding the harness output. A code appears on this site only if it clears every bar: agreement 90%, recall 80%, precision 85%, and at least 8 instances to judge it on.

CodeSupportAgreementPrecisionRecallVerdict
C1299%0%heldsupport 2 < 8 — too rare to evaluate
C24286%64%43%heldagreement 86% < 90%
C316482%86%85%heldagreement 82% < 90%
C44881%0%withdrawnwithdrawn — failed in both directions at once; a numbers-and-units rule catches dosages in the plan's request and misses non-numeric findings, which are most of them
C51398%91%77%heldrecall 77% < 80% — misses real instances
C65390%85%64%heldrecall 64% < 80% — misses real instances
C71594%100%7%heldrecall 7% < 80% — misses real instances
D15585%62%87%heldagreement 85% < 90%
D22798%90%96%shipsclears the threshold
D313587%90%86%heldagreement 87% < 90%
D414100%93%100%shipsclears the threshold
D55578%0%withdrawnwithdrawn — matched 96.8% of all decisions; a code that applies to nearly every decision carries no information
D62697%91%81%shipsclears the threshold
D75185%71%47%heldagreement 85% < 90%
D86088%73%80%heldagreement 88% < 90%
D9499%100%25%heldsupport 4 < 8 — too rare to evaluate

What passed says more than the count. All three shipping codes are named things a reader can point at in the text — a regulatory status, a scored instrument, a statute. Every code requiring a judgement about what the reviewer meant failed, and the near misses fail on recall rather than precision: reviewers name a documentation gap in too many ways for a pattern list to catch two thirds of them. Rule-based extraction recovers named things, not judgements. More hand-reading will not move that.

Three things this validation does not establish. Our own 250-record target is met, but our 100%-agreement bar is not, and nothing came close — the threshold above is deliberately looser and is stated as such rather than presented as the standard. The hand coder and the rule author are the same party; that is mitigated by coding before writing patterns and by coding the second batch blind, but it is not an independent coder and is not presented as one. And the patterns were revised once after seeing the first run — two codes withdrawn because the scores exposed definitions that do not hold up, and two specification bugs corrected. No pattern was widened to capture a decision the harness had been observed to miss; that would be tuning to the reference set, and it would make this table meaningless.

What ships without needing a gate

facts, not judgements

Some fields are read off the record rather than judged. Naming the authority a reviewer applied, counting the reviewers on a panel, or recording that a panel split two to one is closer to reading a label than to interpreting a case, so these carry the site whether or not an interpretive code clears the bar.

Decisions naming at least one authority11,96028.0%
Decisions naming the reviewer's board certification25,64460.0%
Experimental/investigational decisions with a three-reviewer panel7,08116.6%
Decisions publishing a 2–1 split among reviewers2,0024.7%

The published split is the most unexpected of these. Experimental/investigational reviews are decided by three physicians, and where they disagree the decision carries each reviewer’s separate opinion — a published dissent, in 2,002 decisions. Nothing in this project’s own specification anticipated it; it turned up in the hand-read.

The category vocabulary changed in 2026

93% crosswalked

DMHC switched both the diagnosis and treatment coding schemes at report year 2026: 513 decisions use ICD-chapter and procedure-code groupings that appear in no earlier year, and several older classes — Morbid Obesity, Chronic Pain Syndrome, Autism Spectrum — have no successor category at all. No public documentation of the change exists, so a request for it is outstanding.

Diagnosis crosswalks; treatment does not. The 2026 diagnosis scheme is ICD-10 chapters and the legacy one is clinical service lines, which correspond well enough to map with a stated confidence on every row — 476 of 513 2026 decisions carry a crosswalked condition class, and 37 deliberately do not. The 2026 TREATMENT scheme is CPT/HCPCS sections — “Surgery”, “Drug Ad Oth Oral”, “Temp Codes” — against legacy service lines like Pharmacy and DME. A single 2026 “Surgery” bucket spans legacy general surgery, orthopedic, neurosurgery, OB-GYN and cardiovascular procedures, and nothing in the record says which. So there is no treatment crosswalk: inventing one would be the fabrication this dataset exists to avoid.

The mapping is ours, and is labelled as ours. Every row carries a confidence — clean, lossy, or split — and a note saying what the merge cost. A category the crosswalk has never seen makes the build fail rather than landing in an “other” bucket, because that is how a dataset fractures with nobody noticing. Trend lines still stop at 2025: the crosswalk fixes the class, not the fact that 2026 is a part-year.

How the text got richer

the record
EraDecisionsMedian lengthName the reviewer's board
pre-201617,3551,010 chars6.8%
2016-202524,8171,962 chars97%
2026-5381,662 chars99.8%

The jump at 2016 is SB 1410’s field mandate becoming operative. Decisions before it are materially thinner, and 374 across the dataset are too brief to carry any reasoning at all — almost all of them from 2006–2010.

Corrections

open

If a figure here is wrong, or a classification misreads a decision, we would rather know. The corrections process is public and dated, and disputes raised by a state agency get priority review against the source records.

Every rate here counts only denials appealed all the way to independent review. Fewer than 1% of denials are appealed internally at all, so a share on this page is not the share of denials that are wrong — it is the share of contested ones that reviewers did not sustain. California only, and only plans the state regulates: not self-funded employer plans, Medicare or Medicaid.