Written and maintained by CASRAI Editorial Board
Last updated
The short version
Two US states have written into insurance law the single blunt rule that almost no frontier-AI framework contains: a named class of human being, and only that human, may make one specific decision. California SB 1120 (2024) and Texas SB 815 (2025) both target the same decision — whether a requested health care service is medically necessary, and therefore whether a health plan pays for it — and both remove that decision from the authority of any automated system. Neither statute sets a capability threshold, asks for a safety case, or requires an evaluation. They simply say the machine does not get to decide.
This guide walks through what each statute actually says, which regulator enforces it, why the mechanism is insurance regulation rather than health-IT or medical-device law, and why the federal picture is emptier than the state picture. It also corrects two things that are commonly misreported about SB 1120.
California SB 1120: what the chaptered text actually requires
SB 1120 was chaptered on 28 September 2024 as Chapter 879 of the 2024 statutes. It is a short bill that amends exactly two sections, one for each of California’s two health-coverage regulators:
- Health and Safety Code § 1367.01 — the utilization review statute for health care service plans, regulated by the Department of Managed Health Care (DMHC). This covers most Californians, because most California HMO and many PPO products are DMHC-regulated plans.
- Insurance Code § 10123.135 — the parallel utilization review statute for disability (health) insurers, regulated by the California Department of Insurance (CDI).
Into each section the bill adds a new subdivision (k). It applies to a plan or insurer “that uses an artificial intelligence, algorithm, or other software tool for the purpose of utilization review or utilization management functions, based in whole or in part on medical necessity.” Note how broad the trigger is: the statute does not care whether the tool is a large language model, a gradient-boosted classifier, or a deterministic rules engine. “Algorithm, or other software tool” catches the lot. It also does not care whether the tool is bought or built.
Where the tool is in scope, subdivision (k)(1) imposes a list of conditions on it. In substance, the tool must:
- base its determination on the enrollee’s own medical or other clinical history and the individual clinical circumstances as presented by the requesting provider, and not base its determination solely on a group dataset;
- apply criteria that comply with applicable law and with clinical principles;
- not supplant health care provider decisionmaking;
- not discriminate, directly or indirectly, against enrollees in violation of state or federal law;
- be “fairly and equitably applied,” including in accordance with applicable federal Department of Health and Human Services regulations and guidance;
- be open to inspection for audit or compliance reviews;
- have its performance, use, and outcomes periodically reviewed and revised to maximise accuracy and reliability.
Then subdivision (k)(2) does the actual work, and it is worth reading twice, because it is not a procedural safeguard — it is a removal of authority:
“Notwithstanding paragraph (1), the artificial intelligence, algorithm, or other software tool shall not deny, delay, or modify health care services based, in whole or in part, on medical necessity.”
“A determination of medical necessity shall be made only by a licensed physician or a licensed health care professional competent to evaluate the specific clinical issues involved…”
The relationship between (k)(1) and (k)(2) is often misread. Paragraph (1) is not a set of conditions under which a tool may deny a claim; paragraph (2) opens with “notwithstanding paragraph (1)” precisely to foreclose that reading. Satisfying every item in the list does not buy a tool the authority to issue an adverse medical-necessity determination. The list governs how the tool may be used in the process; the prohibition governs who may make the call at the end of it.
Two things commonly misreported about SB 1120
First, the bill is very widely referred to as the “Physicians Make Decisions Act.” The chaptered text contains no short-title section. Its official title is simply “An act to amend Section 1367.01 of the Health and Safety Code, and to amend Section 10123.135 of the Insurance Code, relating to health care coverage.” The popular name comes from the sponsoring campaign, not from the statute, and it will not be found by searching the code. If you are citing this in a policy or compliance document, cite the code sections.
Second, SB 1120 carries no urgency clause, so it took effect on 1 January 2025 under California’s default effective-date rule for statutes chaptered in a regular session — not on the September 2024 chaptering date.
Texas SB 815: a flat prohibition, plus a standing audit power
Texas reached the same target by a different route. SB 815 (89th Legislature, Regular Session), authored by Senator Schwertner, is captioned “Relating to the use of certain automated systems in, and certain adverse determinations made in connection with, the health benefit claims process.” It was signed by the governor on 20 June 2025 and took effect 1 September 2025, applying to health benefit plans delivered, issued for delivery, or renewed on or after 1 January 2026 — so the operative compliance date for most plans is the 2026 plan year, not the 2025 effective date.
The bill adds Section 4201.156 to Subchapter D, Chapter 4201 of the Insurance Code — Chapter 4201 being the utilization review agent chapter, which is the right place to put it, because in Texas the regulated entity is the utilization review agent, whoever that happens to be. It defines the technology twice over:
- Automated decision system: “an algorithm, including an algorithm incorporating an artificial intelligence system, that uses data-based analytics to make, suggest, or recommend certain determinations, decisions, judgments, or conclusions.”
- Artificial intelligence system: “any machine learning-based system that, for any explicit or implicit objective, infers from the inputs the system receives how to generate outputs, including content, decisions, predictions, and recommendations.”
The operative rule is a single sentence: “A utilization review agent may not use an automated decision system to make, wholly or partly, an adverse determination.” The words “wholly or partly” matter more than they look. They rule out the obvious workaround, in which a model produces the denial and a reviewer rubber-stamps it, because the system has still partly made the determination. Whether that holds up in practice depends entirely on enforcement, and the statute supplies the enforcement hook in the next provision: “The commissioner may audit and inspect at any time a utilization review agent’s use of an automated decision system for utilization review.”
That standing, suspicionless inspection power is the part frontier-AI governance readers should notice. Most AI-transparency regimes have to construct an access right from scratch and then argue about its scope — a problem we have covered in which frontier labs let outsiders audit their safety cases. Insurance regulators did not have to construct anything. Examination and market-conduct authority over licensed entities already existed; SB 815 simply names the automated system as a thing within its reach.
The mechanism is insurance regulation, not health-IT or device law
This is the structural point most summaries miss. Neither statute regulates software as a product. Neither creates a registration, a pre-market review, a conformity assessment, or a technical standard. Neither is enforced by anyone with health-IT or medical-device competence.
They regulate a licensed financial entity’s decision process. The enforcement bodies are DMHC and CDI in California and the Texas Department of Insurance in Texas, using the tools those agencies already have: licensure, market conduct examination, corrective action plans, administrative penalties, and in California’s case the existing independent medical review machinery that already sits on top of utilization review decisions. A plan that denies via algorithm is not committing a software offence; it is failing to conduct utilization review in the manner its licence requires.
That is a different posture from the FDA’s device pathway, where the question is whether a tool is a medical device and which submission route it needs — the contrast we draw in foundation models versus the FDA’s Software as a Medical Device framework. A utilization review tool generally is not a device at all: it informs a coverage decision, not a diagnosis or treatment decision. The device framework would never have reached it. Insurance law did, immediately, because the tool sits inside a process that was already regulated.
It is also worth separating this from the ordinary mechanics of payer approval, which have their own vocabulary and their own confusions — see prior authorization versus precertification for how those terms are and are not distinguished. SB 1120 and SB 815 do not change what prior authorization is. They change who is allowed to say no.
What the federal government has not done
There is no parallel federal requirement. It is worth being precise about this, because the state statutes are frequently discussed as though they were implementing something national.
The obvious federal candidate would be the Medicare Advantage and Part D rulemaking, since MA plans conduct utilization review at enormous scale and have been the subject of most of the reporting on algorithmic denials. CMS’s CY2027 MA and Part D final rule (published in the Federal Register of 6 April 2026, 91 FR 17384) does not impose one. On a text search of the rule, artificial intelligence appears only incidentally: in a response to comments touching FDA-cleared autonomous retinal screening, and in a request-for-information question about the use of AI and machine learning in risk-adjustment model calibration. Neither is a utilization-management mandate, and neither names a required human decision-maker. (We note the search method honestly: this is a long rule, and the finding is that no provision creates such a requirement, not a claim that the string appears an exact number of times.)
So the current US position is that a Californian or Texan with commercial coverage has a statutory human decision-maker for medical necessity, and a Medicare Advantage enrollee in the same state relies on the general MA coverage rules and the plan’s own governance. State insurance regulation reached this first, as it reached several other AI questions first — the same dynamic we traced in why no state requires AI liability insurance yet, and on the underwriting side in who underwrites frontier AI risk. Those pages deal with AI as the insured peril. This one is the mirror image: the insurer as the AI deployer.
Why this matters for frontier-AI governance
Almost every framework in this cluster — corporate frontier AI frameworks, the EU AI Act’s GPAI obligations, SB 53, the RAISE Act, NIST’s AI RMF — governs AI by describing the system: how capable it is, what risks it poses, what evaluations were run, what mitigations were applied, what gets disclosed. That approach has a well-known soft spot, which is that every one of those descriptions relies on a term nobody has defined consistently. We have documented the problem repeatedly, most directly in capability thresholds across 14 labs and regulators.
SB 1120 and SB 815 sidestep the whole apparatus. They do not ask how capable the tool is. They do not ask for a threshold, a safety case, an evaluation result, or a disclosure. They identify one decision with a known victim when it goes wrong, and they assign it to a human with a licence — someone who can be sanctioned by a board, sued individually, and asked under oath why they said no.
The trade-offs are real and cut both ways:
- It is enforceable without technical capacity. A regulator auditing SB 815 compliance asks who signed the adverse determination and whether they were a licensed reviewer. It does not need to understand the model, reproduce an evaluation, or referee a dispute about benchmark validity.
- It is narrow by construction. It works because the decision is nameable, the profession is licensed, and the process was already regulated. Very few frontier-AI harms have all three properties. You cannot write “only a licensed person may decide whether this model is safe to deploy,” because there is no licence.
- It does not evaluate anything. A tool that is biased, poorly validated, and wrong most of the time is fully compliant with both statutes as long as it never issues the final determination. The statutes fix authority, not quality. SB 1120’s paragraph (1) conditions gesture at quality — periodic review of performance and outcomes, no reliance on group data alone — but they are stated as standards, with no specified methodology, no published results, and no independent verification.
- Volume pressure is untouched. Neither statute limits how many recommendations a system may hand a reviewer, or how long the reviewer has. The Texas “wholly or partly” language is the only textual handle on rubber-stamping, and whether it bites is an enforcement question, not a drafting one.
The honest summary is that this is a governance pattern with a narrow domain of applicability and unusually high enforceability inside it. It belongs in the toolkit next to the descriptive frameworks, not instead of them. Health systems building internal governance for clinical and administrative AI face exactly this layering problem; see hospital AI governance committee versus corporate AI governance board for how the two structures differ.
Where this touches research administration
The connection is narrower than the general health-AI story, but it is genuine in two places.
Clinical research billing. Academic medical centres bill third-party payers for the routine care costs of clinical trial participants, and whether a given item is billable turns on a coverage-and-medical-necessity analysis performed before the trial opens. That analysis assumes the payer’s own medical-necessity determination is a clinical judgement that can be predicted, documented against, and appealed. When the payer-side determination is produced or shaped by an automated system, the appeal posture changes: in California and Texas, “this denial was issued by your automated system” is now a compliance argument on its own, independent of the clinical merits. Research billing compliance teams doing Medicare coverage analysis for clinical trials should know which of their payer mix sits under these statutes.
Institutions that are themselves utilization review agents. Texas regulates the utilization review agent, not only the insurer. Provider-sponsored health plans, university-affiliated plans, and captive or self-funded arrangements run by academic health systems can fall inside Chapter 4201 in their own right. An institution that has deployed an automated prior-authorisation or claims-triage tool on the payer side of its own operations is in scope for the flat prohibition and for the commissioner’s standing inspection power — the same institution whose research-computing governance committee is elsewhere debating AI acceptable-use policy.
What this is not: it is not an IRB or human-subjects question, and it is not a research security or export control question. If your interest in AI governance is confined to those areas, these statutes do not reach you.
The NIKOLAI angle: a statutory instance of a governance element
NIKOLAI, CASRAI’s independent frontier-AI-safety dictionary, carries an element under track N9 (commitments and governance) called Accountable Decision-Maker and Sign-Off, defined as a record capturing “the named role (and, where published, the named person) who makes or approves a threshold determination, risk-acceptance decision, deployment decision, redaction, or framework change, together with the approval record itself (what was approved, by whom, and when).”
SB 1120 and SB 815 are the closest thing in US law to a mandated instance of that element, and the comparison is instructive in both directions. In frontier-AI frameworks, the accountable decision-maker is nearly always a voluntary internal governance artefact, unnamed in statute — which is the gap we documented in what SB 53 actually requires of accountable decision-makers, where the obligations land on the developer as an entity and no individual role is specified. In utilization review, the reverse holds: the statute names the role, and does so by reference to an existing licensure regime it did not have to invent.
Standard caveat, which matters: NIKOLAI is CASRAI’s own independent and unendorsed dictionary. The N9 element is CASRAI’s synthesis, not an agreed definition from any named organization, and no legislature, regulator, or health plan has adopted it. Any alignment we draw between a NIKOLAI element and a statutory provision is a shadow mapping — our reading, not the organization’s — unless that organization has filed a Mapping Declaration confirming it. None has here. Neither the California nor the Texas legislature has any relationship with NIKOLAI, and neither used its terminology.
A short diligence checklist
If you buy, build, or oversee a tool that touches utilization review in California or Texas:
- Establish which statute reaches you. California turns on whether you are a health care service plan (DMHC, H&S Code § 1367.01) or a disability insurer (CDI, Ins. Code § 10123.135). Texas turns on whether you are a utilization review agent under Insurance Code Chapter 4201 — a status that catches delegated entities and vendors, not only carriers.
- Do not rely on a vendor’s “human in the loop” claim. Ask what the reviewer sees, whether the recommendation is pre-populated, how long the median review takes, and whether the reviewer can see the inputs the model used. Texas’s “wholly or partly” is the operative phrase.
- Confirm reviewer competence, not just licensure. California’s text says a licensed physician or a licensed health care professional “competent to evaluate the specific clinical issues involved.” That is a matching requirement between reviewer specialty and case, and it is auditable.
- Prepare for inspection as a default state. California requires the tool to be “open to inspection for audit or compliance reviews”; Texas gives the commissioner an at-any-time audit power. Neither is triggered by a complaint. If you could not produce model documentation, criteria, and decision logs on request today, you are already out of position.
- Keep the periodic review evidenced. California’s requirement that performance, use, and outcomes be “periodically reviewed and revised to maximize accuracy and reliability” is a documentation obligation in practice. Undocumented review is indistinguishable from no review.
- Watch the Texas application date. Effective 1 September 2025, but applying to plans delivered, issued for delivery, or renewed on or after 1 January 2026.
Frequently asked questions
Does SB 1120 ban health plans from using AI?
No. It permits AI, algorithms, and other software tools in utilization review, subject to conditions, and then bars the tool from making the medical-necessity determination itself. Screening, information gathering, flagging, prioritising, and routing are not prohibited.
Is Texas SB 815 stricter than California SB 1120?
On the face of the text, the Texas rule is the blunter of the two: a single prohibition on using an automated decision system to make an adverse determination “wholly or partly,” attached to a standing audit power. California’s version pairs the same core prohibition with a longer list of use conditions. They are best read as two drafting styles aimed at the same outcome rather than as a strict-versus-lenient pair.
Do these laws apply to Medicare Advantage plans?
Medicare Advantage is federally regulated, and MA organizations are generally not subject to state insurance regulation of their MA products. The CY2027 MA and Part D final rule does not create a parallel federal requirement, so MA enrollees do not get the state statutory protection by this route.
Who enforces these requirements?
The Department of Managed Health Care and the Department of Insurance in California, depending on which entity type is involved, and the Texas Department of Insurance in Texas. They are insurance regulators using licensure and market-conduct tools, not technology regulators.
Does either law require the AI tool to be evaluated or validated?
Not in the sense this cluster usually means. California requires periodic review of performance, use, and outcomes, with no specified methodology, no published results, and no independent evaluator. Texas requires no evaluation at all; it relies on the prohibition plus the commissioner’s inspection power.
Is “Physicians Make Decisions Act” the statute’s real name?
No. The chaptered text of SB 1120 has no short-title section. The name comes from the sponsoring campaign. Cite Health and Safety Code § 1367.01 and Insurance Code § 10123.135.
Primary sources
- California SB 1120 (2023–2024 Regular Session), chaptered 28 September 2024, Chapter 879 — amending Health and Safety Code § 1367.01 and Insurance Code § 10123.135. Full text via California Legislative Information.
- Texas SB 815, 89th Legislature Regular Session — adding Insurance Code § 4201.156; signed 20 June 2025, effective 1 September 2025, applicable to plans delivered, issued for delivery, or renewed on or after 1 January 2026. Bill history and text via Texas Legislature Online.
- CMS, Medicare Advantage and Part D final rule for CY2027, 91 FR 17384 (6 April 2026), via GovInfo — cited here for the absence of a parallel federal requirement.








