● AI Governance Section · Industry deep-dive

AI Governance in Healthcare: A Step-by-Step Case Study

One health system, six AI systems already running, and no governance function. This walks the whole build — ten implementation steps, then six worked scenarios where the framework meets a decision that has to be made on a Tuesday.

Meridian Health System is fictional — a composite. Everything it collides with is real: Section 1557, HTI-1, FDA SaMD, HIPAA, CMS coverage rules, the EU AI Act, and the published research on models that failed in the field.

Step zero

The organization

Governance design is not universal. The same six factors that decide how much process an organization can sustain — size and maturity, industry and regulatory exposure, products and use cases, objectives and risk tolerance — produce a very different answer for a hospital than for a marketing agency. So the profile comes first, because in healthcare almost every fact about the organization drags a specific legal obligation behind it.

4acute-care hospitals, 1,100 beds — one of them the region's safety-net facility
62ambulatory and specialty clinics across two states
14,000employees, including 2,300 employed clinicians
$2.6Bnet patient revenue
48,000lives in Meridian Advantage, its owned Medicare Advantage plan
1research institute, with an EU academic collaboration and a device co-development deal

Meridian takes Medicare and Medicaid, runs a certified EHR, employs clinicians, insures patients, and does research. Each of those is a separate regulatory doorway, and AI walks through all of them at once.

Why each fact matters

Fact about MeridianWhat it triggers
Receives federal financial assistance from HHSIt is a covered entity under Section 1557 of the ACA. Since 1 May 2025 it must make reasonable efforts to identify patient care decision support tools that use race, color, national origin, sex, age or disability as an input variable, and must have policies to mitigate the resulting discrimination risk.
Holds protected health informationHIPAA Privacy and Security Rules: minimum necessary, business associate agreements with every AI vendor touching PHI, and a defensible position on de-identification and secondary use for model training.
Runs a certified EHRONC/ASTP HTI-1 decision support intervention criteria. Predictive DSIs surfaced through the certified system carry 31 source attributes the developer must disclose, plus intervention risk management practices. Meridian is the consumer of that disclosure — and has to actually read it.
Deploys software that diagnoses or informs treatmentFDA jurisdiction over Software as a Medical Device. Some tools are cleared devices, some fall in the clinical decision support carve-out, and some are being used outside the indication they were cleared for — which is where the exposure sits.
Owns a Medicare Advantage planCMS coverage rules. Under 42 CFR 422.101(b)(6) and the February 2024 CMS FAQs, an algorithm may assist a coverage determination but cannot be the sole basis for denying or terminating care; the individual's circumstances must be evaluated.
Operates in two states, one with an active AI statuteState law — generative-AI disclosure in patient communications, limits on AI in utilization review, and algorithmic-discrimination duties. This layer changes fastest and needs a standing watch, not an annual review.
Co-develops an imaging model with an EU market pathEU AI Act. An AI system that is a safety component of a regulated medical device is high-risk by the Annex I route, regardless of Annex III. Meridian also has to work out whether it is a provider or a deployer — the obligations differ sharply.
Accredited hospitalJoint Commission and CHAI published the first Responsible Use of AI in Healthcare guidance on 17 September 2025 — seven elements, from governance structures through to voluntary blinded reporting of AI safety events — with a voluntary certification following in 2026.
The point of this table

Nobody at Meridian has to memorize eight regimes. But somebody has to own the mapping — because a single AI tool can sit inside four of them simultaneously, and the four have different definitions of the same word. "Decision support" means one thing to ONC, another to FDA, and a third to the Office for Civil Rights.

The starting position

What the first inventory found

The governance effort began, as these usually do, with a question from the board after a news story: "How much AI are we actually running?" Nobody could answer. The chief information officer's estimate was "three, maybe four systems." A six-week sweep — EHR configuration review, procurement contract search, vendor security questionnaires, departmental interviews and an amnesty window for shadow tools — found eleven. Six of them mattered.

SystemWhat it doesHow it got hereInitial tier
SentinelSepsis early-warning score in the EHR, firing to nursingSwitched on by the EHR vendor as part of a module upgrade — no local decisionTier 1 · Critical
CareLensPopulation risk score selecting patients for the care-management programBought by population health in 2021, never reviewedTier 1 · Critical
ScribeAmbient documentation — listens to the visit, drafts the noteClinician-led pilot in 3 clinics, spread to 41 by word of mouthTier 2 · High
UM AssistPrioritizes and pre-scores prior-authorization requests in the health planVendor module, configured by the plan's operations teamTier 1 · Critical
PortalDraftGenerative AI drafting replies to patient portal messagesEnabled in a vendor release; on by defaultTier 2 · High
NoduleAILung nodule detection on CT, co-developed with a device partnerResearch collaboration heading toward commercial releaseTier 1 · Critical
Plus five lower-consequence systems: nurse-scheduling optimization, supply forecasting, a revenue-cycle coding assistant, a call-center routing model, and a chatbot on the marketing site.
The three findings that set the agenda

Nobody had decided. Four of the six arrived through a vendor release or a departmental purchase. There was no moment at which Meridian, as an organization, chose to deploy them — which means there was no moment at which anyone assessed them.

Nobody was monitoring. Not one of the six had a defined performance threshold, an owner accountable for it, or a documented review since go-live. Two had been running for over three years.

Nobody could produce evidence. Asked for the Section 1557 file on CareLens — what variables it uses, whether any of them proxy for a protected characteristic, what mitigation exists — the answer was that no such file existed. The compliance date had passed fourteen months earlier.

That last finding is the one that turns this from an IT project into a governance one. Meridian's problem was never that it lacked AI expertise. It was that AI had been entering the organization through six different doors, and none of those doors had a person standing at it.

The build

The ten-step build

Ten steps, in three phases. The first three stand the function up and are done once. The middle four are the gate every system passes through, and get repeated for every tool forever. The last three are the part almost everyone skips, and the part that actually determines whether a model is still safe two years after go-live.

Meridian ran phase A in ten weeks with a working group of eleven people, none of them full-time on it. That is roughly the floor. Compressing it further tends to produce a charter nobody follows.

Phase A · Weeks 1–10 Stand up the function

Done once. The output is a committee with a mandate, a list of what you own, and a rule for how much scrutiny each item gets.

1Give it a mandate, and give it a shape

Effective AI governance is distributed but accountable. The failure mode at both extremes is well documented: appoint one overwhelmed owner and the queue becomes the bottleneck; declare that "everyone owns AI" and nobody does. Meridian's answer was a chartered committee with real decision rights and a named executive sponsor who could be fired for getting it wrong.

The AI Governance Committee was chartered by the board's quality committee — deliberately not by IT — and given three powers written into the charter: it can require validation before deployment, it can suspend a live system, and it can escalate to the board without going through the executive team. Without at least the second of those, the committee is an advisory body, and advisory bodies get routed around.

Who sits on it

  • Chief Medical Officer (chair) — because in a health system, the credible authority to stop a clinical tool is clinical, not technical.
  • Chief Nursing Officer — most clinical AI in a hospital fires at nurses first. Omitting nursing is the most common composition error.
  • Chief Information Officer and Chief Data Officer — the systems and the data lineage.
  • Chief Compliance Officer and Privacy Officer — Section 1557, HIPAA, the OCR relationship.
  • Chief Information Security Officer — model supply chain, PHI egress, vendor security.
  • General Counsel — contracting, liability allocation, malpractice interface.
  • Health Equity lead — subgroup performance is not a side topic here; it is a compliance obligation.
  • Two practising clinicians and one bedside nurse, rotating annually — the people who will actually receive the alerts.
  • A patient or community representative — the only voice in the room with no institutional incentive.

Around the committee, the standard three-lines structure: the deploying department owns the risk day to day (first line), compliance and the governance office set policy and challenge (second line), internal audit tests independently (third line). Internal audit does not sit on the committee — that would compromise the third line.

OwnerBoard quality committee charters it; CMO chairs
InputsExisting committee structure, delegated-authority matrix
OutputSigned charter, RACI, meeting cadence, escalation path
Done whenThe committee has suspended or blocked something at least once
If you skip itGovernance becomes a checklist attached to procurement. It will catch new purchases and miss every tool that arrives in a vendor release — which, at Meridian, was four of the top six.
2Build the inventory — and keep it alive

You cannot govern what you cannot enumerate, and in healthcare the enumeration is genuinely hard, because most AI does not arrive labelled as AI. It arrives as "the new module," "clinical decision support," "the risk score," or "a feature in the release notes."

The five discovery channels Meridian ran in parallel

  1. EHR configuration review — every active predictive model, scoring rule and decision support intervention configured in the production environment, pulled with the vendor's help. This alone found Sentinel.
  2. Contract and procurement search — every agreement since 2018 containing "artificial intelligence," "machine learning," "algorithm," "predictive," "model" or "score." Legal ran it; it surfaced obligations nobody remembered signing.
  3. Network and SaaS discovery — what is actually being talked to, from the security side. This is how shadow tools show up.
  4. Departmental interviews — a standing 30-minute slot with each service line, with one question: "What software makes a suggestion or a prediction that changes what someone does?" Never ask "do you use AI." People say no and then describe an AI.
  5. A shadow-AI amnesty — a 30-day window to declare anything unofficial with an explicit no-blame guarantee, signed by the CEO. Meridian got nine disclosures, including a residency program pasting de-identified case summaries into a consumer chatbot.

The register is a living record, not a spreadsheet produced once. Each entry carries: system name, vendor, version, business owner, clinical owner, what decision it influences, populations affected, data it consumes, whether PHI leaves the environment, regulatory classification, risk tier, validation status, monitoring plan, last review date, and next review date.

The intake trigger that makes it stay current

An inventory decays unless something forces new entries into it. Meridian wired three triggers: no purchase order over $0 clears procurement without an AI screening question; no change request touching clinical decision support clears the change advisory board without a governance reference number; and every vendor release note is reviewed for new AI functionality before the upgrade window. The third one is the one that catches PortalDraft-style surprises.

OwnerAI governance office; CIO accountable for completeness
InputsEHR config, contracts, network discovery, interviews, amnesty
OutputAI register with a named owner per entry
Done whenThree intake triggers are live and adding entries without being asked
If you skip itEvery subsequent step operates on a sample rather than a population. Your risk profile is set by the tools you didn't find, not the ones you did.
3Classify: risk tier first, then regulatory routing

Two different classifications, often confused. Risk tier decides how much internal scrutiny a system gets. Regulatory routing decides which external obligations attach. A tool can be low-tier internally and still be squarely inside Section 1557.

The tiering question

Meridian tiers on two axes — the consequence if the output is wrong, and how much independent human judgment stands between the output and the action. A model that suggests something to a clinician who is going to check it anyway is meaningfully different from a model whose output is executed.

TierTestMeridian examplesWhat it requires
Tier 1Wrong output can cause physical harm, or determines access to care or coverageSentinel, CareLens, UM Assist, NoduleAIFull committee review, local validation, subgroup analysis, board reporting, quarterly monitoring
Tier 2Enters the medical record or reaches a patient, but a clinician reviews before it actsScribe, PortalDraftCommittee review, targeted validation, attestation controls, semi-annual review
Tier 3Operational; affects staff or workflow, not clinical decisions or coverageScheduling, supply forecasting, call routingDelegated review by a subgroup, annual attestation
Tier 4No PHI, no decision influenceMarketing site chatbotRegister entry and a privacy check
The tiering trap

"A human reviews it" is only a mitigation if the human realistically can and does. A nurse receiving 400 alerts a shift is not a meaningful reviewer, and neither is a clinician signing 30 AI-drafted notes at the end of a clinic. Meridian's rule: you may only claim human review as a control if you can state the review rate, the time available per item, and the observed override rate. If you cannot measure it, you cannot claim it.

The routing question

For each system, the committee answers five questions in order, and the answers determine the compliance file that has to exist:

  1. Does it diagnose, treat, or drive a treatment decision? If yes, it is potentially a device — check FDA clearance status and, critically, whether Meridian's use matches the cleared indication.
  2. Does it support clinical decision-making about a patient? If yes, it is a patient care decision support tool under Section 1557 — automated or not — and the identify-and-mitigate obligation applies.
  3. Is it surfaced through the certified EHR? If yes, obtain and read the HTI-1 source attributes. Absence of that disclosure is itself a finding.
  4. Does it influence coverage, payment or utilization? If yes, CMS rules and state utilization-review law apply, and the "not the sole basis" line becomes an operational control, not a slogan.
  5. Does it, or its outputs, reach the EU? If yes, work out provider versus deployer, and which high-risk route applies.
OwnerGovernance office proposes; committee ratifies tier
InputsRegister entry, vendor documentation, intended-use statement
OutputTier assignment plus a regulatory routing sheet per system
Done whenEvery Tier 1 and 2 system has a routing sheet with named obligation owners
If you skip itEverything gets the same review, which means either the trivial tools consume the committee's capacity or the dangerous ones get a trivial review. In practice, both.
Phase B · Per system The gate every system passes through

Repeated for every tool, new or already running. Existing systems are back-fitted in tier order — Tier 1 first, regardless of how long they have been live.

4Intake and pre-deployment review

The intake form is the single highest-leverage artifact in the whole programme, because it forces the requesting department to answer questions they have usually not asked. Meridian's is two pages and deliberately hard to complete without the vendor's help — which is the point, since it surfaces immediately whether the vendor will actually answer.

What the form demands

  • Intended use, stated narrowly. Not "improves sepsis detection" but "generates a score for adult inpatients on medical-surgical units, intended to prompt a nursing assessment." Every downstream control depends on this sentence, and most disputes trace back to it being written vaguely.
  • Out-of-scope uses. Explicitly: paediatrics, obstetrics, ED, ICU, ambulatory — say which are excluded and why. A model validated on adults and rolled out to a paediatric ward is a textbook deployment-context failure.
  • The decision it influences, and who makes it. Named role, not "the care team."
  • Training data provenance — population, timeframe, sites, and demographic composition against Meridian's own patient mix.
  • Performance claims with their evidence — and specifically, where the claimed numbers were measured. Vendor-reported performance on the development population is not evidence about Meridian's population.
  • Input variables, with a flag on any that are protected characteristics or known proxies for them.
  • Data flow — does PHI leave the environment, to whom, under what agreement, retained how long, used for vendor model improvement yes or no.
  • The counterfactual — what happens today without this tool, and what specifically improves. A surprising share of intake requests die honestly at this question.

Review is time-boxed: 15 business days for Tier 1, 10 for Tier 2. A governance function that cannot commit to a turnaround gets bypassed, and a bypassed gate is worse than no gate because it creates a false record of control.

OwnerRequesting department completes; governance office triages
InputsVendor documentation, HTI-1 source attributes, FDA status
OutputCompleted intake, risk assessment, committee decision with reasons
Done whenA written decision exists — approve, approve with conditions, pilot, or decline
If you skip itYou inherit the vendor's intended use, the vendor's performance claims and the vendor's risk assessment — and the vendor is not the entity the Office for Civil Rights will contact.
5Data, privacy and the contract

Three questions decide whether the deal is acceptable, and they are usually settled by people who never speak to each other — so the governance function's job here is mostly convening.

Does PHI leave, and under what terms?

Any vendor processing PHI on Meridian's behalf needs a business associate agreement, and the BAA is where the AI-specific terms have to live because the standard template does not contemplate them. Meridian's additions: the vendor may not use Meridian's data to train or improve models serving other customers without a separate written agreement; the vendor may not attempt to re-identify de-identified data, and must contractually bind its own subprocessors to the same; data is returned or destroyed on termination with certification; and Meridian gets notice of model version changes before they ship. That last clause is the one that turns Scenario 5 from an incident into a scheduled review.

What is the minimum the model actually needs?

Minimum necessary is a HIPAA obligation, not a preference, and model builders default to maximalism. The discipline is to ask, feature by feature, what it is doing for performance — which usually shrinks the feature set and, as a side effect, makes the model easier to explain and less likely to be quietly encoding something you would rather it did not.

Can we use this data for training at all?

Secondary use of clinical data for model development sits at the intersection of HIPAA, the consent under which the data was collected, and institutional research policy. Meridian's route: de-identify to the Safe Harbor standard where possible, use Expert Determination where the required fields make Safe Harbor impossible, and route anything that cannot be de-identified through the IRB. For the multi-site research collaboration, federated learning resolved a hard constraint — partner hospitals could not share patient records, so the model trains locally at each site and only parameter updates are pooled. The model goes to the data; the raw records never move.

The clause that is missing from most healthcare AI contracts

An evidence clause: the vendor will, on request and within a defined window, provide subgroup performance data, the training population's demographic composition, and the model's input variables — in a form Meridian can hand to a regulator. Without it, Meridian's Section 1557 obligation to make reasonable efforts to identify problematic input variables runs into a vendor that treats the feature list as a trade secret. Negotiate this before signature; afterwards there is no leverage.

OwnerPrivacy officer and general counsel; CISO on data egress
InputsData flow diagram, BAA, vendor security assessment
OutputExecuted BAA with AI terms, documented minimum-necessary basis, de-identification determination
Done whenYou can answer "where does our data go and what may they do with it" in one page
If you skip itYou discover at renewal that your patients' records have been training a product you now compete with, and that you agreed to it.
6Local validation — the silent trial

This is the step that separates real healthcare AI governance from paperwork. Vendor performance is a hypothesis about your population, not a measurement of it. Models degrade when moved between institutions because the patients differ, the documentation practices differ, the coding differs, and the care pathways differ.

Meridian's rule for every Tier 1 clinical system: a silent period of no less than 90 days, during which the model runs on live data and its outputs are logged but never shown to anyone. Then compare against the outcome that actually happened.

What gets measured

  • Discrimination — how well the model separates cases from non-cases, in Meridian's population, at the threshold it will actually run at.
  • Calibration — when it says 30%, does roughly 30% happen? A model can rank well and still be badly miscalibrated, which destroys the meaning of any threshold set on it.
  • Sensitivity and specificity at the operating threshold — not at the optimal threshold in a paper.
  • Alert burden — alerts per 100 patients per day, and number needed to evaluate: how many flagged patients a clinician must assess to find one true case. This number, more than any other, predicts whether the tool survives contact with a ward.
  • Incremental value — how many true cases the model catches that the existing process misses, and how much earlier. A model that mostly flags patients the nurse already worried about is not adding safety; it is adding clicks.
  • Subgroup performance — the same metrics, computed separately. Covered in step 7.
This is TEVV, and the last two letters are the ones that get dropped

Testing, evaluation, verification and validation. Verification asks whether the system meets its specification — did we build it right. Validation asks whether it meets the real need in its actual context of use — did we build the right thing, for these patients, on this ward. A tool can pass verification cleanly and fail validation completely, and in healthcare that gap is where patients get hurt.

OwnerClinical informatics and data science; clinical owner signs off
Inputs90 days of shadow-mode output, matched outcome data
OutputLocal validation report with a go / no-go recommendation and a chosen threshold
Done whenThe committee has seen local numbers, not vendor numbers
If you skip itSee Scenario 1. You deploy a model whose real local performance is materially worse than advertised, and you find out from an adverse event review rather than a validation report.
7Bias and equity assessment

In healthcare this is not an ethics exercise appended to the end. Since 1 May 2025 it is a compliance obligation with a named regulator, and the assessment is the evidence that the obligation was met.

The four questions, in order

  1. What is the model actually predicting? Not what it is called — what the label was. This is the single most productive question in healthcare AI bias work, because the most consequential known failure was a label problem, not a model problem.
  2. Does any input variable measure a protected characteristic, or proxy for one? Race, colour, national origin, sex, age and disability directly; and the proxies — ZIP code, insurance type, prior utilization, prior spend, language, and any clinical formula with a race coefficient baked in.
  3. Does performance hold across subgroups? Sensitivity, specificity, calibration and alert rate computed separately by race and ethnicity, sex, age band, primary language, insurance type and disability status where recorded — with pre-agreed thresholds for what counts as a material gap.
  4. Does the resulting allocation hold? Equal accuracy with unequal access is still a problem. Meridian tracks who ends up receiving the intervention, not only how well the model scores.
Removing the variable is not the mitigation

Dropping race from the feature set does not remove race from the model — it removes your ability to see it. The proxies remain, and now the disparity is invisible. Mitigation is measurement plus a corrective: re-label to the outcome you actually care about, recalibrate by subgroup, adjust the threshold, restrict the intended use, or decline to deploy. All five are legitimate; pretending the variable was never there is not.

The output is a written assessment per Tier 1 and Tier 2 system, refreshed on the monitoring cycle, stating what was examined, what was found, what was done about it, and who signed. This document is the file that gets produced if anyone ever asks.

OwnerHealth equity lead with data science; compliance officer countersigns
InputsValidation dataset with demographics, input variable list, allocation data
OutputBias and equity assessment; Section 1557 evidence file
Done whenYou could hand the file to OCR tomorrow without preparing anything
If you skip itSee Scenario 2. The disparity does not announce itself — it operates quietly for years, and the discovery event is usually external.
Phase C · Continuous Live operation

The steps that determine whether a model that was safe at go-live is still safe in year three. This is where most programmes are thinnest.

8Deployment design: the human, the disclosure, the training

A validated model can still fail on the ward if the deployment is designed badly. Three things get designed deliberately.

The human in the loop, specified concretely

"A clinician reviews the output" is not a control until you can say who, with what information, in what time, with what authority to disagree, and with what happens when they do. Meridian specifies all five per system, and requires that override be one action, never a form. If disagreeing with the model costs more effort than complying, you have built automation bias into the workflow and then called it human oversight.

What the patient is told

Disclosure obligations vary by state and by use, and are tightening. Generative AI drafting clinical communications to patients attracts an explicit disclaimer requirement in some states unless a licensed human reviews the message before it goes — which makes the review step a compliance control, not just a quality one. Meridian's policy is a floor above the strictest state it operates in, applied everywhere, because two policies for two states is a policy nobody follows.

Training that is role-specific

AI literacy has moved from good practice to legal obligation — the EU AI Act requires providers and deployers to ensure a sufficient level of AI literacy among staff operating AI on their behalf, proportionate to role and risk, and the Joint Commission guidance expects role-specific training as one of its seven elements. Meridian runs three tracks: 20 minutes for all staff on what AI is in use and how to report a concern; 90 minutes for clinicians who receive AI output, focused on when to distrust it; and a half-day for system owners and the committee.

OwnerClinical owner with informatics; education for training
InputsValidation report, workflow analysis, state disclosure requirements
OutputWorkflow design, override mechanism, disclosure language, training completion
Done whenOverride rate is being measured and is not zero
If you skip itYou get a technically correct system that clinicians route around, or worse, one they trust more than they should.
9Monitoring, drift and incident response

Model performance is not a property; it is a reading taken at a point in time. Populations shift, documentation practices change, a coding update alters an input distribution, a new EHR release changes how a field is captured — and performance moves without anyone touching the model.

What Meridian watches, per Tier 1 system

  • Input drift — has the distribution of what goes in changed against the validation baseline?
  • Output drift — has the score distribution or alert rate moved?
  • Performance — sensitivity, specificity and calibration recomputed against realized outcomes, quarterly.
  • Subgroup performance — the same, disaggregated, on the same cycle. Drift is not evenly distributed; it usually appears in the smallest subgroup first.
  • Human interaction — override rate, response time, alert acknowledgment. A collapsing override rate means either the model got better or the humans stopped reading. Those look identical in the data and require asking.
  • Operational health — latency, availability, failed calls. A clinical model that times out silently is a safety issue.

Every metric has a threshold set in advance and an owner, and crossing it triggers a defined action rather than a discussion about whether it matters.

The incident pathway

AI safety events need a route into the existing patient-safety reporting system, not a parallel one — clinicians will not use a second system. Meridian added an "AI involved?" flag to the standard event report and a four-level severity ladder:

LevelTriggerResponse
Sev 1Patient harm reached the patient, AI output contributedImmediate suspension, executive and board notification within 24h, root cause analysis, regulatory reporting assessment
Sev 2Near miss — caught before reaching the patient — or a confirmed subgroup performance gapCommittee review within 5 days, mitigation plan, consider restricting use
Sev 3Monitoring threshold breached, no known patient impactOwner investigates within 10 days, reports to committee
Sev 4Usability complaint, alert burden, workflow frictionLogged, trended, reviewed in aggregate quarterly

Sev 4 matters more than it looks. Alert-burden complaints are the leading indicator for a Sev 1 — because the mechanism by which a noisy model causes harm is that people stop reading it.

OwnerSystem owner monitors; patient safety owns the incident pathway
InputsValidation baseline, live output logs, outcome data, event reports
OutputMonitoring dashboard, quarterly report to committee, incident record
Done whenA threshold breach has triggered its defined action without anyone escalating manually
If you skip itYour validation report describes a model that no longer exists. Drift is silent by construction — nothing alerts you that a model has quietly become worse.
10Change control, revalidation and decommission

The governance question nobody asks at purchase: what happens when the vendor changes the model? In healthcare AI this is routine and often invisible — a new version ships in a release, and the tool your clinicians trust is not the tool you validated.

Change control

Meridian's rule: a model version change is a change, and goes to the change advisory board with a governance reference. For Tier 1 systems, a version change triggers abbreviated revalidation before it reaches production — a shadow comparison of old versus new on recent data, and confirmation that subgroup performance has not moved. The contractual notice clause from step 5 is what makes this possible; without it, you learn about version changes from your monitoring dashboard, after the fact.

On the device side there is now a formal mechanism for this. FDA's predetermined change control plan lets a manufacturer specify in advance which modifications to an AI-enabled device may be made without a new marketing submission, and how they will be validated. For Meridian as a deployer, the practical question is simple and worth asking every vendor: is there a PCCP, and what does it permit you to change without telling us?

Scheduled revalidation

Independent of any change: Tier 1 annually, Tier 2 every two years. Same protocol as the original validation, against current data. Revalidation regularly finds that a model still works but its clinical context has moved — the care pathway changed, and the model is now optimizing for something the organization stopped caring about.

Decommission

The most under-designed step in AI governance. Retiring a clinical model requires: notice to users with a reason, a defined fallback process, removal from the workflow rather than just disabling the alert, retention of the model's historical outputs where they informed documented clinical decisions, data return or destruction under the contract, and a register update. Meridian retired two of the eleven systems in year one. Neither had a decommission plan, and one of them had outputs embedded in three years of clinical notes that had to be interpretable retrospectively.

OwnerSystem owner; change advisory board enforces the gate
InputsVendor change notice, PCCP scope, revalidation results
OutputChange record, revalidation report, or decommission plan
Done whenA vendor version change has been stopped at the gate at least once
If you skip itSee Scenario 5. Your validated system silently becomes a different system, and the first signal is a complaint.
Scenario analysis

Six worked scenarios

The ten steps are the machine. Scenarios are what the machine is for. Each of the six below is a real decision Meridian's committee had to make, with a date on it and people waiting for an answer — the kind of thing that arrives on a Tuesday and cannot be deferred to the next quarterly meeting.

Each one follows the same shape: the situation, the signal that surfaced it, three options with their actual consequences, what was decided and why, the artifact it produced, and the transferable lesson. The options are not strawmen. In every case, at least two of the three were seriously argued in the room, and in several cases the option that lost is the one most organizations pick.

How to read the branch analysis

The three options in each scenario are genuine forks, not a right answer flanked by two obviously wrong ones. The marked branch is what Meridian chose, given Meridian's facts. Change the facts — a different patient population, a different contract, a different regulator — and a different branch becomes correct. What transfers is the reasoning, not the verdict.

01
Sentinel · sepsis early warning · Tier 1
The model that works everywhere except here
System
Vendor-supplied sepsis early-warning score, embedded in the EHR, firing to bedside nurses. Live on all four hospitals for 26 months before governance existed.
Vendor claim
AUC 0.83 in the vendor's published development cohort; marketing material cites “up to 20% reduction in sepsis mortality.”
Local result
AUC 0.64 across 38,900 adult admissions in a 90-day silent period. Sensitivity 34% at the vendor's default threshold.
Trigger
Step 6 local validation, run because the system was already live and had never been checked.
Regulatory hooks
HTI-1 predictive DSI source attributes; Joint Commission RUAIH pillars 4 and 6; internal patient safety reporting.
The situation

Sentinel was bought in a hurry during a quality improvement push, configured by the EHR team, and switched on. Nobody validated it locally. Nobody was monitoring it. When the governance committee's inventory sweep found it, the clinical sponsor who had championed it had left the organization eighteen months earlier, and the current owner of record was a build analyst who had never seen the vendor's performance claims.

The committee did not turn it off. It ran the Step 6 silent period against the live system: the model kept scoring, the alerts kept firing, and a parallel adjudication compared what Sentinel said against a chart-review gold standard on a stratified sample.

This was not a fishing expedition. There was a published reason to look. An external validation of a widely deployed proprietary sepsis model at an academic health system found an AUC of 0.63 against a vendor-reported 0.76–0.83, 33% sensitivity, alerts on 18% of all hospitalizations, and 67% of sepsis cases missed entirely — with roughly 109 alerts needed to identify one case of sepsis developing within four hours.1 Meridian wanted to know whether it had bought the same problem.

The signal

It had. The 90-day silent period produced numbers that were, if anything, slightly worse:

What was measuredVendor claimMeridian, 90 daysWhat it means operationally
Discrimination (AUC)0.830.64Barely better than a coin flip weighted by how sick someone looks
Sensitivity at default thresholdnot stated34%Two out of three sepsis cases never generate an alert
Alert ratenot stated17% of admissions (6,613 alerts)Roughly one alert per nurse per shift, on top of everything else
Number needed to evaluatenot stated9696 bedside evaluations to find one true early sepsis case
Median lead time vs. existing screen“hours earlier”42 minutesReal, but small
Cases where a nurse had already escalated first71%In most alerts, the human got there first anyway

The subgroup analysis found something the aggregate number hid. At the flagship academic hospital, AUC was 0.67. At the safety-net hospital, it was 0.58. The model leaned heavily on the cadence of structured vital signs and lab results, and the safety-net hospital charted vitals less frequently and ordered fewer routine labs. The model was not reading sickness there. It was reading documentation density — and reading it as health.

That is the finding that changed the conversation. A model that performs worst at the hospital serving the most vulnerable patients is not a performance problem. It is an equity problem wearing a performance problem's clothes.

The three options on the table
Option A
Leave it on at the vendor default. It has been running for two years; turning it off is its own risk.

6,613 alerts a quarter at 96 evaluations per true case. Alert fatigue is not confined to the alert that causes it — response rates degrade across every alert in the system, including the ones that work.

And the committee has now documented, in writing, that it knew. Knowing and doing nothing is a materially worse legal position than not knowing.

Rejected — the validation created a duty
Option B
Raise the alert threshold to cut the noise. Same model, fewer alerts, less fatigue.

Modelled at threshold 8: alerts fall to 6% of admissions, but sensitivity falls from 34% to 19%. You buy quiet by missing four out of five cases instead of two out of three.

This is the most commonly chosen option in practice, because it makes the complaint stop. It does not make the model work; it makes the model quieter about not working.

Rejected — trades detection for silence
Option C — chosen
Withdraw it house-wide. Redeploy only where the counterfactual is real, with a locally re-fit threshold and a hard sunset.

Alerts stop firing on 62 units where nurses were already catching it first. The model is re-fit on Meridian data with the vendor contractually on the hook, and returns only on the two night-shift units with the thinnest staffing ratios.

Alerts route to the rapid response nurse, not the bedside nurse. Six-month sunset: it comes back to committee or it switches off automatically.

Chosen
What the committee decided, and why

Sentinel was suspended house-wide within nine days of the validation report, using the committee's standing authority to suspend a live system without waiting for a quarterly meeting. That authority, written into the charter in Step 1 for exactly this reason, is what made a nine-day timeline possible instead of a ninety-day one.

The reasoning that carried the room was not the AUC. It was the 71%. A model whose alerts mostly arrive after a human has already acted is not adding safety; it is adding interruptions. The committee reframed the question from “is this model accurate?” to “what does a clinician do differently because of it, and is that different thing better?” On 62 units the answer was nothing. On two night-shift units with one nurse to seven patients, the answer was plausibly something — so that is where it went back, and only there.

  • Vendor was notified in writing that local validation did not reproduce the published performance, and asked for the training population demographics and input feature list under the evidence clause from Step 5.
  • The safety-net subgroup gap was logged as a patient safety finding and reported to the board quality committee, not just recorded in the AI register.
  • The re-fit model must clear a fresh 90-day silent period before it fires a single alert on the two remaining units.
Artifact this produced The local validation report template — now mandatory for every Tier 1 system — and one new question on the intake form that did not exist before: “In what fraction of cases would the clinician have acted anyway, without this output?” Nobody can answer it at intake. That is the point: it forces the silent period, because the silent period is the only way to find out.
Transferable lesson

A vendor's AUC is a claim about someone else's patients, measured on someone else's documentation habits. Discrimination is the least useful number in the deployment decision. The ones that decide it are alert burden, number needed to evaluate, and incremental value over what the humans already do — and none of those can be read off a datasheet.

Discussion questions
  1. Sentinel ran for 26 months before anyone checked. What in your organization has been running longest without a local validation, and who would you have to ask to find out?
  2. If the safety-net gap had been 0.67 versus 0.65 instead of 0.67 versus 0.58, at what point does a subgroup difference stop being noise and become a finding you must act on? Write the number down before you need it.
  3. Option B is what most organizations choose. What would have to be true in your governance process for Option B to be visibly unacceptable rather than obviously reasonable?
02
CareLens · population health risk score · Tier 1
The model with no protected variable that discriminated anyway
System
Vendor risk-stratification score ranking the 48,000-life Medicare Advantage population and the ACO attributed population for enrollment into intensive care management. Roughly 3,100 enrollment slots a year.
Label
Predicted total cost of care over the next 12 months.
Inputs
Claims, utilization, diagnosis codes, pharmacy fills, prior cost. Race, ethnicity and sex are not input variables.
Trigger
Step 7 bias and equity assessment, prompted by a care manager's observation.
Regulatory hooks
ACA Section 1557 patient care decision support tools (compliance date 1 May 2025); CMS Medicare Advantage rules; Joint Commission RUAIH pillar 6.
The situation

CareLens decides who gets a care manager. A care manager is a nurse who calls you, arranges your transport, reconciles your medications, and notices when you have stopped filling a prescription. For a patient with four chronic conditions and no car, it is the difference between managed illness and a series of admissions. There are 3,100 slots and roughly 61,000 attributed lives, so the ranking is the entire allocation mechanism.

The system was procured by the health plan, not the hospitals, and had never been reviewed as a clinical tool because on paper it is not one. It does not diagnose. It does not treat. It ranks.

The signal

A care manager mentioned in a huddle that her panel had almost nobody from the county the safety-net hospital serves. An analyst pulled the numbers to check. Patients from that catchment were 31% of the attributed population and 14% of the enrolled.

The first instinct in the room was to check the inputs for race. Race was not an input. Neither was ZIP code, in any direct form. By the standard the organization had been applying — we do not feed it protected characteristics — CareLens was clean.

The Step 7 assessment asks its questions in a specific order, and question one is not about inputs. It is “what is the model actually predicting?” The answer was cost. Not illness, not need, not deterioration. Cost.

This is a known and quantified failure. A widely used commercial risk algorithm applied to roughly 200 million people in the United States used healthcare cost as a proxy for healthcare need. At any given risk score, Black patients turned out to have 26.3% more chronic illnesses than white patients with the same score. Correcting the label — predicting illness rather than expenditure — raised the share of Black patients who would receive additional help from 17.7% to 46.5%.2

Meridian ran the same test on its own data. At an identical CareLens score, patients in the lowest income quartile of the attributed population had on average 1.9 more active chronic conditions and 2.4 more uncontrolled-condition indicators than patients in the highest quartile. The model was systematically ranking sicker poor patients below healthier affluent ones, and it was doing so correctly — because affluent patients with better access do consume more care, and therefore do cost more. The model was not broken. It was answering the question it had been given.

Under Section 1557, a patient care decision support tool is any automated or non-automated tool used to support clinical decision-making. Deciding who receives intensive care management is squarely inside that. The rule requires reasonable efforts to identify tools that use race, color, national origin, sex, age or disability as input variables, and policies to mitigate the risk of discrimination those tools create. Meridian had satisfied the first duty trivially — no protected variables — and completely failed the second.

The three options on the table
Option A
Add a race-based correction factor to the output, adjusting scores upward for underserved groups.

There is nothing to remove — race was never an input — so this means adding an explicit race-based adjustment to a resource allocation tool. That is a considerably harder thing to defend to a regulator than the problem it fixes.

It also treats the symptom. The score still means “expected cost.” You have applied a patch on top of a number that measures the wrong thing.

Rejected — new exposure, old mechanism
Option B
Leave the model alone; impose an enrollment quota so each service area gets slots proportional to its share of the population.

The disparity metric goes green immediately. The dashboard looks fixed at the next board meeting.

But within each service area the ranking is still by cost, so you are still enrolling the wrong patients — now with a fairness metric certifying that you are not. The measurement has been repaired and the mechanism has not.

Rejected — fixes the metric, not the harm
Option C — chosen
Change the label. Predict avoidable deterioration and active uncontrolled chronic conditions instead of cost, then re-rank and remediate backwards.

Requires the vendor to re-target the model, or replacement. It took seven months and a contract renegotiation. Two of the three shortlisted vendors could not say what their label was.

Re-scoring the prior twelve months under the new label identified 640 patients who would have qualified and were never enrolled. All 640 got outreach.

Chosen
What the committee decided, and why

The committee treated the label as the defect, because the label was the defect. Every downstream mitigation — reweighting, quotas, corrections, thresholds — is an attempt to talk a model out of answering the question it was trained on. The only durable fix is to change the question.

The harder decision was the backward look. Nothing compelled Meridian to re-score the past. The committee did it anyway, on the reasoning that if you discover an allocation mechanism has been misallocating, the people it misallocated away from are identifiable, still your patients, and still sick. The 640 outreach calls were the most expensive line item in the remediation and the only one that helped anyone who had already been harmed.

  • Section 1557 documentation was written to record both duties: the identification exercise (no protected inputs) and the mitigation analysis (proxy discrimination via the target variable, found and corrected).
  • The disparity check was made a standing quarterly monitor, not a one-time assessment — enrollment rate by service area, income quartile and language, with a defined threshold that triggers review.
  • Procurement standard added: a vendor who cannot state the model's target variable in one sentence does not proceed to evaluation.
Artifact this produced The label interrogation, now question one of every intake: “What is the label, and what did the label have to stand in for?” Cost stands in for need. Prior diagnosis stands in for disease. Length of stay stands in for severity. Being referred stands in for needing a referral. Every proxy carries the history of who had access, and the model learns that history as if it were biology.
Transferable lesson

Removing the protected variable is not mitigation, and its absence is not a defence. This model had no race input, no ZIP code, and no ethnicity flag, and it still allocated care by proxy for race and income — because the outcome it was trained to predict had already absorbed decades of unequal access. Audit the target variable first. It is where the bias lives, and it is the one thing an input audit cannot see.

Discussion questions
  1. Name three models in your organization and state each one's target variable in one sentence. If you cannot, that is the finding.
  2. Meridian re-scored the past and made 640 calls. What is your threshold for retrospective remediation — and who has the authority to authorise the cost of it?
  3. Option B produces a green dashboard. How would your governance process distinguish a metric that improved because the harm stopped from one that improved because the measurement changed?
03
Scribe · ambient clinical documentation · Tier 2
The note that said “denies chest pain”
System
Ambient AI scribe listening to the encounter and drafting the clinical note. 610 clinicians, roughly 2,800 encounters a day.
Why clinicians love it
Measured 47 minutes a day returned per clinician; the highest-satisfaction technology deployment in the organization's history.
Trigger
A single urgent care physician who read her draft note carefully before signing it.
Regulatory hooks
HIPAA (audio is PHI); state recording-consent law across two states; medical record integrity and attestation; vendor BAA and secondary-use terms.
The situation

Scribe was the popular one. It was the system nobody wanted governance anywhere near, because it was working, clinicians were grateful, and the organization had spent a decade making documentation worse. It was classified Tier 2 rather than Tier 1 on the reasoning that it does not make a clinical recommendation — it transcribes and organizes, and a licensed human signs every output.

That reasoning was correct about what the system does and wrong about what the signature means.

The signal

An urgent care physician was reviewing a draft note for a patient who had come in with an ankle injury. The review of systems read: “Patient denies chest pain, shortness of breath, or palpitations.” She had never asked. The conversation, all six minutes of it, was about an ankle. The audio confirmed it — no one said any of those words.

This was not a transcription error. The model had produced a fluent, clinically plausible, conventionally formatted pertinent-negative statement because that is what a review of systems looks like in its training data. The sentence was grammatical, appropriate to the setting, and written in the physician's own documentation voice. It was also entirely invented.

The committee ordered an audit: 500 randomly sampled signed notes reconciled against retained audio. The result:

21notes of 500 (4.2%) contained at least one clinical assertion unsupported by the audio
8fabricated pertinent negatives — symptoms “denied” that were never discussed
7physical exam findings the clinician had not performed or verbalized
6history attributed to the patient that a family member in the room had described
21of the 21 had been read, attested and signed by a licensed clinician

The last number is the finding. Every fabrication had passed through the control that was supposed to catch it. Median time from note presentation to signature was 11 seconds. The human in the loop was in the loop and was not reading.

Two further problems surfaced once the committee had the contract in front of it. The vendor's default configuration retained encounter audio indefinitely — a growing archive of recorded clinical conversations, PHI in the most sensitive form the organization holds, with no retention limit anyone had agreed to. And the standard terms granted the vendor rights to use customer audio for model improvement. Nobody at Meridian had ever decided that patient conversations would be used to train a commercial product. Somebody had signed a contract that said so.

The three options on the table
Option A
Suspend it. A system that invents clinical findings in the legal medical record cannot run while you think about it.

610 clinicians lose 47 minutes a day. The measured improvement in documentation burden reverses overnight, and governance becomes the function that took away the one thing that helped.

It also solves nothing structural. Attestation-as-a-click is a problem for every draft-generating system Meridian will ever deploy, and suspending this one leaves that intact.

Rejected — proportionate to the alarm, not the risk
Option B
Keep it. Add a warning banner above every draft and retrain all 610 clinicians on their attestation obligation.

This is the intervention organizations reach for because it is fast, cheap, documentable and feels responsible. It generates a training completion rate to report.

It relies on sustained vigilance against a failure mode engineered to be invisible. The fabricated text is fluent, plausible and in the reader's own voice. Banners habituate in about a week.

Rejected — a control that depends on nobody getting tired
Option C — chosen
Keep it, and change what the model is permitted to write. The note may assert only what was said; anything inferred goes in a block that cannot be signed until it is resolved.

Generated pertinent negatives and unverbalized exam findings suppressed entirely at the vendor level. Speaker attribution required before any history is recorded as the patient's.

Median attestation time moved from 11 seconds to 48. Clinicians still overwhelmingly preferred it to typing — adoption fell by under 3%.

Chosen
What the committee decided, and why

The committee's insight was that the defect was not in the model's accuracy but in the shape of its output. A draft that mixes what was heard with what was inferred, in identical formatting, makes verification cognitively impossible at speed — and clinical documentation always happens at speed. The fix was to make the two categories visually and structurally different, and to make the second one block the signature until a human touched each item.

The one-action override principle from Step 8 was deliberately inverted here. Everywhere else, the committee insists that overriding the AI must take one action and never a form, because friction on the override is friction on the human's judgment. Here the friction is the control: accepting an inference the model made up should be harder than deleting it.

  • Audio retention set to 30 days and then destruction, written into the amended BAA rather than left as a configuration setting a vendor can change.
  • Secondary use for model training prohibited outright. The vendor agreed; the clause had been boilerplate nobody had read.
  • Patient notice at registration and verbally at the start of the encounter, with opt-out honoured without any change to how the visit proceeds — and no requirement that the patient explain why.
  • Reclassified from Tier 2 to Tier 1. A system that writes the legal medical record shapes every downstream decision made by anyone who reads it.
Artifact this produced The attestation design standard, which now applies to every system that drafts something a human signs: generated content must be distinguishable from captured content; inferred content must require an explicit per-item action; and the median time-to-signature is monitored as a control-effectiveness metric. If attestation time collapses, the control has failed, whether or not anything has gone wrong yet.
Transferable lesson

“A licensed human reviews every output” is the most common control in healthcare AI governance and the most frequently untrue. A signature box is not a control; it is a place to record that a control was supposed to happen. If the interface makes it possible to attest without reading, the interface has decided the outcome and your policy is documentation of an intention.

Discussion questions
  1. Where in your organization does a human “review” an AI output? Measure how long that review actually takes. What would you do if the answer is eleven seconds?
  2. Scribe was classified Tier 2 because it makes no recommendation. Is writing the record that everyone else's decisions are based on a lower risk than making one recommendation, or a higher one?
  3. Nobody decided that patient conversations could train a vendor's model — somebody signed it. Which of your existing AI contracts have you actually read the secondary-use clause in?
04
UM Assist · prior authorization pre-scoring · Tier 1
The word “assists” is doing a lot of work
System
Pre-scores incoming prior authorization requests in the health plan and sorts them into likely meets criteria and unlikely to meet criteria before a nurse reviewer opens the case. Roughly 1,900 requests a week.
Stated design
Decision support only. A licensed human makes every determination. The vendor's documentation says the model “assists” reviewers.
Trigger
Step 9 monitoring — the appeal overturn rate crossed a threshold nobody had set until governance set one.
Regulatory hooks
42 CFR 422.101(b)(6) and CMS's February 2024 guidance;3 state utilization review law restricting medical necessity determinations to licensed clinicians; Section 1557; ERISA and appeal rights.
The situation

Meridian sits on both sides of this table. It runs a 48,000-life Medicare Advantage plan that uses AI to triage authorization requests, and it is a provider whose own requests are triaged by other payers' algorithms. The committee found the second fact clarifying. Several of the people arguing that UM Assist was obviously fine had spent the previous year complaining, at length, about exactly the same technology pointed at them.

The compliance position was considered settled. CMS is explicit that an algorithm may be used to assist in coverage decisions but cannot serve as the sole basis to deny or terminate coverage, and that a determination must rest on the individual patient's circumstances rather than on what a larger dataset says about people like them. UM Assist never issued a denial. A nurse reviewer did, and a medical director signed the ones that required it. Sole basis, satisfied.

The signal

The overturn rate. Of the denials that patients or providers appealed, 58% were being overturned — up from 34% in the eighteen months before UM Assist went live. An overturn is the plan telling itself, in writing, that its own determination was wrong.

The committee pulled the reviewer timestamps and found the mechanism:

Reviewer behaviourCase flagged “likely meets criteria”Case flagged “unlikely to meet criteria”
Median review time4 min3 min
Median review time, cases with no flag shown (control period)14 min
Reviewer disagreed with the flag9% of cases6% of cases
Additional clinical records requested before deciding11%8%
Denials later overturned on appeal58%

Three minutes. On a case the model had already labelled unlikely to meet criteria, a nurse reviewer spent three minutes and agreed 94% of the time. The model was not assisting the determination. It was making the determination and the reviewer was ratifying it — which is precisely the arrangement the rule exists to prohibit, arrived at without anyone deciding to do it.

The distinction between assists and decides is not settled by the policy document, the vendor's product description or the org chart. It is settled by the timestamp log. Meridian's said decides.

The three options on the table
Option A
Keep it as configured. A licensed human signs every determination, which is what the regulation requires.

Technically defensible right up to the moment someone requests the timestamps. The three-minute median and the 58% overturn rate are already in the plan's own systems and would be produced in any audit or litigation.

The overturn rate is also a patient harm measure. Every overturned denial is a patient who was told no, and who only got to yes by having the resources and persistence to appeal.

Rejected — the evidence contradicts the policy
Option B
Remove UM Assist entirely and return to unassisted human review.

Review capacity drops roughly 40% against fixed regulatory turnaround deadlines. Turnaround failures are their own patient harm — a delayed approval is a delayed treatment.

It also discards something that genuinely works. On the approval side the model is accurate, and faster approvals hurt nobody.

Rejected — throws away the half that helps
Option C — chosen
Make the automation asymmetric. The model may accelerate approvals. It may never appear anywhere on the denial path.

The “unlikely to meet criteria” flag was removed from the reviewer interface entirely — not de-emphasised, not accompanied by a caution, removed. A reviewer working a case that may end in denial sees no model output at all.

Review capacity is preserved, because most requests are approvals. Median review time on potential-denial cases returned to 15 minutes. Overturn rate fell to 31% over the following two quarters.

Chosen
What the committee decided, and why

The argument that won was about the asymmetry of the errors. A wrong approval costs the plan money. A wrong denial costs a patient their treatment, and lands hardest on the patients least equipped to appeal. When the consequences of the two error types are that different, the automation should be that different too. Meridian's rule became: a model may make the benign outcome faster; it may not make the harmful outcome easier.

The committee also rejected the framing that removing the flag was a loss of functionality. Anchoring is not a bug in the reviewer. It is how human judgment works, reliably, in every profession that has been measured. Designing a workflow that shows a busy clinician a confident-looking verdict and then asks for independent judgment is asking for something people cannot do.

  • Every denial must document the individual patient circumstances considered, by name, in a free-text field that cannot be templated — the individualised-determination requirement made operational rather than asserted.
  • Medical necessity determinations restricted in policy to licensed clinicians reviewing the individual record, mirroring the standard several states have now written into statute and the design CMS used in its own WISeR prior-authorization model, where every non-affirmation recommendation must come from an appropriately licensed clinician.
  • Appeal overturn rate promoted to a Sev-2 monitored metric with a defined threshold, reported quarterly to the board quality committee alongside clinical quality measures rather than buried in a plan operations report.
  • The same asymmetry test applied retroactively to every other system in the register, which is how the committee found two more places it applied.
Artifact this produced The asymmetric automation test, now applied at intake to every system: “What does a false positive cost, what does a false negative cost, and who pays each one?” Where the answers differ materially, the deployment must differ correspondingly — and where the cost falls on the patient rather than the organization, the burden of proof sits with the proposal.
Transferable lesson

“The AI only assists” is a claim about behaviour, and behaviour is measurable. If your reviewers spend three minutes on flagged cases and fourteen on unflagged ones, the model is deciding and your policy is describing an organization you do not have. Instrument the humans, not just the model.

Discussion questions
  1. Pick a system where your policy says a human decides. What measurement would falsify that claim, and do you currently collect it?
  2. Meridian removed the flag rather than adding a caution next to it. When is removing information from a human the more responsible design?
  3. The overturn rate was visible in the plan's own reporting for eighteen months. Why does a number that exists not count as a signal until somebody owns it?
05
PortalDraft · generative replies to patient messages · Tier 2
The model changed over a weekend and nobody told anyone
System
Drafts replies to patient portal messages for clinician review and sending. Roughly 4,100 messages a day across primary care and specialty clinics.
What changed
The vendor upgraded the underlying foundation model on a Saturday night. Routine, from their perspective. Covered by their service terms.
How Meridian found out
A patient telephoned to complain that the reply she received “wasn't from a person.”
Trigger
Patient complaint. Not vendor notice, not monitoring, not the change advisory board. A complaint.
Regulatory hooks
State generative-AI disclosure law for patient communications; medical record integrity; contractual change control; EU AI Act Article 50 for the institute's European-facing systems.
The situation

PortalDraft had passed intake cleanly. It was validated, the drafts were good, clinician satisfaction was high, and the human-review control looked solid: no message reaches a patient without a clinician sending it. It was the least controversial system in the register.

It was also the system whose validation had the shortest half-life, and nobody had noticed because nothing in the process asked how long a validation stays true.

The signal

The complaint was vague and the patient was right. Reconciling the week before and the week after the weekend upgrade:

Output characteristicBeforeAfterMeridian's standard
Reading level (Flesch–Kincaid grade)8.212.66th–8th grade
Median draft length94 words212 words
Drafts appending a generic “consult your physician” closing3%76%
Drafts declining to address the question directly2%14%
Drafts sent with no clinician edit61%61%

The last row is the one that should worry you. The system changed materially and the human review layer did not register it at all — the same 61% went out untouched, now at a twelfth-grade reading level, telling patients who had just messaged their physician to consult their physician.

Meridian's validation had been performed against a model that, by Monday morning, no longer existed. The contract said nothing about notice, because nobody had thought to ask for it. To the vendor this was a version bump. To Meridian it was an unvalidated system in production, communicating with patients, for eleven days before anyone looked.

The disclosure question came up in the same review. California's AB 3030 requires generative-AI-produced clinical patient communications to carry a disclaimer and instructions for reaching a human — at the outset of a written message and persistently through a continuous online interaction — and exempts communications that a licensed human provider reads and reviews first. Meridian operates in neither California nor any state with an equivalent statute in force. The committee adopted the standard anyway, for two reasons: running two message pipelines with different disclosure rules is an error factory, and the exemption itself is the uncomfortable part. It turns on a human genuinely reading the message. Meridian had just established, at 61% sent unedited, that it could not evidence that for the majority of its traffic.

The three options on the table
Option A
Accept the new model and update the internal patient communication standard to match its output.

Fastest path, and it was seriously proposed — the new model was in most respects more capable.

It also concedes that the vendor's release schedule sets Meridian's patient communication policy. Health literacy standards exist because a twelfth-grade reply to a patient with an eighth-grade reading level is a failed communication regardless of how sophisticated it is.

Rejected — lets the vendor write the policy
Option B
Contractually pin the model version. Nothing changes without our written approval, ever.

Superficially the strongest control and the least sustainable. Vendors do not maintain arbitrary old versions indefinitely; you eventually sit on something unpatched and unsupported.

It also blocks improvements you want, including safety fixes, and creates pressure to grant blanket exceptions that quietly restore the original problem.

Rejected — unsustainable, and it decays into Option A
Option C — chosen
Notice plus a validation window in the contract, and output monitoring that does not depend on the vendor telling you anything.

Thirty days' written notice of any change to the model, its version, its training data or its system prompt; a fifteen-business-day validation window before the change reaches production; and a right to reject.

Crucially, an automated output monitor — reading level, length, refusal rate, disclaimer presence — running daily, because the contract only protects you against vendors who comply with it.

Chosen
What the committee decided, and why

Both halves of Option C were treated as load-bearing, and the committee was explicit that the monitor mattered more than the clause. A contract is a remedy after the fact. Monitoring is how you find out. Meridian had a patient complaint as its detection mechanism, which means the detection mechanism was a patient being poorly served and caring enough to telephone about it.

The reclassification argument was harder. PortalDraft was Tier 2 because a clinician sends every message. Having watched that control absorb a complete model swap without noticing, the committee moved it to Tier 1 — not because the system got riskier, but because the control it was relying on had been measured and found weaker than assumed.

  • Every AI-drafted patient message carries a disclosure line and instructions for reaching a human, in every state, regardless of local requirement.
  • Reading level enforced at generation rather than checked afterwards; drafts above grade 8 are regenerated before a clinician ever sees them.
  • The change notice clause was retrofitted into every AI vendor contract at renewal, and became non-negotiable for new procurements. Two vendors refused; one of those was not renewed.
  • Output drift monitoring extended to every generative system in the register, including the ones with no known problem.
Artifact this produced The model change notice clause, and the output drift monitor that assumes the clause will fail. Together they answer a question the original intake process never asked: “How would we know if this system stopped being the system we validated?” If the answer is “the vendor would tell us,” there is no answer.
Transferable lesson

A validated system is a validated version. Generative systems can be swapped underneath you between a Friday and a Monday, and the swap is invisible to every control that depends on a human noticing that the text reads differently. Monitor the outputs, not the release notes — and treat the day you cannot detect a version change as the day your validation expired.

Discussion questions
  1. For each generative system you run, what would you measure daily to detect that the model changed? If you have no such measure, how many days could it run unvalidated before someone complained?
  2. Meridian adopted the strictest state disclosure standard everywhere. When is voluntarily exceeding the requirement cheaper than tracking which requirement applies where?
  3. Two vendors refused the change notice clause. What is your organization's actual willingness to walk away from an AI vendor, and who has the authority to do it?
06
NoduleAI · lung nodule detection on chest CT · Tier 1
The same organization is a deployer, a provider, and did not know it
System
FDA-cleared, CE-marked computer-aided detection for pulmonary nodules on chest CT, running in radiology at all four hospitals.
The complication
The research institute is co-developing the next version with the manufacturer, using Meridian data plus data from a European academic hospital, and intends to deploy the modified model at the European site.
Trigger
Step 4 intake on the research protocol — a lawyer asked which role Meridian occupies in Europe, and nobody could answer.
Regulatory hooks
FDA software as a medical device and predetermined change control plans; EU AI Act provider versus deployer obligations, Annex I routing, Article 50 transparency, Article 4 AI literacy; GDPR and cross-border transfer; the EU Medical Device Regulation.
The situation

NoduleAI was the least alarming system in the register. Regulated device, cleared by FDA, CE-marked, radiologist reads every study, well-understood technology with a decade of literature. It sailed through intake in the commercial deployment.

The research protocol was a different document, and it arrived at the committee only because Step 4 requires every new AI use to come through intake, including research uses that a research ethics committee has already approved. That rule felt bureaucratic when it was written. It is the reason this was caught.

The signal

The lawyer's question was simple and nobody in the room could answer it: in Europe, is Meridian a deployer or a provider?

The answer turned out to be both, in different capacities, with different obligations, on different timelines:

CapacityRoleRouteDate that governsWhat it requires
Clinical use at Meridian's US hospitalsDeployerFDA / US lawNowLocal validation, monitoring, human oversight, HTI-1 source attributes from the EHR developer
Clinical use at the European partner siteDeployerEU AI Act, Annex I — safety component of a regulated device2 August 20284Use per instructions, human oversight by competent staff, input data relevance, log retention, incident reporting
The modified model built by the collaborationProviderEU AI Act, Annex I2 August 2028Conformity assessment, risk management system, technical documentation, data governance, post-market monitoring — the full provider stack
The institute's EU-facing trial participant assistantDeployerEU AI Act, Article 50 transparency2 August 2026Tell people they are interacting with an AI system

Two findings came out of that table. The first is the one the committee went looking for: substantially modifying a high-risk system, or putting your name on it, makes you its provider. The research institute had been thinking of itself as a data contributor to someone else's product. Under the EU framework it was building a new high-risk medical device system and planning to put it into service.

The second finding was an accident. While confirming that the 2028 date was right, the committee checked whether anything else in the register touched the EU regime and found the trial participant assistant — a generative chatbot answering questions from European trial participants, stood up by the research institute nine months earlier, never entered in the register, and squarely inside Article 50's transparency obligation. Article 50's timeline was not moved by the Digital Omnibus amendments that pushed the high-risk dates back. It applies from 2 August 2026.

The committee reviewed this on 24 July 2026. The obligation was nine days away, on a system nobody had known existed, discovered while researching a deadline two years out.

The question that should be asked of every regulated AI vendor

The committee also asked the manufacturer whether NoduleAI had a predetermined change control plan, and what it permitted. It did: retraining on additional data, threshold recalibration, and performance improvements within the cleared indication — all without a new marketing submission. That is exactly what a PCCP is designed to allow, and it is entirely legitimate. It also means the model in Meridian's radiology department can change materially, lawfully, and with no regulatory event to notice. The only notice Meridian gets is whatever the contract requires, and the contract required nothing. Scenario 5's clause was retrofitted here the same week.

The three options on the table
Option A
Treat it as the manufacturer's problem. It is their device, their CE mark, their conformity assessment. We are a hospital.

Correct for the commercial deployment and wrong for the collaboration. The provider obligation attaches to substantial modification and to putting the system into service under your own name, not to who wrote the original code.

It would also have left the Article 50 exposure entirely undiscovered, since nobody would have looked.

Rejected — right answer to the wrong question
Option B
Withdraw from the European collaboration. The provider obligations are disproportionate to a research partnership.

Ends a genuinely valuable program to avoid an obligation that is demanding but well-defined, with a two-year runway to meet it.

This is governance as an avoidance function, and it is how governance loses the standing to be consulted early. Say no to the hard thing once and the next protocol will not come to intake at all.

Rejected — buys compliance with credibility
Option C — chosen
Separate the roles explicitly, staff each one, and put a hard gate before the modified model reaches any patient.

Deployer obligations assigned to clinical operations; provider obligations assigned to the research institute with named accountability and a budget, because provider obligations are a programme, not a checklist.

The modified model may be developed and evaluated. It may not be put into service at any site until the provider-side documentation is complete and the committee has approved it as a separate decision.

Chosen
What the committee decided, and why

The committee treated role clarity as the deliverable. Most of the confusion in cross-border AI governance comes from organizations that assume they occupy one role for all purposes, when in practice a health system is a deployer of most things, a provider of a few, and occasionally both for the same underlying model in different capacities. Writing the roles into a table, capacity by capacity, resolved arguments that had been circling for weeks.

The collaboration was redesigned to train the modified model by federated learning — the model travels to each site, trains locally, and only parameter updates are exchanged. No European patient data leaves its jurisdiction, which simplifies the transfer analysis considerably and was, the institute's own investigators pointed out, better science as well, since it allowed a wider partner network than any data-sharing agreement would have supported.

  • The trial participant assistant was brought into the register, given a disclosure notice at the start of every conversation, and assigned an owner — inside the nine days.
  • Article 4 AI literacy treated as a legal obligation rather than a training nicety: staff who operate or oversee an AI system in an EU context must be demonstrably competent to do so, and that competence has to be evidenced.
  • A standing horizon-scanning item added to the quarterly agenda, owned by legal, because the committee's near-miss was caused by tracking one date and not the others.
  • Every research protocol involving AI now routes through both the research ethics committee and AI governance. Neither substitutes for the other; they are asking different questions.
Artifact this produced The role determination worksheet — run per system, per jurisdiction, per capacity — and a regulatory calendar that holds every applicable date in one place with an owner against each. Meridian's near-miss was not caused by ignorance of the EU AI Act. It was caused by knowing one date confidently and never checking whether the others had moved differently.
Transferable lesson

Provider and deployer are capacities, not identities. The same organization can hold both for the same model, and modifying a system or putting your name on it moves you across the line whether or not you intended to become a manufacturer. And the far-off deadline is rarely the dangerous one — the dangerous one is the near date on the system nobody entered in the register.

Discussion questions
  1. List the jurisdictions your organization touches, including through research collaborations, telehealth and patients who travel. For each, which role do you occupy for which systems?
  2. The trial participant assistant existed for nine months without governance knowing. What is your equivalent, and which of Step 2's five discovery channels would have surfaced it?
  3. Option B is the defensible-sounding refusal. What does it cost a governance function, over two years, to be the group that says no to the interesting work?
Running these as an exercise

Each scenario is written so it can be handed to a group with the decision removed. Give them the situation, the signal and the three options, ask for a decision and the reasoning behind it, then reveal what Meridian chose. The disagreements are more instructive than the answers — particularly on Scenario 4, where the option most groups pick first is the one the timestamps refute, and on Scenario 1, where the technically-minded reach for the threshold adjustment almost every time.

The map

The regulatory map

Healthcare AI does not have one law. It has a stack of regimes written at different times for different purposes, none of which were designed with each other in mind, and most of which predate the technology they now govern. A single system routinely sits inside four or five of them at once. Sentinel is simultaneously a clinical decision support intervention under HTI-1, a quality and safety matter under Joint Commission expectations, a potential Section 1557 concern because of the safety-net subgroup gap, and a patient safety reporting obligation. None of those regimes will tell you about the others.

What follows is the map Meridian built. It is organised by what triggers each regime rather than by which agency issued it, because the trigger is the useful part when a new system arrives at intake.

RegimeWhat pulls you inWhat it makes you doDate
ACA Section 1557
patient care decision support tools
Any automated or non-automated tool used to support clinical decision-making. Scheduling, supply chain and staffing tools are outside it; anything shaping what care a patient is offered is inside. Reasonable efforts to identify tools using race, color, national origin, sex, age or disability as input variables, and policies to mitigate the resulting discrimination risk. Enforcement weighs entity size and resources, whether the tool was used as intended, how it was customised, developer guidance, and whether you have a governance process at all. Compliance date 1 May 2025
ONC / ASTP HTI-1
decision support interventions
Certified health IT that supplies decision support. Falls on the EHR developer, but the hospital is the party that needs the information. Predictive DSI must ship 31 source attributes across 9 categories — intervention details and output, purpose, out-of-scope cautions, development and input features, fairness assurance, external validation, performance metrics, ongoing maintenance, and the update and revalidation schedule. Evidence-based DSI requires 13. Plus intervention risk management: analysis, mitigation, governance. Developer delivery 31 Dec 2024; in effect 1 January 2025
FDA
software as a medical device
Software intended to diagnose, treat, prevent or mitigate disease. Clinical decision software that a clinician cannot independently review the basis of is generally in scope; the marketing language is often the tell. Clearance or approval for the device itself. For the deployer, the operative question is the predetermined change control plan: what has the manufacturer pre-authorised itself to change — retraining, recalibration, performance improvements — without a new marketing submission? In force
HIPAA Any use or disclosure of protected health information, including audio, images and anything handed to a vendor's model. Business associate agreements with AI-specific terms: no secondary use for model training, defined retention and destruction, subcontractor and sub-processor disclosure, breach notification that covers model outputs. De-identification via Safe Harbor or Expert Determination, with Expert Determination the realistic route for rich clinical data. In force
CMS Medicare Advantage
42 CFR 422.101(b)(6)
Algorithms used in coverage, utilization management or medical necessity determinations by a Medicare Advantage organization. An algorithm may assist but cannot be the sole basis to deny, discontinue or terminate coverage. The determination must rest on the individual patient's circumstances, not on what a larger dataset predicts about similar patients. MAOs must also not use algorithms in a way that violates Section 1557. FAQ guidance 6 February 2024
CMS WISeR model Original Medicare fee-for-service prior authorization in six states — New Jersey, Ohio, Oklahoma, Texas, Arizona and Washington — for a defined service list including skin and tissue substitutes, electrical nerve stimulator implantation, and knee arthroscopy for osteoarthritis. Inpatient-only, emergency, and services where delay creates risk are excluded. AI or machine learning may be used, but every non-affirmation recommendation must be made by an appropriately licensed clinician. Participation is voluntary. Worth reading even if you are outside the six states: it is CMS designing a compliant AI-assisted review process, which makes it a useful reference standard. 1 January 2026 – 31 December 2031
EU AI Act
Annex III — standalone high-risk
Standalone high-risk systems listed in Annex III. For healthcare organizations this typically catches employment, access to essential services and creditworthiness uses rather than clinical devices. Full high-risk obligations, allocated by whether you are the provider or the deployer. Moved to 2 December 20274
EU AI Act
Annex I — embedded / safety component
AI that is a safety component of a product already covered by Union harmonisation legislation. Medical devices route here, which is where most clinical AI lands. Providers: conformity assessment, risk management system, technical documentation, data governance, post-market monitoring. Deployers: use per instructions, human oversight by competent staff, input data relevance, log retention, incident reporting. Moved to 2 August 20284
EU AI Act
Article 50 transparency & Article 4 AI literacy
Systems that interact directly with people, or generate synthetic content, where the output is used in the Union. Article 4 applies to anyone operating or overseeing an AI system. Tell people they are interacting with an AI system; mark synthetic content. Ensure staff have a sufficient level of AI literacy — a legal obligation, not a training aspiration, and one you have to be able to evidence. 2 August 2026 — not moved by the Digital Omnibus amendments
Joint Commission & CHAI
Responsible Use of AI in Healthcare
Accredited organizations. Guidance rather than a standard, with a voluntary certification following it. Seven pillars: AI policies and governance structures; patient privacy and transparency; data security and data use protections; ongoing quality monitoring; voluntary blinded reporting of AI safety events; risk and bias assessment; and education and training. Guidance 17 September 2025; certification 2026
State law Varies enormously and changes fast. California AB 3030 (generative AI in clinical patient communications must disclose and give human contact instructions, unless a licensed provider reads and reviews first) and SB 1120 (AI may not determine medical necessity; only a licensed physician or qualified professional may, using the individual's data) are the pattern others are copying. Colorado's original AI Act is being replaced by legislation applying to consequential decisions from 1 January 2027. Disclosure, human review, individualised determination, documentation retention, adverse-outcome notification. The specifics differ by state and by year. Volatile — verify before relying
Treat state law as the moving part

Federal regimes change on multi-year cycles with notice-and-comment in between. State AI law in 2025 and 2026 has changed on a scale of months, including statutes that took effect and were then paused by litigation or repealed and replaced before enforcement began. Build your load-bearing controls on the federal and accreditation requirements, which are stable, and treat state obligations as a layer you re-check quarterly. Do not architect a workflow around a single state statute without an owner assigned to watch it.

One further point about the map, which Meridian's committee reached the hard way in Scenario 6: the dangerous date is almost never the famous one. Everybody in healthcare AI governance knows the EU AI Act high-risk dates. Far fewer noticed that the transparency obligations were carved out of the delay and kept their original timeline, or that the obligation attaches to a chatbot the research institute stood up without telling anyone. Track every applicable date in one calendar, with a named owner against each, and re-check the whole calendar rather than the entry you were already worried about.

Sequence

The 12-month rollout

Meridian did not do the ten steps in order and neither will you. Steps 1 to 3 have to come first because everything else depends on them, but after that the sequence is driven by what the inventory found, and the inventory will find something that cannot wait. What follows is roughly how the first year actually went, including the part where the plan changed in month four.

Months 0–1 · Mandate
Charter the committee and give it real powers

Board quality committee charters the function, not IT. Three powers written down explicitly: require validation before deployment, suspend a live system, escalate to the board directly. CMO chairs. CNO on it, because clinical AI reaches nurses first. Internal audit deliberately excluded, to preserve an independent third line.

  • Deliverable: signed charter, named members, meeting cadence, escalation path
  • The test: can this group suspend a live system this week without asking anyone? If not, you have an advisory board.
Months 1–3 · Discovery
Build the inventory, and expect the number to be wrong

Five discovery channels run in parallel: EHR configuration review, contract database search, network and SaaS discovery, departmental interviews, and a CEO-signed 30-day amnesty for shadow AI. Leadership estimated three or four systems. The sweep found eleven.

  • Never ask “do you use AI?” — ask “what software makes a suggestion or a prediction that changes what someone does?”
  • Deliverable: the register, with an owner named against every entry. An entry with no owner is not an entry.
Months 2–4 · Triage
Classify, then validate the worst one first

Risk tiers assigned on consequence times independent human judgment, plus the five-question regulatory routing. Six of the eleven systems matter. Sentinel goes first — highest tier, longest running, most patients, no validation ever performed.

  • Resist the temptation to start with the easy one. Starting with the hardest system is what establishes that the gate is real.
Month 4 · The plan changes
Scenario 1 lands and consumes the quarter

Sentinel's validation comes back at AUC 0.64 with a safety-net subgroup gap. The system is suspended house-wide within nine days. Everything else slips a month, the committee earns its credibility in a single decision, and intake volume triples the following quarter because people have now seen that the process does something.

Months 4–7 · The gate
Stand up intake, contracts and validation as routine

Intake form, time-boxed review (15 business days Tier 1, 10 days Tier 2), BAA addendum with the evidence clause and the change notice clause, and the silent-period protocol. CareLens equity assessment runs here and produces Scenario 2.

  • The time box matters more than the checklist. A gate with no deadline becomes the thing people route around.
Months 6–9 · Deployment design
Make human oversight specific, and train three audiences

Human-in-the-loop specified concretely per system: who, what information, at what time, with what authority, and what happens on disagreement. Override is one action, never a form. Patient disclosure standardised. Training runs at three depths — 20 minutes for all staff, 90 minutes for clinicians, half a day for system owners. Scenario 3 surfaces here and rewrites the attestation standard.

Months 8–12 · Monitoring
Turn on the part nobody budgets for

Six monitoring dimensions per Tier 1 system: input drift, output drift, performance, subgroup performance, human interaction, operational health. Four-severity incident ladder with defined thresholds and on-call ownership. Scenarios 4 and 5 are both detected by monitoring rather than by complaint — which is the entire point of building it.

  • Sev-4 alert-burden complaints are the leading indicator for Sev-1 patient harm. Treat the annoyance reports as data, not noise.
Months 10–12 · Change control
Close the loop that makes validation durable

Version changes route to the change advisory board. Every vendor is asked whether there is a predetermined change control plan and what it permits them to change without telling you. Tier 1 revalidates annually, Tier 2 every two years. Decommission procedures written — the most under-designed step in almost every programme. Scenario 6 lands here and adds the jurisdictional role analysis.

Month 13 onward
What steady state looks like

Intake is routine and mostly uncontroversial. The register is current because three intake triggers keep it current, not because someone runs an annual sweep. Monitoring catches things before patients do. The committee spends most of its time on genuinely hard cases rather than on discovering systems it did not know existed. And the second inventory sweep, run twelve months after the first, finds two more systems — which is a healthy number, not a failure.

What you actually build

Artifacts and templates

A governance programme is, concretely, about a dozen documents that exist and get used. Everything else is meetings. This is the set Meridian ended the year with, mapped to the step that produced it and the person who owns it.

ArtifactWhat is in itOwnerFrom
Committee charterMandate, membership, the three powers, escalation path, meeting cadence, quorum, conflict-of-interest rulesBoard quality committeeStep 1
AI registerEvery system, its owner, tier, intended use, out-of-scope uses, validation status, monitoring status, last review date, next revalidation dateGovernance leadStep 2
Risk tiering rubricConsequence × independent human judgment matrix, plus the five-question regulatory routingCommitteeStep 3
Intake formIntended use in one sentence, out-of-scope uses, the label question, the counterfactual question, the asymmetry question, data sources, affected populationsGovernance leadStep 4
BAA AI addendumNo secondary use for training, retention and destruction, sub-processor disclosure, the evidence clause, the model change notice clause, audit rightsLegal & privacyStep 5
Validation report templateSilent-period design, discrimination, calibration, sensitivity and specificity at the operating threshold, alert burden, number needed to evaluate, incremental value, subgroup performanceClinical informaticsStep 6
Equity assessmentThe four questions in order, starting with what the model is actually predicting; subgroup results; proxy analysis; mitigation and its evidenceHealth equity & clinical informaticsStep 7
Human oversight specificationPer system: who, what information, at what time, with what authority, and what happens on disagreementSystem ownerStep 8
Attestation design standardGenerated content distinguishable from captured; inferred content requires per-item action; time-to-signature monitored as control effectivenessClinical informaticsScenario 3
Monitoring planSix dimensions, thresholds, frequency, who reads it, what happens when a threshold tripsSystem ownerStep 9
Incident procedureFour severity levels, definitions, on-call ownership, suspension authority, patient notification triggers, reporting obligationsPatient safety & governance leadStep 9
Change control recordVersion, what changed, PCCP scope, revalidation performed, approval, rollback planChange advisory boardStep 10
Role determination worksheetPer system, per jurisdiction, per capacity: provider or deployer, and the obligations that followLegalScenario 6
Regulatory calendarEvery applicable date, in one place, with a named owner and a quarterly re-checkLegalScenario 6
Decommission procedureNotice to users, data disposition, record retention, what happens to outputs already embedded in the record, vendor offboardingSystem ownerStep 10
The evidence test

For each artifact, ask whether you could hand it to a regulator, an accreditor or a plaintiff's expert tomorrow with no preparation. If the honest answer is that it would need to be written up first, it does not exist — you have a practice and a memory of a practice, which is a different thing. Governance that cannot be evidenced is indistinguishable, from the outside, from governance that never happened.

Measurement

Knowing if it's working

Most AI governance metrics measure the governance function's activity rather than its effect. Number of systems reviewed, policies published, training completions — these tell you the committee met. They do not tell you whether anything is safer. The ones below are chosen because each has a plausible way of going wrong that you would want to know about.

Register coverage
Systems in the register divided by systems believed to exist. Measured properly by running a fresh discovery sweep and counting what it finds that you did not already have.
Bad sign: a sweep that finds nothing new. It usually means the sweep was weak, not that coverage is complete.
Time from intake to decision
Median business days, by tier, against the 15-day and 10-day commitments.
Bad sign: the median is fine and the tail is enormous. The tail is where people learn to route around you.
Systems rejected, modified or suspended
The count of times the gate changed an outcome. Zero over a year means the gate is a formality.
Bad sign: a high rejection rate with rising shadow deployments. You are not preventing use, only preventing visibility.
Silent-period completion rate
Share of Tier 1 systems that completed a full local validation before firing anything at a human.
Bad sign: a growing list of documented exceptions. Exceptions are how a standard becomes a suggestion.
Largest subgroup performance gap
Per Tier 1 system, the widest gap between any two defined subgroups on the primary performance measure, tracked over time.
Bad sign: the gap is only reported when someone asks. It should have a threshold and a trigger, set before you need them.
Alert burden and number needed to evaluate
Per alerting system: alerts per clinician per shift, and how many evaluations produce one true finding.
Bad sign: override rate climbing while alert volume is flat. The humans have stopped believing it and nobody has told you.
Time-to-signature on attested outputs
Median seconds between an AI-drafted output being presented and a human signing it.
Bad sign: anything that looks like eleven seconds. The control has failed whether or not harm has occurred yet.
Incident count and severity mix
Incidents by severity per quarter. Read the mix, not the total.
Bad sign: no Sev-3 or Sev-4 incidents at all. Minor incidents are the leading indicator; their absence means people are not reporting, not that nothing is happening.
Revalidation currency
Share of Tier 1 systems whose last validation is within twelve months, and Tier 2 within twenty-four.
Bad sign: currency is high but no vendor version change has ever been caught. You are revalidating on schedule and missing the changes in between.
Appeal, override and reversal rates
Wherever an AI-influenced decision can be contested: how often it is, and how often the contest succeeds.
Bad sign: a high overturn rate treated as an appeals-process statistic rather than as evidence that the original decisions were wrong.

Three of these — override rate, time-to-signature, and the Sev-4 complaint stream — are the leading indicators. They move before patients are harmed. Everything else on the list tells you about something that has already happened. If you can only build three monitors in the first year, build those, and put a name against each one so somebody is obliged to look.

Honest assessment

How this goes wrong

Meridian's programme worked, which is partly why it is worth being explicit about the ways it nearly did not. Every failure below is one the committee either walked into or came close to, and most of them look like success from the inside for a considerable period before they announce themselves.

The committee without power. The most common failure and the hardest to reverse. A group that can recommend but not require, chartered under IT rather than under the board, is an advisory body that will be consulted after the contract is signed. Meridian's nine-day Sentinel suspension was only possible because the authority to suspend had been written down before anyone needed it. Retrofit that authority during a crisis and you will spend the crisis negotiating instead of acting.

The inventory that decays. A one-time sweep produces a register that is accurate for about a quarter. Without intake triggers wired into procurement, contract renewal and EHR configuration change, the register becomes a historical document that everyone treats as current. Meridian's second-year sweep found two systems that had appeared since the first — a good outcome, because two is what a working process leaks, and eleven is what an absent one accumulates.

Validation theatre. A validation report that reports AUC and stops. It is the number vendors supply, the number that sounds most scientific, and the number least connected to whether deploying the system will help anyone. If your validation template does not force alert burden, number needed to evaluate, subgroup performance and incremental value, it will produce documents that satisfy an auditor and tell a clinician nothing.

Human-in-the-loop that is not. The single most over-claimed control in healthcare AI. “A licensed clinician reviews every output” is a statement about workflow design, and it is false by default in any workflow that presents a confident-looking output to a busy person under time pressure. Three minutes on a flagged prior authorization; eleven seconds on a drafted note. Neither was a policy failure. Both were interface decisions that determined the outcome long before any policy applied.

Monitoring nobody reads. Building the dashboard is the funded part; reading it every week is not. A monitoring plan without a named person, a cadence and a defined action when a threshold trips is a data collection exercise. Meridian's rule became that any monitor without a trigger and an owner gets switched off, on the reasoning that an unwatched monitor is worse than none — it creates the belief that someone is watching.

Equity assessment as an artifact rather than a practice. A bias assessment performed once at intake, filed, and never repeated. Populations shift, care patterns shift, and models drift; a fairness result from eighteen months ago describes a system that no longer exists. The CareLens disparity would not have been found by a document. It was found because a care manager noticed her panel looked wrong and somebody took her seriously enough to pull the numbers.

The pilot that never ends. Systems deployed as pilots to avoid the governance gate, then quietly becoming permanent. The defence is a hard sunset: a pilot without an expiry date and a named decision-maker at the end of it is a deployment with better branding.

Decommission never designed. Almost every programme has an intake process and almost none has an exit process. What happens to outputs already embedded in the medical record? Who tells the users? What is retained, for how long, and under whose retention schedule? Meridian wrote its decommission procedure in month eleven, having already suspended a system in month four without one.

Governance as an avoidance function. The subtlest failure. A committee that says no to hard things stops being asked about hard things, and the interesting work relocates to wherever the committee is not. Meridian's most important decision in Scenario 6 may have been declining to exit the European collaboration — not because the compliance analysis demanded it, but because a governance function that only ever subtracts capability has a finite lifespan.

One date in the calendar. Knowing the headline regulatory deadline confidently and never checking whether the others moved differently. Meridian was nine days from missing a live transparency obligation on a system it did not know it operated, and found it by accident while researching a date two years out.

Facilitation

Running this as a workshop

This case study is built to be taught, not only read. The scenarios have their decisions separable from their setups precisely so a group can be made to commit to an answer before seeing what Meridian chose — which is where the learning is. Two formats work.

Half day, roughly three hours. Thirty minutes on the organization and what the first inventory found, which is enough to establish that everyone in the room is probably underestimating their own count. Then three scenarios at forty minutes each: Scenario 1 for validation and incremental value, Scenario 2 for the target variable, and Scenario 4 for the gap between the policy and the timestamps. Close with the metrics section and ask each participant to name the one leading indicator they could stand up within a quarter.

Full day. All six scenarios at forty-five minutes each, with the ten-step build presented in the morning as the framework the scenarios test. Give groups the situation and the signal, hand them the three options with the consequences redacted, and have them write the consequences themselves before revealing them. The redacted-consequence version generates far better discussion than the complete one, because arguing about what an option would actually cost is the skill the job requires.

A few facilitation notes drawn from running it. Scenario 4 is the reliable one for a mixed clinical and administrative audience, because both groups arrive certain that a human signing the determination settles the question, and the timestamp table lands hard. Scenario 2 works best with a group that includes someone from analytics, who will usually be the first to see that the label is the defect and can then explain it to the room more persuasively than a facilitator can. Scenario 3 is the one that changes behaviour after people leave, because everyone present has signed something in eleven seconds.

The most productive question to hold in reserve, for the point where a group has settled comfortably on the right answer: what would have to be different about your organization for the option you just rejected to be correct? It is the question that turns a case study into a transferable method rather than a set of remembered verdicts.

A note on the fiction

Meridian Health System does not exist. Its size, structure and system portfolio are a composite, and its internal numbers — the 38,900 admissions, the 4.2% audit finding, the 58% overturn rate — are constructed to be realistic rather than reported. The regulatory regimes, the published research findings, and the failure modes are real, cited, and verifiable. When you use this with a group, be explicit about which is which. A case study whose facts are treated as citable when they are illustrative does more harm than good.

Grounding

Sources for the cited findings

The numbers attributed to published research, as distinct from Meridian's constructed internal figures, come from the following.

  1. External validation of a widely deployed proprietary sepsis prediction model. Wong et al., JAMA Internal Medicine, 2021. 27,697 patients across 38,455 hospitalizations at a single academic health system. AUC 0.63 (95% CI 0.62–0.64) against a vendor-reported 0.76–0.83; 33% sensitivity at the recommended threshold; alerts on 18% of hospitalizations; 67% of sepsis cases missed; roughly 109 alerts to identify one case developing within four hours.
  2. Dissecting racial bias in an algorithm used to manage the health of populations. Obermeyer, Powers, Vogeli & Mullainathan, Science, 2019. A commercial risk algorithm using healthcare cost as a proxy for healthcare need; at the same risk score, Black patients had 26.3% more chronic illnesses; correcting the label raised the share of Black patients receiving additional help from 17.7% to 46.5%. Algorithms of this class were applied to roughly 200 million people in the United States.
  3. CMS guidance on algorithms in Medicare Advantage coverage decisions. 42 CFR 422.101(b)(6) and the CMS frequently asked questions of 6 February 2024: an algorithm may assist but cannot be the sole basis for denying, discontinuing or terminating coverage; determinations must rest on the individual's circumstances rather than on a larger dataset; use must not violate Section 1557.
  4. EU AI Act Digital Omnibus amendments. Provisional political agreement 6 May 2026, confirmed by the Council 13 May 2026. Annex III standalone high-risk obligations moved from 2 August 2026 to 2 December 2027; Annex I embedded and safety-component obligations, which cover medical devices, moved from 2 August 2027 to 2 August 2028. Both are fixed dates rather than conditional triggers. Article 50 transparency obligations were not moved and apply from 2 August 2026.

Regulatory positions described elsewhere on this page — Section 1557 patient care decision support tools, ONC/ASTP HTI-1 decision support intervention source attributes, FDA software as a medical device and predetermined change control plans, the CMS WISeR model, the Joint Commission and CHAI guidance of 17 September 2025, and the California and Colorado statutes — are summarised from the primary rules, guidance documents and legislation as they stood in mid-2026. Verify current status before relying on any of it for a compliance decision; the state-level positions in particular have changed repeatedly.

Go deeper

Where to go next

kashifz.com