One health system, six AI systems already running, and no governance function. This walks the whole build — ten implementation steps, then six worked scenarios where the framework meets a decision that has to be made on a Tuesday.
Meridian Health System is fictional — a composite. Everything it collides with is real: Section 1557, HTI-1, FDA SaMD, HIPAA, CMS coverage rules, the EU AI Act, and the published research on models that failed in the field.
Governance design is not universal. The same six factors that decide how much process an organization can sustain — size and maturity, industry and regulatory exposure, products and use cases, objectives and risk tolerance — produce a very different answer for a hospital than for a marketing agency. So the profile comes first, because in healthcare almost every fact about the organization drags a specific legal obligation behind it.
Meridian takes Medicare and Medicaid, runs a certified EHR, employs clinicians, insures patients, and does research. Each of those is a separate regulatory doorway, and AI walks through all of them at once.
| Fact about Meridian | What it triggers |
|---|---|
| Receives federal financial assistance from HHS | It is a covered entity under Section 1557 of the ACA. Since 1 May 2025 it must make reasonable efforts to identify patient care decision support tools that use race, color, national origin, sex, age or disability as an input variable, and must have policies to mitigate the resulting discrimination risk. |
| Holds protected health information | HIPAA Privacy and Security Rules: minimum necessary, business associate agreements with every AI vendor touching PHI, and a defensible position on de-identification and secondary use for model training. |
| Runs a certified EHR | ONC/ASTP HTI-1 decision support intervention criteria. Predictive DSIs surfaced through the certified system carry 31 source attributes the developer must disclose, plus intervention risk management practices. Meridian is the consumer of that disclosure — and has to actually read it. |
| Deploys software that diagnoses or informs treatment | FDA jurisdiction over Software as a Medical Device. Some tools are cleared devices, some fall in the clinical decision support carve-out, and some are being used outside the indication they were cleared for — which is where the exposure sits. |
| Owns a Medicare Advantage plan | CMS coverage rules. Under 42 CFR 422.101(b)(6) and the February 2024 CMS FAQs, an algorithm may assist a coverage determination but cannot be the sole basis for denying or terminating care; the individual's circumstances must be evaluated. |
| Operates in two states, one with an active AI statute | State law — generative-AI disclosure in patient communications, limits on AI in utilization review, and algorithmic-discrimination duties. This layer changes fastest and needs a standing watch, not an annual review. |
| Co-develops an imaging model with an EU market path | EU AI Act. An AI system that is a safety component of a regulated medical device is high-risk by the Annex I route, regardless of Annex III. Meridian also has to work out whether it is a provider or a deployer — the obligations differ sharply. |
| Accredited hospital | Joint Commission and CHAI published the first Responsible Use of AI in Healthcare guidance on 17 September 2025 — seven elements, from governance structures through to voluntary blinded reporting of AI safety events — with a voluntary certification following in 2026. |
Nobody at Meridian has to memorize eight regimes. But somebody has to own the mapping — because a single AI tool can sit inside four of them simultaneously, and the four have different definitions of the same word. "Decision support" means one thing to ONC, another to FDA, and a third to the Office for Civil Rights.
The governance effort began, as these usually do, with a question from the board after a news story: "How much AI are we actually running?" Nobody could answer. The chief information officer's estimate was "three, maybe four systems." A six-week sweep — EHR configuration review, procurement contract search, vendor security questionnaires, departmental interviews and an amnesty window for shadow tools — found eleven. Six of them mattered.
| System | What it does | How it got here | Initial tier |
|---|---|---|---|
| Sentinel | Sepsis early-warning score in the EHR, firing to nursing | Switched on by the EHR vendor as part of a module upgrade — no local decision | Tier 1 · Critical |
| CareLens | Population risk score selecting patients for the care-management program | Bought by population health in 2021, never reviewed | Tier 1 · Critical |
| Scribe | Ambient documentation — listens to the visit, drafts the note | Clinician-led pilot in 3 clinics, spread to 41 by word of mouth | Tier 2 · High |
| UM Assist | Prioritizes and pre-scores prior-authorization requests in the health plan | Vendor module, configured by the plan's operations team | Tier 1 · Critical |
| PortalDraft | Generative AI drafting replies to patient portal messages | Enabled in a vendor release; on by default | Tier 2 · High |
| NoduleAI | Lung nodule detection on CT, co-developed with a device partner | Research collaboration heading toward commercial release | Tier 1 · Critical |
| Plus five lower-consequence systems: nurse-scheduling optimization, supply forecasting, a revenue-cycle coding assistant, a call-center routing model, and a chatbot on the marketing site. | |||
Nobody had decided. Four of the six arrived through a vendor release or a departmental purchase. There was no moment at which Meridian, as an organization, chose to deploy them — which means there was no moment at which anyone assessed them.
Nobody was monitoring. Not one of the six had a defined performance threshold, an owner accountable for it, or a documented review since go-live. Two had been running for over three years.
Nobody could produce evidence. Asked for the Section 1557 file on CareLens — what variables it uses, whether any of them proxy for a protected characteristic, what mitigation exists — the answer was that no such file existed. The compliance date had passed fourteen months earlier.
That last finding is the one that turns this from an IT project into a governance one. Meridian's problem was never that it lacked AI expertise. It was that AI had been entering the organization through six different doors, and none of those doors had a person standing at it.
Ten steps, in three phases. The first three stand the function up and are done once. The middle four are the gate every system passes through, and get repeated for every tool forever. The last three are the part almost everyone skips, and the part that actually determines whether a model is still safe two years after go-live.
Meridian ran phase A in ten weeks with a working group of eleven people, none of them full-time on it. That is roughly the floor. Compressing it further tends to produce a charter nobody follows.
Done once. The output is a committee with a mandate, a list of what you own, and a rule for how much scrutiny each item gets.
Effective AI governance is distributed but accountable. The failure mode at both extremes is well documented: appoint one overwhelmed owner and the queue becomes the bottleneck; declare that "everyone owns AI" and nobody does. Meridian's answer was a chartered committee with real decision rights and a named executive sponsor who could be fired for getting it wrong.
The AI Governance Committee was chartered by the board's quality committee — deliberately not by IT — and given three powers written into the charter: it can require validation before deployment, it can suspend a live system, and it can escalate to the board without going through the executive team. Without at least the second of those, the committee is an advisory body, and advisory bodies get routed around.
Around the committee, the standard three-lines structure: the deploying department owns the risk day to day (first line), compliance and the governance office set policy and challenge (second line), internal audit tests independently (third line). Internal audit does not sit on the committee — that would compromise the third line.
You cannot govern what you cannot enumerate, and in healthcare the enumeration is genuinely hard, because most AI does not arrive labelled as AI. It arrives as "the new module," "clinical decision support," "the risk score," or "a feature in the release notes."
The register is a living record, not a spreadsheet produced once. Each entry carries: system name, vendor, version, business owner, clinical owner, what decision it influences, populations affected, data it consumes, whether PHI leaves the environment, regulatory classification, risk tier, validation status, monitoring plan, last review date, and next review date.
An inventory decays unless something forces new entries into it. Meridian wired three triggers: no purchase order over $0 clears procurement without an AI screening question; no change request touching clinical decision support clears the change advisory board without a governance reference number; and every vendor release note is reviewed for new AI functionality before the upgrade window. The third one is the one that catches PortalDraft-style surprises.
Two different classifications, often confused. Risk tier decides how much internal scrutiny a system gets. Regulatory routing decides which external obligations attach. A tool can be low-tier internally and still be squarely inside Section 1557.
Meridian tiers on two axes — the consequence if the output is wrong, and how much independent human judgment stands between the output and the action. A model that suggests something to a clinician who is going to check it anyway is meaningfully different from a model whose output is executed.
| Tier | Test | Meridian examples | What it requires |
|---|---|---|---|
| Tier 1 | Wrong output can cause physical harm, or determines access to care or coverage | Sentinel, CareLens, UM Assist, NoduleAI | Full committee review, local validation, subgroup analysis, board reporting, quarterly monitoring |
| Tier 2 | Enters the medical record or reaches a patient, but a clinician reviews before it acts | Scribe, PortalDraft | Committee review, targeted validation, attestation controls, semi-annual review |
| Tier 3 | Operational; affects staff or workflow, not clinical decisions or coverage | Scheduling, supply forecasting, call routing | Delegated review by a subgroup, annual attestation |
| Tier 4 | No PHI, no decision influence | Marketing site chatbot | Register entry and a privacy check |
"A human reviews it" is only a mitigation if the human realistically can and does. A nurse receiving 400 alerts a shift is not a meaningful reviewer, and neither is a clinician signing 30 AI-drafted notes at the end of a clinic. Meridian's rule: you may only claim human review as a control if you can state the review rate, the time available per item, and the observed override rate. If you cannot measure it, you cannot claim it.
For each system, the committee answers five questions in order, and the answers determine the compliance file that has to exist:
Repeated for every tool, new or already running. Existing systems are back-fitted in tier order — Tier 1 first, regardless of how long they have been live.
The intake form is the single highest-leverage artifact in the whole programme, because it forces the requesting department to answer questions they have usually not asked. Meridian's is two pages and deliberately hard to complete without the vendor's help — which is the point, since it surfaces immediately whether the vendor will actually answer.
Review is time-boxed: 15 business days for Tier 1, 10 for Tier 2. A governance function that cannot commit to a turnaround gets bypassed, and a bypassed gate is worse than no gate because it creates a false record of control.
Three questions decide whether the deal is acceptable, and they are usually settled by people who never speak to each other — so the governance function's job here is mostly convening.
Any vendor processing PHI on Meridian's behalf needs a business associate agreement, and the BAA is where the AI-specific terms have to live because the standard template does not contemplate them. Meridian's additions: the vendor may not use Meridian's data to train or improve models serving other customers without a separate written agreement; the vendor may not attempt to re-identify de-identified data, and must contractually bind its own subprocessors to the same; data is returned or destroyed on termination with certification; and Meridian gets notice of model version changes before they ship. That last clause is the one that turns Scenario 5 from an incident into a scheduled review.
Minimum necessary is a HIPAA obligation, not a preference, and model builders default to maximalism. The discipline is to ask, feature by feature, what it is doing for performance — which usually shrinks the feature set and, as a side effect, makes the model easier to explain and less likely to be quietly encoding something you would rather it did not.
Secondary use of clinical data for model development sits at the intersection of HIPAA, the consent under which the data was collected, and institutional research policy. Meridian's route: de-identify to the Safe Harbor standard where possible, use Expert Determination where the required fields make Safe Harbor impossible, and route anything that cannot be de-identified through the IRB. For the multi-site research collaboration, federated learning resolved a hard constraint — partner hospitals could not share patient records, so the model trains locally at each site and only parameter updates are pooled. The model goes to the data; the raw records never move.
An evidence clause: the vendor will, on request and within a defined window, provide subgroup performance data, the training population's demographic composition, and the model's input variables — in a form Meridian can hand to a regulator. Without it, Meridian's Section 1557 obligation to make reasonable efforts to identify problematic input variables runs into a vendor that treats the feature list as a trade secret. Negotiate this before signature; afterwards there is no leverage.
This is the step that separates real healthcare AI governance from paperwork. Vendor performance is a hypothesis about your population, not a measurement of it. Models degrade when moved between institutions because the patients differ, the documentation practices differ, the coding differs, and the care pathways differ.
Meridian's rule for every Tier 1 clinical system: a silent period of no less than 90 days, during which the model runs on live data and its outputs are logged but never shown to anyone. Then compare against the outcome that actually happened.
Testing, evaluation, verification and validation. Verification asks whether the system meets its specification — did we build it right. Validation asks whether it meets the real need in its actual context of use — did we build the right thing, for these patients, on this ward. A tool can pass verification cleanly and fail validation completely, and in healthcare that gap is where patients get hurt.
In healthcare this is not an ethics exercise appended to the end. Since 1 May 2025 it is a compliance obligation with a named regulator, and the assessment is the evidence that the obligation was met.
Dropping race from the feature set does not remove race from the model — it removes your ability to see it. The proxies remain, and now the disparity is invisible. Mitigation is measurement plus a corrective: re-label to the outcome you actually care about, recalibrate by subgroup, adjust the threshold, restrict the intended use, or decline to deploy. All five are legitimate; pretending the variable was never there is not.
The output is a written assessment per Tier 1 and Tier 2 system, refreshed on the monitoring cycle, stating what was examined, what was found, what was done about it, and who signed. This document is the file that gets produced if anyone ever asks.
The steps that determine whether a model that was safe at go-live is still safe in year three. This is where most programmes are thinnest.
A validated model can still fail on the ward if the deployment is designed badly. Three things get designed deliberately.
"A clinician reviews the output" is not a control until you can say who, with what information, in what time, with what authority to disagree, and with what happens when they do. Meridian specifies all five per system, and requires that override be one action, never a form. If disagreeing with the model costs more effort than complying, you have built automation bias into the workflow and then called it human oversight.
Disclosure obligations vary by state and by use, and are tightening. Generative AI drafting clinical communications to patients attracts an explicit disclaimer requirement in some states unless a licensed human reviews the message before it goes — which makes the review step a compliance control, not just a quality one. Meridian's policy is a floor above the strictest state it operates in, applied everywhere, because two policies for two states is a policy nobody follows.
AI literacy has moved from good practice to legal obligation — the EU AI Act requires providers and deployers to ensure a sufficient level of AI literacy among staff operating AI on their behalf, proportionate to role and risk, and the Joint Commission guidance expects role-specific training as one of its seven elements. Meridian runs three tracks: 20 minutes for all staff on what AI is in use and how to report a concern; 90 minutes for clinicians who receive AI output, focused on when to distrust it; and a half-day for system owners and the committee.
Model performance is not a property; it is a reading taken at a point in time. Populations shift, documentation practices change, a coding update alters an input distribution, a new EHR release changes how a field is captured — and performance moves without anyone touching the model.
Every metric has a threshold set in advance and an owner, and crossing it triggers a defined action rather than a discussion about whether it matters.
AI safety events need a route into the existing patient-safety reporting system, not a parallel one — clinicians will not use a second system. Meridian added an "AI involved?" flag to the standard event report and a four-level severity ladder:
| Level | Trigger | Response |
|---|---|---|
| Sev 1 | Patient harm reached the patient, AI output contributed | Immediate suspension, executive and board notification within 24h, root cause analysis, regulatory reporting assessment |
| Sev 2 | Near miss — caught before reaching the patient — or a confirmed subgroup performance gap | Committee review within 5 days, mitigation plan, consider restricting use |
| Sev 3 | Monitoring threshold breached, no known patient impact | Owner investigates within 10 days, reports to committee |
| Sev 4 | Usability complaint, alert burden, workflow friction | Logged, trended, reviewed in aggregate quarterly |
Sev 4 matters more than it looks. Alert-burden complaints are the leading indicator for a Sev 1 — because the mechanism by which a noisy model causes harm is that people stop reading it.
The governance question nobody asks at purchase: what happens when the vendor changes the model? In healthcare AI this is routine and often invisible — a new version ships in a release, and the tool your clinicians trust is not the tool you validated.
Meridian's rule: a model version change is a change, and goes to the change advisory board with a governance reference. For Tier 1 systems, a version change triggers abbreviated revalidation before it reaches production — a shadow comparison of old versus new on recent data, and confirmation that subgroup performance has not moved. The contractual notice clause from step 5 is what makes this possible; without it, you learn about version changes from your monitoring dashboard, after the fact.
On the device side there is now a formal mechanism for this. FDA's predetermined change control plan lets a manufacturer specify in advance which modifications to an AI-enabled device may be made without a new marketing submission, and how they will be validated. For Meridian as a deployer, the practical question is simple and worth asking every vendor: is there a PCCP, and what does it permit you to change without telling us?
Independent of any change: Tier 1 annually, Tier 2 every two years. Same protocol as the original validation, against current data. Revalidation regularly finds that a model still works but its clinical context has moved — the care pathway changed, and the model is now optimizing for something the organization stopped caring about.
The most under-designed step in AI governance. Retiring a clinical model requires: notice to users with a reason, a defined fallback process, removal from the workflow rather than just disabling the alert, retention of the model's historical outputs where they informed documented clinical decisions, data return or destruction under the contract, and a register update. Meridian retired two of the eleven systems in year one. Neither had a decommission plan, and one of them had outputs embedded in three years of clinical notes that had to be interpretable retrospectively.
The ten steps are the machine. Scenarios are what the machine is for. Each of the six below is a real decision Meridian's committee had to make, with a date on it and people waiting for an answer — the kind of thing that arrives on a Tuesday and cannot be deferred to the next quarterly meeting.
Each one follows the same shape: the situation, the signal that surfaced it, three options with their actual consequences, what was decided and why, the artifact it produced, and the transferable lesson. The options are not strawmen. In every case, at least two of the three were seriously argued in the room, and in several cases the option that lost is the one most organizations pick.
The three options in each scenario are genuine forks, not a right answer flanked by two obviously wrong ones. The marked branch is what Meridian chose, given Meridian's facts. Change the facts — a different patient population, a different contract, a different regulator — and a different branch becomes correct. What transfers is the reasoning, not the verdict.
Sentinel was bought in a hurry during a quality improvement push, configured by the EHR team, and switched on. Nobody validated it locally. Nobody was monitoring it. When the governance committee's inventory sweep found it, the clinical sponsor who had championed it had left the organization eighteen months earlier, and the current owner of record was a build analyst who had never seen the vendor's performance claims.
The committee did not turn it off. It ran the Step 6 silent period against the live system: the model kept scoring, the alerts kept firing, and a parallel adjudication compared what Sentinel said against a chart-review gold standard on a stratified sample.
This was not a fishing expedition. There was a published reason to look. An external validation of a widely deployed proprietary sepsis model at an academic health system found an AUC of 0.63 against a vendor-reported 0.76–0.83, 33% sensitivity, alerts on 18% of all hospitalizations, and 67% of sepsis cases missed entirely — with roughly 109 alerts needed to identify one case of sepsis developing within four hours.1 Meridian wanted to know whether it had bought the same problem.
The signalIt had. The 90-day silent period produced numbers that were, if anything, slightly worse:
| What was measured | Vendor claim | Meridian, 90 days | What it means operationally |
|---|---|---|---|
| Discrimination (AUC) | 0.83 | 0.64 | Barely better than a coin flip weighted by how sick someone looks |
| Sensitivity at default threshold | not stated | 34% | Two out of three sepsis cases never generate an alert |
| Alert rate | not stated | 17% of admissions (6,613 alerts) | Roughly one alert per nurse per shift, on top of everything else |
| Number needed to evaluate | not stated | 96 | 96 bedside evaluations to find one true early sepsis case |
| Median lead time vs. existing screen | “hours earlier” | 42 minutes | Real, but small |
| Cases where a nurse had already escalated first | — | 71% | In most alerts, the human got there first anyway |
The subgroup analysis found something the aggregate number hid. At the flagship academic hospital, AUC was 0.67. At the safety-net hospital, it was 0.58. The model leaned heavily on the cadence of structured vital signs and lab results, and the safety-net hospital charted vitals less frequently and ordered fewer routine labs. The model was not reading sickness there. It was reading documentation density — and reading it as health.
That is the finding that changed the conversation. A model that performs worst at the hospital serving the most vulnerable patients is not a performance problem. It is an equity problem wearing a performance problem's clothes.
The three options on the table6,613 alerts a quarter at 96 evaluations per true case. Alert fatigue is not confined to the alert that causes it — response rates degrade across every alert in the system, including the ones that work.
And the committee has now documented, in writing, that it knew. Knowing and doing nothing is a materially worse legal position than not knowing.
Modelled at threshold 8: alerts fall to 6% of admissions, but sensitivity falls from 34% to 19%. You buy quiet by missing four out of five cases instead of two out of three.
This is the most commonly chosen option in practice, because it makes the complaint stop. It does not make the model work; it makes the model quieter about not working.
Alerts stop firing on 62 units where nurses were already catching it first. The model is re-fit on Meridian data with the vendor contractually on the hook, and returns only on the two night-shift units with the thinnest staffing ratios.
Alerts route to the rapid response nurse, not the bedside nurse. Six-month sunset: it comes back to committee or it switches off automatically.
Sentinel was suspended house-wide within nine days of the validation report, using the committee's standing authority to suspend a live system without waiting for a quarterly meeting. That authority, written into the charter in Step 1 for exactly this reason, is what made a nine-day timeline possible instead of a ninety-day one.
The reasoning that carried the room was not the AUC. It was the 71%. A model whose alerts mostly arrive after a human has already acted is not adding safety; it is adding interruptions. The committee reframed the question from “is this model accurate?” to “what does a clinician do differently because of it, and is that different thing better?” On 62 units the answer was nothing. On two night-shift units with one nurse to seven patients, the answer was plausibly something — so that is where it went back, and only there.
A vendor's AUC is a claim about someone else's patients, measured on someone else's documentation habits. Discrimination is the least useful number in the deployment decision. The ones that decide it are alert burden, number needed to evaluate, and incremental value over what the humans already do — and none of those can be read off a datasheet.
CareLens decides who gets a care manager. A care manager is a nurse who calls you, arranges your transport, reconciles your medications, and notices when you have stopped filling a prescription. For a patient with four chronic conditions and no car, it is the difference between managed illness and a series of admissions. There are 3,100 slots and roughly 61,000 attributed lives, so the ranking is the entire allocation mechanism.
The system was procured by the health plan, not the hospitals, and had never been reviewed as a clinical tool because on paper it is not one. It does not diagnose. It does not treat. It ranks.
The signalA care manager mentioned in a huddle that her panel had almost nobody from the county the safety-net hospital serves. An analyst pulled the numbers to check. Patients from that catchment were 31% of the attributed population and 14% of the enrolled.
The first instinct in the room was to check the inputs for race. Race was not an input. Neither was ZIP code, in any direct form. By the standard the organization had been applying — we do not feed it protected characteristics — CareLens was clean.
The Step 7 assessment asks its questions in a specific order, and question one is not about inputs. It is “what is the model actually predicting?” The answer was cost. Not illness, not need, not deterioration. Cost.
This is a known and quantified failure. A widely used commercial risk algorithm applied to roughly 200 million people in the United States used healthcare cost as a proxy for healthcare need. At any given risk score, Black patients turned out to have 26.3% more chronic illnesses than white patients with the same score. Correcting the label — predicting illness rather than expenditure — raised the share of Black patients who would receive additional help from 17.7% to 46.5%.2
Meridian ran the same test on its own data. At an identical CareLens score, patients in the lowest income quartile of the attributed population had on average 1.9 more active chronic conditions and 2.4 more uncontrolled-condition indicators than patients in the highest quartile. The model was systematically ranking sicker poor patients below healthier affluent ones, and it was doing so correctly — because affluent patients with better access do consume more care, and therefore do cost more. The model was not broken. It was answering the question it had been given.
Under Section 1557, a patient care decision support tool is any automated or non-automated tool used to support clinical decision-making. Deciding who receives intensive care management is squarely inside that. The rule requires reasonable efforts to identify tools that use race, color, national origin, sex, age or disability as input variables, and policies to mitigate the risk of discrimination those tools create. Meridian had satisfied the first duty trivially — no protected variables — and completely failed the second.
The three options on the tableThere is nothing to remove — race was never an input — so this means adding an explicit race-based adjustment to a resource allocation tool. That is a considerably harder thing to defend to a regulator than the problem it fixes.
It also treats the symptom. The score still means “expected cost.” You have applied a patch on top of a number that measures the wrong thing.
The disparity metric goes green immediately. The dashboard looks fixed at the next board meeting.
But within each service area the ranking is still by cost, so you are still enrolling the wrong patients — now with a fairness metric certifying that you are not. The measurement has been repaired and the mechanism has not.
Requires the vendor to re-target the model, or replacement. It took seven months and a contract renegotiation. Two of the three shortlisted vendors could not say what their label was.
Re-scoring the prior twelve months under the new label identified 640 patients who would have qualified and were never enrolled. All 640 got outreach.
The committee treated the label as the defect, because the label was the defect. Every downstream mitigation — reweighting, quotas, corrections, thresholds — is an attempt to talk a model out of answering the question it was trained on. The only durable fix is to change the question.
The harder decision was the backward look. Nothing compelled Meridian to re-score the past. The committee did it anyway, on the reasoning that if you discover an allocation mechanism has been misallocating, the people it misallocated away from are identifiable, still your patients, and still sick. The 640 outreach calls were the most expensive line item in the remediation and the only one that helped anyone who had already been harmed.
Removing the protected variable is not mitigation, and its absence is not a defence. This model had no race input, no ZIP code, and no ethnicity flag, and it still allocated care by proxy for race and income — because the outcome it was trained to predict had already absorbed decades of unequal access. Audit the target variable first. It is where the bias lives, and it is the one thing an input audit cannot see.
Scribe was the popular one. It was the system nobody wanted governance anywhere near, because it was working, clinicians were grateful, and the organization had spent a decade making documentation worse. It was classified Tier 2 rather than Tier 1 on the reasoning that it does not make a clinical recommendation — it transcribes and organizes, and a licensed human signs every output.
That reasoning was correct about what the system does and wrong about what the signature means.
The signalAn urgent care physician was reviewing a draft note for a patient who had come in with an ankle injury. The review of systems read: “Patient denies chest pain, shortness of breath, or palpitations.” She had never asked. The conversation, all six minutes of it, was about an ankle. The audio confirmed it — no one said any of those words.
This was not a transcription error. The model had produced a fluent, clinically plausible, conventionally formatted pertinent-negative statement because that is what a review of systems looks like in its training data. The sentence was grammatical, appropriate to the setting, and written in the physician's own documentation voice. It was also entirely invented.
The committee ordered an audit: 500 randomly sampled signed notes reconciled against retained audio. The result:
The last number is the finding. Every fabrication had passed through the control that was supposed to catch it. Median time from note presentation to signature was 11 seconds. The human in the loop was in the loop and was not reading.
Two further problems surfaced once the committee had the contract in front of it. The vendor's default configuration retained encounter audio indefinitely — a growing archive of recorded clinical conversations, PHI in the most sensitive form the organization holds, with no retention limit anyone had agreed to. And the standard terms granted the vendor rights to use customer audio for model improvement. Nobody at Meridian had ever decided that patient conversations would be used to train a commercial product. Somebody had signed a contract that said so.
The three options on the table610 clinicians lose 47 minutes a day. The measured improvement in documentation burden reverses overnight, and governance becomes the function that took away the one thing that helped.
It also solves nothing structural. Attestation-as-a-click is a problem for every draft-generating system Meridian will ever deploy, and suspending this one leaves that intact.
This is the intervention organizations reach for because it is fast, cheap, documentable and feels responsible. It generates a training completion rate to report.
It relies on sustained vigilance against a failure mode engineered to be invisible. The fabricated text is fluent, plausible and in the reader's own voice. Banners habituate in about a week.
Generated pertinent negatives and unverbalized exam findings suppressed entirely at the vendor level. Speaker attribution required before any history is recorded as the patient's.
Median attestation time moved from 11 seconds to 48. Clinicians still overwhelmingly preferred it to typing — adoption fell by under 3%.
The committee's insight was that the defect was not in the model's accuracy but in the shape of its output. A draft that mixes what was heard with what was inferred, in identical formatting, makes verification cognitively impossible at speed — and clinical documentation always happens at speed. The fix was to make the two categories visually and structurally different, and to make the second one block the signature until a human touched each item.
The one-action override principle from Step 8 was deliberately inverted here. Everywhere else, the committee insists that overriding the AI must take one action and never a form, because friction on the override is friction on the human's judgment. Here the friction is the control: accepting an inference the model made up should be harder than deleting it.
“A licensed human reviews every output” is the most common control in healthcare AI governance and the most frequently untrue. A signature box is not a control; it is a place to record that a control was supposed to happen. If the interface makes it possible to attest without reading, the interface has decided the outcome and your policy is documentation of an intention.
Meridian sits on both sides of this table. It runs a 48,000-life Medicare Advantage plan that uses AI to triage authorization requests, and it is a provider whose own requests are triaged by other payers' algorithms. The committee found the second fact clarifying. Several of the people arguing that UM Assist was obviously fine had spent the previous year complaining, at length, about exactly the same technology pointed at them.
The compliance position was considered settled. CMS is explicit that an algorithm may be used to assist in coverage decisions but cannot serve as the sole basis to deny or terminate coverage, and that a determination must rest on the individual patient's circumstances rather than on what a larger dataset says about people like them. UM Assist never issued a denial. A nurse reviewer did, and a medical director signed the ones that required it. Sole basis, satisfied.
The signalThe overturn rate. Of the denials that patients or providers appealed, 58% were being overturned — up from 34% in the eighteen months before UM Assist went live. An overturn is the plan telling itself, in writing, that its own determination was wrong.
The committee pulled the reviewer timestamps and found the mechanism:
| Reviewer behaviour | Case flagged “likely meets criteria” | Case flagged “unlikely to meet criteria” |
|---|---|---|
| Median review time | 4 min | 3 min |
| Median review time, cases with no flag shown (control period) | 14 min | |
| Reviewer disagreed with the flag | 9% of cases | 6% of cases |
| Additional clinical records requested before deciding | 11% | 8% |
| Denials later overturned on appeal | — | 58% |
Three minutes. On a case the model had already labelled unlikely to meet criteria, a nurse reviewer spent three minutes and agreed 94% of the time. The model was not assisting the determination. It was making the determination and the reviewer was ratifying it — which is precisely the arrangement the rule exists to prohibit, arrived at without anyone deciding to do it.
The distinction between assists and decides is not settled by the policy document, the vendor's product description or the org chart. It is settled by the timestamp log. Meridian's said decides.
The three options on the tableTechnically defensible right up to the moment someone requests the timestamps. The three-minute median and the 58% overturn rate are already in the plan's own systems and would be produced in any audit or litigation.
The overturn rate is also a patient harm measure. Every overturned denial is a patient who was told no, and who only got to yes by having the resources and persistence to appeal.
Review capacity drops roughly 40% against fixed regulatory turnaround deadlines. Turnaround failures are their own patient harm — a delayed approval is a delayed treatment.
It also discards something that genuinely works. On the approval side the model is accurate, and faster approvals hurt nobody.
The “unlikely to meet criteria” flag was removed from the reviewer interface entirely — not de-emphasised, not accompanied by a caution, removed. A reviewer working a case that may end in denial sees no model output at all.
Review capacity is preserved, because most requests are approvals. Median review time on potential-denial cases returned to 15 minutes. Overturn rate fell to 31% over the following two quarters.
The argument that won was about the asymmetry of the errors. A wrong approval costs the plan money. A wrong denial costs a patient their treatment, and lands hardest on the patients least equipped to appeal. When the consequences of the two error types are that different, the automation should be that different too. Meridian's rule became: a model may make the benign outcome faster; it may not make the harmful outcome easier.
The committee also rejected the framing that removing the flag was a loss of functionality. Anchoring is not a bug in the reviewer. It is how human judgment works, reliably, in every profession that has been measured. Designing a workflow that shows a busy clinician a confident-looking verdict and then asks for independent judgment is asking for something people cannot do.
“The AI only assists” is a claim about behaviour, and behaviour is measurable. If your reviewers spend three minutes on flagged cases and fourteen on unflagged ones, the model is deciding and your policy is describing an organization you do not have. Instrument the humans, not just the model.
PortalDraft had passed intake cleanly. It was validated, the drafts were good, clinician satisfaction was high, and the human-review control looked solid: no message reaches a patient without a clinician sending it. It was the least controversial system in the register.
It was also the system whose validation had the shortest half-life, and nobody had noticed because nothing in the process asked how long a validation stays true.
The signalThe complaint was vague and the patient was right. Reconciling the week before and the week after the weekend upgrade:
| Output characteristic | Before | After | Meridian's standard |
|---|---|---|---|
| Reading level (Flesch–Kincaid grade) | 8.2 | 12.6 | 6th–8th grade |
| Median draft length | 94 words | 212 words | — |
| Drafts appending a generic “consult your physician” closing | 3% | 76% | — |
| Drafts declining to address the question directly | 2% | 14% | — |
| Drafts sent with no clinician edit | 61% | 61% | — |
The last row is the one that should worry you. The system changed materially and the human review layer did not register it at all — the same 61% went out untouched, now at a twelfth-grade reading level, telling patients who had just messaged their physician to consult their physician.
Meridian's validation had been performed against a model that, by Monday morning, no longer existed. The contract said nothing about notice, because nobody had thought to ask for it. To the vendor this was a version bump. To Meridian it was an unvalidated system in production, communicating with patients, for eleven days before anyone looked.
The disclosure question came up in the same review. California's AB 3030 requires generative-AI-produced clinical patient communications to carry a disclaimer and instructions for reaching a human — at the outset of a written message and persistently through a continuous online interaction — and exempts communications that a licensed human provider reads and reviews first. Meridian operates in neither California nor any state with an equivalent statute in force. The committee adopted the standard anyway, for two reasons: running two message pipelines with different disclosure rules is an error factory, and the exemption itself is the uncomfortable part. It turns on a human genuinely reading the message. Meridian had just established, at 61% sent unedited, that it could not evidence that for the majority of its traffic.
The three options on the tableFastest path, and it was seriously proposed — the new model was in most respects more capable.
It also concedes that the vendor's release schedule sets Meridian's patient communication policy. Health literacy standards exist because a twelfth-grade reply to a patient with an eighth-grade reading level is a failed communication regardless of how sophisticated it is.
Superficially the strongest control and the least sustainable. Vendors do not maintain arbitrary old versions indefinitely; you eventually sit on something unpatched and unsupported.
It also blocks improvements you want, including safety fixes, and creates pressure to grant blanket exceptions that quietly restore the original problem.
Thirty days' written notice of any change to the model, its version, its training data or its system prompt; a fifteen-business-day validation window before the change reaches production; and a right to reject.
Crucially, an automated output monitor — reading level, length, refusal rate, disclaimer presence — running daily, because the contract only protects you against vendors who comply with it.
Both halves of Option C were treated as load-bearing, and the committee was explicit that the monitor mattered more than the clause. A contract is a remedy after the fact. Monitoring is how you find out. Meridian had a patient complaint as its detection mechanism, which means the detection mechanism was a patient being poorly served and caring enough to telephone about it.
The reclassification argument was harder. PortalDraft was Tier 2 because a clinician sends every message. Having watched that control absorb a complete model swap without noticing, the committee moved it to Tier 1 — not because the system got riskier, but because the control it was relying on had been measured and found weaker than assumed.
A validated system is a validated version. Generative systems can be swapped underneath you between a Friday and a Monday, and the swap is invisible to every control that depends on a human noticing that the text reads differently. Monitor the outputs, not the release notes — and treat the day you cannot detect a version change as the day your validation expired.
NoduleAI was the least alarming system in the register. Regulated device, cleared by FDA, CE-marked, radiologist reads every study, well-understood technology with a decade of literature. It sailed through intake in the commercial deployment.
The research protocol was a different document, and it arrived at the committee only because Step 4 requires every new AI use to come through intake, including research uses that a research ethics committee has already approved. That rule felt bureaucratic when it was written. It is the reason this was caught.
The signalThe lawyer's question was simple and nobody in the room could answer it: in Europe, is Meridian a deployer or a provider?
The answer turned out to be both, in different capacities, with different obligations, on different timelines:
| Capacity | Role | Route | Date that governs | What it requires |
|---|---|---|---|---|
| Clinical use at Meridian's US hospitals | Deployer | FDA / US law | Now | Local validation, monitoring, human oversight, HTI-1 source attributes from the EHR developer |
| Clinical use at the European partner site | Deployer | EU AI Act, Annex I — safety component of a regulated device | 2 August 20284 | Use per instructions, human oversight by competent staff, input data relevance, log retention, incident reporting |
| The modified model built by the collaboration | Provider | EU AI Act, Annex I | 2 August 2028 | Conformity assessment, risk management system, technical documentation, data governance, post-market monitoring — the full provider stack |
| The institute's EU-facing trial participant assistant | Deployer | EU AI Act, Article 50 transparency | 2 August 2026 | Tell people they are interacting with an AI system |
Two findings came out of that table. The first is the one the committee went looking for: substantially modifying a high-risk system, or putting your name on it, makes you its provider. The research institute had been thinking of itself as a data contributor to someone else's product. Under the EU framework it was building a new high-risk medical device system and planning to put it into service.
The second finding was an accident. While confirming that the 2028 date was right, the committee checked whether anything else in the register touched the EU regime and found the trial participant assistant — a generative chatbot answering questions from European trial participants, stood up by the research institute nine months earlier, never entered in the register, and squarely inside Article 50's transparency obligation. Article 50's timeline was not moved by the Digital Omnibus amendments that pushed the high-risk dates back. It applies from 2 August 2026.
The committee reviewed this on 24 July 2026. The obligation was nine days away, on a system nobody had known existed, discovered while researching a deadline two years out.
The committee also asked the manufacturer whether NoduleAI had a predetermined change control plan, and what it permitted. It did: retraining on additional data, threshold recalibration, and performance improvements within the cleared indication — all without a new marketing submission. That is exactly what a PCCP is designed to allow, and it is entirely legitimate. It also means the model in Meridian's radiology department can change materially, lawfully, and with no regulatory event to notice. The only notice Meridian gets is whatever the contract requires, and the contract required nothing. Scenario 5's clause was retrofitted here the same week.
Correct for the commercial deployment and wrong for the collaboration. The provider obligation attaches to substantial modification and to putting the system into service under your own name, not to who wrote the original code.
It would also have left the Article 50 exposure entirely undiscovered, since nobody would have looked.
Ends a genuinely valuable program to avoid an obligation that is demanding but well-defined, with a two-year runway to meet it.
This is governance as an avoidance function, and it is how governance loses the standing to be consulted early. Say no to the hard thing once and the next protocol will not come to intake at all.
Deployer obligations assigned to clinical operations; provider obligations assigned to the research institute with named accountability and a budget, because provider obligations are a programme, not a checklist.
The modified model may be developed and evaluated. It may not be put into service at any site until the provider-side documentation is complete and the committee has approved it as a separate decision.
The committee treated role clarity as the deliverable. Most of the confusion in cross-border AI governance comes from organizations that assume they occupy one role for all purposes, when in practice a health system is a deployer of most things, a provider of a few, and occasionally both for the same underlying model in different capacities. Writing the roles into a table, capacity by capacity, resolved arguments that had been circling for weeks.
The collaboration was redesigned to train the modified model by federated learning — the model travels to each site, trains locally, and only parameter updates are exchanged. No European patient data leaves its jurisdiction, which simplifies the transfer analysis considerably and was, the institute's own investigators pointed out, better science as well, since it allowed a wider partner network than any data-sharing agreement would have supported.
Provider and deployer are capacities, not identities. The same organization can hold both for the same model, and modifying a system or putting your name on it moves you across the line whether or not you intended to become a manufacturer. And the far-off deadline is rarely the dangerous one — the dangerous one is the near date on the system nobody entered in the register.
Each scenario is written so it can be handed to a group with the decision removed. Give them the situation, the signal and the three options, ask for a decision and the reasoning behind it, then reveal what Meridian chose. The disagreements are more instructive than the answers — particularly on Scenario 4, where the option most groups pick first is the one the timestamps refute, and on Scenario 1, where the technically-minded reach for the threshold adjustment almost every time.
Healthcare AI does not have one law. It has a stack of regimes written at different times for different purposes, none of which were designed with each other in mind, and most of which predate the technology they now govern. A single system routinely sits inside four or five of them at once. Sentinel is simultaneously a clinical decision support intervention under HTI-1, a quality and safety matter under Joint Commission expectations, a potential Section 1557 concern because of the safety-net subgroup gap, and a patient safety reporting obligation. None of those regimes will tell you about the others.
What follows is the map Meridian built. It is organised by what triggers each regime rather than by which agency issued it, because the trigger is the useful part when a new system arrives at intake.
| Regime | What pulls you in | What it makes you do | Date |
|---|---|---|---|
| ACA Section 1557 patient care decision support tools |
Any automated or non-automated tool used to support clinical decision-making. Scheduling, supply chain and staffing tools are outside it; anything shaping what care a patient is offered is inside. | Reasonable efforts to identify tools using race, color, national origin, sex, age or disability as input variables, and policies to mitigate the resulting discrimination risk. Enforcement weighs entity size and resources, whether the tool was used as intended, how it was customised, developer guidance, and whether you have a governance process at all. | Compliance date 1 May 2025 |
| ONC / ASTP HTI-1 decision support interventions |
Certified health IT that supplies decision support. Falls on the EHR developer, but the hospital is the party that needs the information. | Predictive DSI must ship 31 source attributes across 9 categories — intervention details and output, purpose, out-of-scope cautions, development and input features, fairness assurance, external validation, performance metrics, ongoing maintenance, and the update and revalidation schedule. Evidence-based DSI requires 13. Plus intervention risk management: analysis, mitigation, governance. | Developer delivery 31 Dec 2024; in effect 1 January 2025 |
| FDA software as a medical device |
Software intended to diagnose, treat, prevent or mitigate disease. Clinical decision software that a clinician cannot independently review the basis of is generally in scope; the marketing language is often the tell. | Clearance or approval for the device itself. For the deployer, the operative question is the predetermined change control plan: what has the manufacturer pre-authorised itself to change — retraining, recalibration, performance improvements — without a new marketing submission? | In force |
| HIPAA | Any use or disclosure of protected health information, including audio, images and anything handed to a vendor's model. | Business associate agreements with AI-specific terms: no secondary use for model training, defined retention and destruction, subcontractor and sub-processor disclosure, breach notification that covers model outputs. De-identification via Safe Harbor or Expert Determination, with Expert Determination the realistic route for rich clinical data. | In force |
| CMS Medicare Advantage 42 CFR 422.101(b)(6) |
Algorithms used in coverage, utilization management or medical necessity determinations by a Medicare Advantage organization. | An algorithm may assist but cannot be the sole basis to deny, discontinue or terminate coverage. The determination must rest on the individual patient's circumstances, not on what a larger dataset predicts about similar patients. MAOs must also not use algorithms in a way that violates Section 1557. | FAQ guidance 6 February 2024 |
| CMS WISeR model | Original Medicare fee-for-service prior authorization in six states — New Jersey, Ohio, Oklahoma, Texas, Arizona and Washington — for a defined service list including skin and tissue substitutes, electrical nerve stimulator implantation, and knee arthroscopy for osteoarthritis. Inpatient-only, emergency, and services where delay creates risk are excluded. | AI or machine learning may be used, but every non-affirmation recommendation must be made by an appropriately licensed clinician. Participation is voluntary. Worth reading even if you are outside the six states: it is CMS designing a compliant AI-assisted review process, which makes it a useful reference standard. | 1 January 2026 – 31 December 2031 |
| EU AI Act Annex III — standalone high-risk |
Standalone high-risk systems listed in Annex III. For healthcare organizations this typically catches employment, access to essential services and creditworthiness uses rather than clinical devices. | Full high-risk obligations, allocated by whether you are the provider or the deployer. | Moved to 2 December 20274 |
| EU AI Act Annex I — embedded / safety component |
AI that is a safety component of a product already covered by Union harmonisation legislation. Medical devices route here, which is where most clinical AI lands. | Providers: conformity assessment, risk management system, technical documentation, data governance, post-market monitoring. Deployers: use per instructions, human oversight by competent staff, input data relevance, log retention, incident reporting. | Moved to 2 August 20284 |
| EU AI Act Article 50 transparency & Article 4 AI literacy |
Systems that interact directly with people, or generate synthetic content, where the output is used in the Union. Article 4 applies to anyone operating or overseeing an AI system. | Tell people they are interacting with an AI system; mark synthetic content. Ensure staff have a sufficient level of AI literacy — a legal obligation, not a training aspiration, and one you have to be able to evidence. | 2 August 2026 — not moved by the Digital Omnibus amendments |
| Joint Commission & CHAI Responsible Use of AI in Healthcare |
Accredited organizations. Guidance rather than a standard, with a voluntary certification following it. | Seven pillars: AI policies and governance structures; patient privacy and transparency; data security and data use protections; ongoing quality monitoring; voluntary blinded reporting of AI safety events; risk and bias assessment; and education and training. | Guidance 17 September 2025; certification 2026 |
| State law | Varies enormously and changes fast. California AB 3030 (generative AI in clinical patient communications must disclose and give human contact instructions, unless a licensed provider reads and reviews first) and SB 1120 (AI may not determine medical necessity; only a licensed physician or qualified professional may, using the individual's data) are the pattern others are copying. Colorado's original AI Act is being replaced by legislation applying to consequential decisions from 1 January 2027. | Disclosure, human review, individualised determination, documentation retention, adverse-outcome notification. The specifics differ by state and by year. | Volatile — verify before relying |
Federal regimes change on multi-year cycles with notice-and-comment in between. State AI law in 2025 and 2026 has changed on a scale of months, including statutes that took effect and were then paused by litigation or repealed and replaced before enforcement began. Build your load-bearing controls on the federal and accreditation requirements, which are stable, and treat state obligations as a layer you re-check quarterly. Do not architect a workflow around a single state statute without an owner assigned to watch it.
One further point about the map, which Meridian's committee reached the hard way in Scenario 6: the dangerous date is almost never the famous one. Everybody in healthcare AI governance knows the EU AI Act high-risk dates. Far fewer noticed that the transparency obligations were carved out of the delay and kept their original timeline, or that the obligation attaches to a chatbot the research institute stood up without telling anyone. Track every applicable date in one calendar, with a named owner against each, and re-check the whole calendar rather than the entry you were already worried about.
Meridian did not do the ten steps in order and neither will you. Steps 1 to 3 have to come first because everything else depends on them, but after that the sequence is driven by what the inventory found, and the inventory will find something that cannot wait. What follows is roughly how the first year actually went, including the part where the plan changed in month four.
Board quality committee charters the function, not IT. Three powers written down explicitly: require validation before deployment, suspend a live system, escalate to the board directly. CMO chairs. CNO on it, because clinical AI reaches nurses first. Internal audit deliberately excluded, to preserve an independent third line.
Five discovery channels run in parallel: EHR configuration review, contract database search, network and SaaS discovery, departmental interviews, and a CEO-signed 30-day amnesty for shadow AI. Leadership estimated three or four systems. The sweep found eleven.
Risk tiers assigned on consequence times independent human judgment, plus the five-question regulatory routing. Six of the eleven systems matter. Sentinel goes first — highest tier, longest running, most patients, no validation ever performed.
Sentinel's validation comes back at AUC 0.64 with a safety-net subgroup gap. The system is suspended house-wide within nine days. Everything else slips a month, the committee earns its credibility in a single decision, and intake volume triples the following quarter because people have now seen that the process does something.
Intake form, time-boxed review (15 business days Tier 1, 10 days Tier 2), BAA addendum with the evidence clause and the change notice clause, and the silent-period protocol. CareLens equity assessment runs here and produces Scenario 2.
Human-in-the-loop specified concretely per system: who, what information, at what time, with what authority, and what happens on disagreement. Override is one action, never a form. Patient disclosure standardised. Training runs at three depths — 20 minutes for all staff, 90 minutes for clinicians, half a day for system owners. Scenario 3 surfaces here and rewrites the attestation standard.
Six monitoring dimensions per Tier 1 system: input drift, output drift, performance, subgroup performance, human interaction, operational health. Four-severity incident ladder with defined thresholds and on-call ownership. Scenarios 4 and 5 are both detected by monitoring rather than by complaint — which is the entire point of building it.
Version changes route to the change advisory board. Every vendor is asked whether there is a predetermined change control plan and what it permits them to change without telling you. Tier 1 revalidates annually, Tier 2 every two years. Decommission procedures written — the most under-designed step in almost every programme. Scenario 6 lands here and adds the jurisdictional role analysis.
Intake is routine and mostly uncontroversial. The register is current because three intake triggers keep it current, not because someone runs an annual sweep. Monitoring catches things before patients do. The committee spends most of its time on genuinely hard cases rather than on discovering systems it did not know existed. And the second inventory sweep, run twelve months after the first, finds two more systems — which is a healthy number, not a failure.
A governance programme is, concretely, about a dozen documents that exist and get used. Everything else is meetings. This is the set Meridian ended the year with, mapped to the step that produced it and the person who owns it.
| Artifact | What is in it | Owner | From |
|---|---|---|---|
| Committee charter | Mandate, membership, the three powers, escalation path, meeting cadence, quorum, conflict-of-interest rules | Board quality committee | Step 1 |
| AI register | Every system, its owner, tier, intended use, out-of-scope uses, validation status, monitoring status, last review date, next revalidation date | Governance lead | Step 2 |
| Risk tiering rubric | Consequence × independent human judgment matrix, plus the five-question regulatory routing | Committee | Step 3 |
| Intake form | Intended use in one sentence, out-of-scope uses, the label question, the counterfactual question, the asymmetry question, data sources, affected populations | Governance lead | Step 4 |
| BAA AI addendum | No secondary use for training, retention and destruction, sub-processor disclosure, the evidence clause, the model change notice clause, audit rights | Legal & privacy | Step 5 |
| Validation report template | Silent-period design, discrimination, calibration, sensitivity and specificity at the operating threshold, alert burden, number needed to evaluate, incremental value, subgroup performance | Clinical informatics | Step 6 |
| Equity assessment | The four questions in order, starting with what the model is actually predicting; subgroup results; proxy analysis; mitigation and its evidence | Health equity & clinical informatics | Step 7 |
| Human oversight specification | Per system: who, what information, at what time, with what authority, and what happens on disagreement | System owner | Step 8 |
| Attestation design standard | Generated content distinguishable from captured; inferred content requires per-item action; time-to-signature monitored as control effectiveness | Clinical informatics | Scenario 3 |
| Monitoring plan | Six dimensions, thresholds, frequency, who reads it, what happens when a threshold trips | System owner | Step 9 |
| Incident procedure | Four severity levels, definitions, on-call ownership, suspension authority, patient notification triggers, reporting obligations | Patient safety & governance lead | Step 9 |
| Change control record | Version, what changed, PCCP scope, revalidation performed, approval, rollback plan | Change advisory board | Step 10 |
| Role determination worksheet | Per system, per jurisdiction, per capacity: provider or deployer, and the obligations that follow | Legal | Scenario 6 |
| Regulatory calendar | Every applicable date, in one place, with a named owner and a quarterly re-check | Legal | Scenario 6 |
| Decommission procedure | Notice to users, data disposition, record retention, what happens to outputs already embedded in the record, vendor offboarding | System owner | Step 10 |
For each artifact, ask whether you could hand it to a regulator, an accreditor or a plaintiff's expert tomorrow with no preparation. If the honest answer is that it would need to be written up first, it does not exist — you have a practice and a memory of a practice, which is a different thing. Governance that cannot be evidenced is indistinguishable, from the outside, from governance that never happened.
Most AI governance metrics measure the governance function's activity rather than its effect. Number of systems reviewed, policies published, training completions — these tell you the committee met. They do not tell you whether anything is safer. The ones below are chosen because each has a plausible way of going wrong that you would want to know about.
Three of these — override rate, time-to-signature, and the Sev-4 complaint stream — are the leading indicators. They move before patients are harmed. Everything else on the list tells you about something that has already happened. If you can only build three monitors in the first year, build those, and put a name against each one so somebody is obliged to look.
Meridian's programme worked, which is partly why it is worth being explicit about the ways it nearly did not. Every failure below is one the committee either walked into or came close to, and most of them look like success from the inside for a considerable period before they announce themselves.
The committee without power. The most common failure and the hardest to reverse. A group that can recommend but not require, chartered under IT rather than under the board, is an advisory body that will be consulted after the contract is signed. Meridian's nine-day Sentinel suspension was only possible because the authority to suspend had been written down before anyone needed it. Retrofit that authority during a crisis and you will spend the crisis negotiating instead of acting.
The inventory that decays. A one-time sweep produces a register that is accurate for about a quarter. Without intake triggers wired into procurement, contract renewal and EHR configuration change, the register becomes a historical document that everyone treats as current. Meridian's second-year sweep found two systems that had appeared since the first — a good outcome, because two is what a working process leaks, and eleven is what an absent one accumulates.
Validation theatre. A validation report that reports AUC and stops. It is the number vendors supply, the number that sounds most scientific, and the number least connected to whether deploying the system will help anyone. If your validation template does not force alert burden, number needed to evaluate, subgroup performance and incremental value, it will produce documents that satisfy an auditor and tell a clinician nothing.
Human-in-the-loop that is not. The single most over-claimed control in healthcare AI. “A licensed clinician reviews every output” is a statement about workflow design, and it is false by default in any workflow that presents a confident-looking output to a busy person under time pressure. Three minutes on a flagged prior authorization; eleven seconds on a drafted note. Neither was a policy failure. Both were interface decisions that determined the outcome long before any policy applied.
Monitoring nobody reads. Building the dashboard is the funded part; reading it every week is not. A monitoring plan without a named person, a cadence and a defined action when a threshold trips is a data collection exercise. Meridian's rule became that any monitor without a trigger and an owner gets switched off, on the reasoning that an unwatched monitor is worse than none — it creates the belief that someone is watching.
Equity assessment as an artifact rather than a practice. A bias assessment performed once at intake, filed, and never repeated. Populations shift, care patterns shift, and models drift; a fairness result from eighteen months ago describes a system that no longer exists. The CareLens disparity would not have been found by a document. It was found because a care manager noticed her panel looked wrong and somebody took her seriously enough to pull the numbers.
The pilot that never ends. Systems deployed as pilots to avoid the governance gate, then quietly becoming permanent. The defence is a hard sunset: a pilot without an expiry date and a named decision-maker at the end of it is a deployment with better branding.
Decommission never designed. Almost every programme has an intake process and almost none has an exit process. What happens to outputs already embedded in the medical record? Who tells the users? What is retained, for how long, and under whose retention schedule? Meridian wrote its decommission procedure in month eleven, having already suspended a system in month four without one.
Governance as an avoidance function. The subtlest failure. A committee that says no to hard things stops being asked about hard things, and the interesting work relocates to wherever the committee is not. Meridian's most important decision in Scenario 6 may have been declining to exit the European collaboration — not because the compliance analysis demanded it, but because a governance function that only ever subtracts capability has a finite lifespan.
One date in the calendar. Knowing the headline regulatory deadline confidently and never checking whether the others moved differently. Meridian was nine days from missing a live transparency obligation on a system it did not know it operated, and found it by accident while researching a date two years out.
This case study is built to be taught, not only read. The scenarios have their decisions separable from their setups precisely so a group can be made to commit to an answer before seeing what Meridian chose — which is where the learning is. Two formats work.
Half day, roughly three hours. Thirty minutes on the organization and what the first inventory found, which is enough to establish that everyone in the room is probably underestimating their own count. Then three scenarios at forty minutes each: Scenario 1 for validation and incremental value, Scenario 2 for the target variable, and Scenario 4 for the gap between the policy and the timestamps. Close with the metrics section and ask each participant to name the one leading indicator they could stand up within a quarter.
Full day. All six scenarios at forty-five minutes each, with the ten-step build presented in the morning as the framework the scenarios test. Give groups the situation and the signal, hand them the three options with the consequences redacted, and have them write the consequences themselves before revealing them. The redacted-consequence version generates far better discussion than the complete one, because arguing about what an option would actually cost is the skill the job requires.
A few facilitation notes drawn from running it. Scenario 4 is the reliable one for a mixed clinical and administrative audience, because both groups arrive certain that a human signing the determination settles the question, and the timestamp table lands hard. Scenario 2 works best with a group that includes someone from analytics, who will usually be the first to see that the label is the defect and can then explain it to the room more persuasively than a facilitator can. Scenario 3 is the one that changes behaviour after people leave, because everyone present has signed something in eleven seconds.
The most productive question to hold in reserve, for the point where a group has settled comfortably on the right answer: what would have to be different about your organization for the option you just rejected to be correct? It is the question that turns a case study into a transferable method rather than a set of remembered verdicts.
Meridian Health System does not exist. Its size, structure and system portfolio are a composite, and its internal numbers — the 38,900 admissions, the 4.2% audit finding, the 58% overturn rate — are constructed to be realistic rather than reported. The regulatory regimes, the published research findings, and the failure modes are real, cited, and verifiable. When you use this with a group, be explicit about which is which. A case study whose facts are treated as citable when they are illustrative does more harm than good.
The numbers attributed to published research, as distinct from Meridian's constructed internal figures, come from the following.
Regulatory positions described elsewhere on this page — Section 1557 patient care decision support tools, ONC/ASTP HTI-1 decision support intervention source attributes, FDA software as a medical device and predetermined change control plans, the CMS WISeR model, the Joint Commission and CHAI guidance of 17 September 2025, and the California and Colorado statutes — are summarised from the primary rules, guidance documents and legislation as they stood in mid-2026. Verify current status before relying on any of it for a compliance decision; the state-level positions in particular have changed repeatedly.