flowchart TB
A([AI Deployment\nProposal]) --> B[Privacy and\nSecurity Screening]
B --> C[Equity Impact\nAssessment\nSubgroup analysis design]
C --> D[AISC Ethical Review\nConsent architecture · IP · Liability]
D --> E{Approved?}
E -->|No| F[Revise or Decline]
E -->|Yes| G[Community\nEngagement Plan]
G --> H[Pilot Deployment\nReal-time equity monitoring]
H --> I[Governance Report\nSubgroup performance · Incidents]
I --> J{Performance\nEquitable?}
J -->|Yes| K([Continue with\nperiodic review])
J -->|No| L[Suspend · Investigate · Redesign]
L --> C
16 Ethics, Equity, and Institutional Accountability
The ethical challenges of AI in the AMC are not primarily about individual decisions by individual clinicians. They are structural. A predictive model that systematically underestimates the health needs of Black patients does not fail because the clinician using it is biased; it fails because it was trained on data that encodes decades of inequitable access to care, and deployed without monitoring that would detect the systematic underperformance (Obermeyer et al. 2019). If an ambient documentation system turns out to transcribe some patients less accurately than others, it will not be because the clinician was careless; it will be because the model was trained on speech that overrepresented certain patterns and deployed without anyone stratifying its performance. The pattern is consistent: the ethical failures that have actually occurred in deployed healthcare AI are structural failures, predictable in advance, and correctable through governance — if the governance exists.
This chapter argues that AMC AI ethics requires a structural turn. The question is not “does this AI tool respect individual patient autonomy?” but “does the process by which this institution deploys, monitors, and governs AI tools systematically protect patient equity and institutional accountability?” Individual ethical review is necessary but not sufficient. Structural governance is the mechanism through which individual ethical commitments become institutional practice.
16.1 Algorithmic Bias as a Structural Problem
The most cited demonstration of algorithmic bias in healthcare involves a commercial risk stratification algorithm used by major health systems to identify high-risk patients who would benefit from care management programs (Obermeyer et al. 2019). The algorithm used healthcare costs as a proxy for health need — a reasonable proxy if access to care were uniformly distributed, which it is not. Black patients with the same actual health burden as white patients had systematically lower costs, because they had systematically less access to care. The algorithm therefore scored them as lower risk, directing care management resources away from patients who needed them more.
The finding was not that the algorithm was malicious. It was that the algorithm was trained on data that encoded an existing inequity, using a proxy variable that faithfully reproduced that inequity at scale. Race was never an input. The authors estimated that changing the prediction target from cost to illness would raise the share of Black patients flagged for additional help from 17.7% to 46.5% (Obermeyer et al. 2019). Subsequent work has documented the same shape of problem in other tools: race correction factors built directly into kidney function estimates and cardiac risk calculators (Vyas et al. 2020), and dermatology imaging models that lose accuracy on darker skin tones (Daneshjou et al. 2022).
The structural governance response requires three elements that individual ethical review cannot provide alone. First, demographic stratification of performance metrics as a standard part of AI validation — not an optional audit, but a required table in every model evaluation. Second, monitoring that continues after deployment, because bias can emerge over time as the population served changes or as the model is applied to use cases outside the validation set. Third, a reporting structure that routes performance stratification findings to clinical leadership and the governance committee, not just to the informatics team.
16.2 Health Equity as a Performance Metric
The framework proposed by Badal and colleagues (Badal et al. 2023), introduced in Chapter 6, includes the alleviation of health disparities as the first principle of responsible clinical AI. In practice, this requires operationalizing “equity” as a quantitative performance dimension alongside accuracy. A model that achieves 85% accuracy overall but 70% accuracy for the subpopulation with the highest disease burden is not a high-performing model — it is a model that performs best for the patients who need it least.
The practical implementation of equity monitoring requires demographic data in the validation and monitoring datasets. This is straightforward in principle and difficult in practice, because demographic data in EHRs is often missing, inconsistent, or coded in ways that do not capture the granularity needed for subgroup analysis. AMCs that serve diverse populations should treat demographic data quality as an AI readiness issue: the return on investment for improving race, ethnicity, and language data completeness comes precisely when the institution needs to assess whether its AI tools are performing equitably.
The FUTURE-AI international consensus guideline, developed by 117 experts across 50 countries and published in 2025, establishes fairness as one of six guiding principles for trustworthy healthcare AI, alongside universality, traceability, usability, robustness, and explainability (Lekadir et al. 2025). These principles map to operational AMC practices: fairness requires subgroup performance evaluation; traceability requires audit logging of model outputs; explainability requires that clinicians can access the reasoning behind a model recommendation. The NIST AI RMF (National Institute of Standards and Technology 2023) organizes its own trustworthiness characteristics into a Govern/Map/Measure/Manage structure rather than FUTURE-AI’s, and neither document maps itself onto the other. They are compatible enough in practice that an institution can run the FUTURE-AI principles as technical requirements inside the NIST functions, but that pairing is a recommendation of this book rather than something either framework prescribes.
16.3 Informed Consent in the Continuous AI Era
The consent framework governing clinical AI does not fit the technology. The traditional consent model is episodic and discrete: a patient is informed about a specific intervention at a specific moment and chooses whether to accept it. Clinical AI tools are continuous and ambient: a predictive model runs on every patient in the database, continuously, updating as new data arrives. The patient is not present when the model generates a risk score; there is no discrete moment at which they can meaningfully consent or decline.
Ambient documentation creates a more visible consent challenge — the patient is in the room when the AI is active — but the design solution is tractable. The verbal consent model described in Chapter 13 and Chapter 12 provides explicit, encounter- specific authorization. The harder case is the background AI: the readmission risk model, the sepsis predictor, the no-show prediction algorithm. These tools affect every patient encounter without any patient-facing disclosure.
The position this book takes, and it is a position rather than a settled consensus, is that “invisible” background AI calls for institutional-level disclosure rather than individual patient consent: the institution publishes a clear statement of which AI tools are used in patient care, how they affect clinical decisions, and what patients can do to learn more. This does not resolve the ethical tension between efficient deployment and individual autonomy, but it addresses the transparency gap. The 2021 WHO guidance on the ethics and governance of AI for health treats meaningful transparency about automated decision-making as a prerequisite for ethical deployment (World Health Organization 2021).
16.5 The Regulatory Turn: HHS Section 1557 and the Duty to Mitigate
For most of the 2010s, AI equity was a governance aspiration — a principle that showed up in AMC values statements and academic papers but carried no specific legal obligation. The 2024 HHS Section 1557 final rule changed that (U.S. Department of Health and Human Services, Office for Civil Rights 2024). Under 45 C.F.R. § 92.210, covered entities are required to take reasonable steps to identify and mitigate discrimination in patient care decision support tools that are used to make, recommend, or facilitate clinical decisions (U.S. Department of Health and Human Services 2024). The rule explicitly covers algorithmic and AI-assisted tools, and compliance was required by May 1, 2025, three hundred days after the rule took effect.
What “reasonable steps” means in practice is not fully defined by the rule. But the class of tool it targets is not mysterious: a care management algorithm that scores patients on a proxy for need is exactly the pattern Obermeyer and colleagues documented, and it is the pattern the rule’s identify-and-mitigate duty is shaped around. An institution deploying a readmission risk model, a care gap identification tool, or a utilization management system without having assessed its performance across demographic groups cannot demonstrate compliance. The assessment does not need to be a clinical trial; it needs to be documented evidence that someone looked. The equity audit process in Section 16.11 is what that documented evidence looks like.
The risk in Section 1557 is not just regulatory penalty. It is reputational and evidentiary. If a patient files a discrimination complaint and the institution cannot produce documentation that it evaluated whether its AI tools affected that patient’s demographic group differently, the absence of documentation is itself evidence of unreasonable practice. Building the equity audit function is not a compliance checkbox — it is the institutional record that will matter when a complaint arrives.
16.6 Beyond Obermeyer: Recent Cases of Algorithmic Bias
The Obermeyer 2019 finding — that a commercial risk stratification algorithm used a cost proxy that systematically underestimated the health needs of Black patients — is the most cited demonstration of algorithmic bias in healthcare, and it risks becoming a comfortable historical example that lets institutions off the hook for examining what their own deployed tools are doing right now.
The 2022 to 2025 literature documents the pattern continuing. Daneshjou and colleagues built the Diverse Dermatology Images dataset, the first publicly available image set that is both pathologically confirmed and curated for skin tone diversity, and showed that state-of-the-art dermatology models degrade substantially on dark skin tones and on uncommon diseases. Fine-tuning those models on the new images closed the gap between light and dark skin, which places the fault in the training data rather than in the task (Daneshjou et al. 2022). Ambient documentation is the obvious place to look for a speech-based analogue, and an AMC can measure that locally, but the published evidence on clinical dictation accuracy stratified by accent or dialect is thin. Treat it as a hypothesis to test in your own deployment, not as an established finding.
The 2024 Senate Permanent Subcommittee on Investigations report on Medicare Advantage prior authorization deserves to be read precisely, because it is routinely cited for more than it says (U.S. Senate Permanent Subcommittee on Investigations 2024). The subcommittee found that in 2022 UnitedHealthcare and CVS denied prior authorization for post-acute care at roughly three times their overall denial rates, and Humana at more than sixteen times its overall rate. That is a comparison between service lines within each insurer, not a comparison between algorithmic and human review, and the report is explicit that the evidence it obtained did not establish the extent of Humana’s use of automation at all. The number does not show what it is usually quoted to show.
What the report does document about automation is narrower and, for governance purposes, more useful. An internal UnitedHealthcare committee reviewing an automated authorization model in early 2021 recorded that testing had produced faster handling times along with “an increase in adverse determination rate,” and voted to approve the model the following month. An institution watched automation raise its own denial rate, wrote that down, and proceeded. The governance failure there is not that the tool was biased. It is that a measured increase in adverse decisions was treated as an acceptable cost of speed.
The ProPublica investigation into Cigna’s PxDx system found company physicians rejecting claims at an average of 1.2 seconds each, more than 300,000 denials over two months (Rucker et al. 2023). PxDx is not a learning model; it is an algorithm that flags mismatches between diagnosis codes and the procedures Cigna considers acceptable for them, after which physicians sign off on denials in batches. The distinction matters less than it appears. What made the system work was not the sophistication of the matching but the volume it created, which reduced physician review to a formality. That is not human-in-the-loop review. It is human-in-the-loop theater. For an AMC using algorithmic or AI-assisted prior authorization and utilization management, the governance question is not only whether those tools carry demographic bias, which to some degree they almost certainly do, but whether the human review attached to them is substantive enough to catch and override the bias when it appears.
16.7 State Privacy Laws Beyond HIPAA
HIPAA remains the dominant privacy framework for clinical AI, but it is no longer sufficient as a complete governance guide. A patchwork of state laws has emerged in the 2022 to 2025 period that creates obligations for AI use at AMCs that operate in, or serve patients from, specific states.
Washington’s My Health MY Data Act regulates consumer health data that falls outside HIPAA’s scope, meaning data collected by apps, wellness tools, and AI systems whose operators are not covered entities (Washington State Legislature 2023). It was signed in 2023, but the dates that matter are staggered: the geofencing prohibition took effect in July 2023, and the main obligations on most regulated entities on March 31, 2024. The Act requires separate opt-in consent for collection and for sharing, gives consumers a right to withdraw consent and to demand deletion, and bans geofencing around healthcare facilities, which reaches location-based AI tools and mobile health applications. Because the Act’s definition of “consumer health data” is broad enough to capture AI-generated health inferences, an AMC deploying patient-facing AI tools that touch Washington residents needs to analyze the Act’s requirements specifically.
Colorado HB 26-1139, signed on June 2, 2026 and effective January 1, 2027, bars entities performing utilization review from denying coverage on medical necessity grounds “solely on the output of an AI system without human review by a licensed clinician or physician or other competent regulated professional” (HB 26-1139 2026). The statute also requires that any AI used in utilization review base its determinations on the patient’s own clinical history rather than on group data alone, that it comply with anti-discrimination law, and that it be reviewed periodically for accuracy. Note the bill number carefully: HB 24-1139, from the 2024 session, is an unrelated workers’ compensation measure. For AMCs that operate health plans or run care management programs with AI-assisted utilization management, the 2026 statute is a direct Colorado compliance obligation, and the January 2027 effective date means the work of documenting where AI touches utilization review needs to happen now.
Illinois BIPA’s health care exemption is broader than plaintiffs had argued and narrower than institutions sometimes assume. In Mosby v. Ingalls Memorial Hospital, decided November 30, 2023, the Illinois Supreme Court held that the exemption is not limited to patients’ biometric information, and covers nurses’ fingerprint scans used to access medication and supply cabinets in the course of patient treatment (Illinois Supreme Court 2023). That matters for AMCs using ambient audio, retinal scans, or other biometric identification in clinical workflows. What Mosby did not do is exempt biometric data collected for security access, time and attendance, or administrative identification, and those uses remain exposed to the Act’s consent and notice requirements.
The broader pattern is that HIPAA compliance is a floor, not a ceiling. Each state where an AMC operates patients, employs staff, or deploys patient-facing digital tools may impose additional requirements on AI-related data handling, consent, and human review. The institutional legal review process for AI deployments needs to include state law analysis, not just HIPAA review.
16.8 The Workforce and Labor Dimension
An ethics chapter about AI in the AMC that does not address what happens to the people whose work AI changes is incomplete. The institutional ethics question here is not whether to deploy AI tools that make some existing roles redundant — that is already happening — but how the institution manages the human consequences of that displacement.
The roles most directly affected by AI automation in the current wave are not clinical roles requiring complex judgment. They are roles involving high-volume, structured, repetitive cognitive work: medical coders whose work is partially automated by AI-assisted coding tools; prior authorization specialists whose decisions are increasingly pre-populated or reviewed by AI; transcriptionists who have seen their role transformed or eliminated by ambient documentation; certain radiology reading functions where AI handles high-volume, lower-complexity cases.
The institution that deploys AI tools that reduce the need for these roles without an explicit workforce transition program — retraining, reassignment, severance, outplacement — is making an ethical choice, whether or not it acknowledges it as one. AMCs that have invested in the relationships with their frontline staff that make clinical quality possible should not treat AI-driven workforce changes as a pure efficiency calculation. The social compact that allows an AMC to function as a clinical and community institution is relevant to how it manages the people affected by AI-driven change, not just to how it treats patients.
16.10 Liability, the Standard of Care, and the Duty to Use
Liability for clinical AI is developing rather than settled, and the honest summary is that almost all of the settled part runs one way. Clinicians and institutions face recognized exposure for harms caused by following AI recommendations without adequate oversight. The mirror-image exposure, liability for failing to use a tool, has not been established by any court, and this section should not be read as reporting that it has.
What the legal literature does establish is narrower and more interesting. Price, Gerke, and Cohen laid out the template that most subsequent analysis works from: under existing tort doctrine, a physician’s liability turns on whether the course taken matched the standard of care, not on whether an algorithm was consulted (Price et al. 2019). The practical consequence is that current law protects the physician who follows standard care and declines the AI, and offers the least protection precisely where AI would add the most value, when the model correctly identifies that the standard course is wrong for this patient. On that reading, tort law is a brake on AI adoption rather than an engine of it.
The experimental evidence complicates that picture. Tobia, Nielsen, and Stremitzer put four vignettes to a nationally representative sample of 2,000 U.S. adults, varying whether an AI system recommended standard or nonstandard care and whether the physician accepted or rejected the recommendation, with harm resulting in every case. Potential jurors judged physicians who accepted standard-care AI advice less harshly than those who rejected it, and there was no comparable protection from rejecting nonstandard advice in favor of the standard course. The authors concluded that tort law is unlikely to undermine the use of medical AI and may encourage it (Tobia et al. 2021). A later replication extended the design to German physicians and German adults, testing whether the pattern survives in a regime where court-appointed experts rather than lay jurors determine liability. It did (Tacconelli et al. 2026).
Read together, these are evidence about how decisions get evaluated after the fact, not a holding that a duty to use AI exists. The extrapolation this book makes from them is its own: as tools for specific tasks accumulate evidence of performance at or above specialist-level accuracy, and as the people who judge reasonableness come to treat consulting those tools as the ordinary thing a careful clinician does, the gap between ethical argument and professional expectation narrows. A radiologist who does not use an FDA-cleared tool for pneumothorax detection, where that tool has demonstrated sensitivity superior to unassisted reading, is in a position that will be harder to defend in 2032 than in 2026. That is a prediction. Institutions should plan for it without treating it as current law.
The prudent governance posture is to document, for each deployed clinical AI tool, the institutional reasoning about when and how it should be used — not just the existence of the tool, but the clinical judgment about its appropriate role. When a clinician overrides an AI recommendation, that decision should be documentable. When a clinician relies on an AI recommendation, that reliance should be documentable. The medical record is the primary liability defense; it should reflect the clinician’s engagement with AI tools, not hide it.
16.11 Where to Start
16.11.1 Starter Project 1: Equity Audit of Deployed Clinical AI
What it is: A structured retrospective audit of performance stratification for the two or three highest-impact clinical AI tools currently deployed, assessing whether performance metrics vary significantly by race, ethnicity, age, insurance status, and language.
Why now: HHS Section 1557 requires that covered entities not deploy discriminatory patient care decision-support tools. The section 1557 final rule is in effect. An institution that has not assessed its clinical AI tools for demographic performance variation cannot certify compliance, and more importantly, cannot know whether its tools are harming the patients most at risk.
How to execute: Work with the clinical informatics team to extract retrospective performance data for each tool, stratified by available demographic dimensions. Identify subgroups with statistically significant performance differences. Assess whether the difference is clinically meaningful and whether it reflects a correctable bias in the model or an irreducible clinical population difference. Report findings to clinical leadership and the governance committee. For tools with significant performance disparities, develop a remediation plan.
Buy vs. build: Analytical work using existing institutional data. Commercial AI governance and model monitoring platforms are marketed for this purpose and may save engineering time, but none of them is a prerequisite, and the subgroup comparisons at the core of the audit are ordinary statistics your analysts can already do.
16.11.2 Starter Project 2: Clinical AI Ethics and Accountability Policy
What it is: A published institutional policy on the ethical deployment of clinical AI that addresses the four structural elements described in this chapter: equity monitoring requirements, consent architecture for background AI, IP and authorship accountability, and documentation requirements for AI-assisted clinical decisions.
Why now: Without a published policy, there is no institutional standard to hold deployments to, no governance anchor for the ethics review pipeline in Figure 16.1, and no document to point to when a patient asks why an AI tool was used in their care.
How to execute: Draft using the NIST AI RMF as the governance scaffold and the FUTURE-AI principles as the technical requirements framework. Review with legal (liability and IP), compliance (Section 1557, applicable state AI law, which for Colorado institutions means SB 26-189 from January 2027 (SB 26-189 2026) and HB 26-1139 from the same date if the institution touches utilization review (HB 26-1139 2026), and the rest of Chapter 10), clinical leadership (standard of care implications), and patient representatives (consent and disclosure language). Publish as institutional policy with a defined review cycle aligned with the annual AI governance report.