CANNOT BE
OUTSOURCED.
The hospital needs a complete chain of authority.
Background and Objective
Clinical artificial intelligence can influence diagnosis, prioritization, treatment, documentation, and resource allocation, but a model does not assume the hospital’s obligations to patients. Risk arises through the combined behavior of data, software, interfaces, clinicians, vendors, policies, and operating conditions. This review examines how hospital leaders can govern decision rights, bias, malpractice exposure, and lifecycle assurance without treating either regulatory authorization or nominal human oversight as sufficient proof of safe use.
Methods
A targeted PubMed/MEDLINE narrative search was completed on August 13, 2026 for English-language peer-reviewed research published from January 2000 through August 13, 2026. Eligible evidence addressed clinical-AI validation, implementation, human factors, automation bias, subgroup performance, liability perceptions, patient notice, drift, incidents, or governance. Primary statutes, regulations, and current official agency materials were reviewed separately to describe applicable legal and regulatory structures. News, preprints, vendor marketing, and unsupported commentary were excluded as evidence.
Key Content and Findings
Peer-reviewed studies show that model performance can change across institutions, patient groups, thresholds, workflows, and time. Inaccurate advice can degrade clinician judgment, and explanations do not reliably prevent harmful reliance. Bias can enter through target selection, labels, hidden proxies, missingness, deployment, or unequal response. Generative systems add prompt, version, hallucination, and source-verification risks. Hospital assurance therefore requires a use-case inventory, named decision owners, local and subgroup validation, human-factors testing, controlled deployment, patient-facing rules appropriate to purpose and risk, vendor transparency, continuous monitoring, incident investigation, change control, and enforceable stop authority.
Conclusions
Hospitals should govern clinical AI as a clinical service across its full lifecycle. The model may advise, but the organization must retain authority, evidence, monitoring, and remedy. Authority should remain visible through approval, deployment, material change, incident review, corrective action, suspension, and final retirement of every clinical use case.
Keywords: artificial intelligence; clinical decision support; algorithmic bias; malpractice; hospital governance
Introduction: The Algorithm Advises; The Hospital Answers
Clinical artificial intelligence is often purchased as software, evaluated as a model, and presented to clinicians as a recommendation. In care delivery, however, it functions as a distributed clinical service. Its performance depends on source data, interfaces, thresholds, staffing, training, workflow, downstream capacity, and the authority given to its output. A model can rank a patient, draft a note, identify an image finding, predict deterioration, or recommend an action. The organization decides where the output appears, who sees it, whether it interrupts work, which response is expected, and how an error is detected.
Adoption has moved faster than assurance. In a national hospital survey, 65% of responding nonfederal acute-care hospitals reported using predictive models. Among model users, 61% reported locally evaluating most or all models for accuracy and 44% reported doing so for bias [24]. A subsequent national survey found that 31.5% of responding hospitals reported electronic-health-record-integrated generative AI use in 2024, while 24.7% planned use within one year [25]. These self-reported findings do not establish effectiveness or harm. They identify a management gap: clinical influence is expanding while local evaluation remains uneven.
The gap matters because keeping a clinician nominally “in the loop” does not guarantee independent judgment. Experimental research shows that inaccurate advice can lower clinician performance and that explanations do not reliably neutralize systematically biased output [3-6]. External-validation studies likewise show that performance and workload can change substantially across institutions and thresholds [16-18]. The hospital therefore cannot treat a vendor metric, regulatory status, or professional license as a substitute for evidence about the combined human-AI system in the local environment.
This review advances an operating premise: the algorithm advises; the hospital answers. The premise does not mean that a hospital will be legally responsible for every model-associated injury. Malpractice, product, organizational, and contractual responsibility remain fact-specific and jurisdiction-dependent [7]. It means that the hospital controls essential safeguards and should be able to prove how authority, evidence, monitoring, and remedy operated. The objective is to synthesize peer-reviewed research and separately reviewed primary legal authorities into an executive framework for intended use, validation, human judgment, equity, malpractice risk, vendor governance, monitoring, change control, and board assurance. This article is presented in accordance with a narrative review reporting checklist.
Methods
This narrative review used a purposive and reproducible search designed for hospital-management relevance. Literature published from January 1, 2000 through August 13, 2026 was eligible. Targeted PubMed/MEDLINE blocks were rerun on August 13, 2026 and generally emphasized 2020 onward; the large-language-model block began in 2022. Earlier foundational studies were identified through reference and related-record review. Search concepts combined hospital, health system, clinician, and clinical decision support with artificial intelligence, machine learning, predictive model, large language model, generative AI, validation, calibration, generalizability, automation bias, human factors, algorithmic bias, subgroup performance, malpractice, liability, patient notice, governance, procurement, monitoring, model drift, incident response, and change control.
Eligible evidence included empirical studies, systematic or scoping reviews, prospective or retrospective validation studies, comparative implementation evaluations, and peer-reviewed qualitative or mixed-methods research that informed a hospital decision, safeguard, or measurable risk. Priority was given to national, multisite, prospective, systematic, recent, or methodologically informative evidence. Preprints, news, conference abstracts, vendor marketing, corporate white papers, law-firm blogs, and unsupported commentary were excluded as evidence. Peer-reviewed conceptual frameworks were used to organize governance questions, not to claim effectiveness. Reference lists and related records were reviewed to identify additional eligible studies.
The author conducted relevance screening, selection, extraction, synthesis, and citation verification. Findings were organized by use-case authority, validation, human-AI interaction, equity, liability, patient-facing duties, vendor control, monitoring, incidents, and lifecycle assurance. Because the literature included experiments, retrospective validations, surveys, implementation studies, document analyses, and systematic reviews with different outcomes, no pooled estimate was attempted. No duplicate independent screening or formal risk-of-bias instrument was used.
For each included source, the extraction recorded design, setting, population, task, model or intervention, comparator, reported performance or outcome, limitation, and governance relevance. Numerical results were retained only when tied to the studied population, model version, threshold, and design. Experimental judgments were not treated as clinical outcomes, public regulatory-document gaps were not treated as proof that testing never occurred, and survey reports were not treated as verified organizational performance. These distinctions were carried into the synthesis and tables.
Legal and regulatory research was conducted separately so that statements of law would not be confused with empirical findings. Sources included applicable federal statutes and regulations, current FDA guidance, health IT certification requirements, nondiscrimination provisions, privacy and security requirements, and device-reporting rules [35-42]. Guidance was identified as nonbinding where applicable. This review does not provide a fifty-state survey. Malpractice, informed-consent, corporate-negligence, evidence, and product-liability questions require analysis of the facts and governing jurisdiction.
| Element | Approach |
|---|---|
| Date of search | August 13, 2026 |
| Database and other sources | PubMed/MEDLINE; reference-list and related-record screening; primary legal authorities reviewed separately |
| Timeframe | January 1, 2000 through August 13, 2026 |
| Core concepts | Clinical AI; machine learning; predictive models; generative AI; validation; calibration; automation bias; human factors; equity; liability; consent; governance; monitoring; drift; incidents; change control |
| Inclusion | English-language peer-reviewed empirical research, systematic or scoping reviews, validation studies, implementation evaluations, and relevant mixed-methods or qualitative research |
| Exclusion | Preprints; news; conference abstracts; vendor marketing; corporate white papers; law-firm blogs; unsupported commentary; conceptual proposals used as proof of effectiveness |
| Selection process | Single-author relevance screening, extraction, narrative synthesis, and citation verification; reference and related-record review |
| Synthesis | Purposive narrative synthesis; no formal risk-of-bias instrument or pooled estimate |
| Legal sources | Primary statutes and regulations plus current official agency materials, reviewed separately from empirical evidence |
Accountability Begins With the Clinical Decision
An inventory of products is not an inventory of clinical risk. The governable unit is the use case: a defined model version using specified inputs for a particular population, setting, user, output, timing, and expected action. The same software may be administrative in one workflow and clinically consequential in another. A summarization tool that prepares an internal draft differs from a system that prioritizes emergency patients, withholds an alert, or recommends treatment. Governance should begin with the decision that may change, not with the vendor category.
Hospitals have adopted different organizational models. Interviews with early-adopter systems identified decentralized translation, information-technology-led governance, and dedicated AI teams, each with different tradeoffs [22]. A single-site implementation study similarly found that infrastructure, frontline co-design, training, trust, role clarity, and workflow ownership were central to moving a sepsis model into routine care [23]. Neither study establishes one superior organizational chart. Together, they show that technical deployment and clinical accountability are inseparable.
Every use case should therefore identify at least four distinct rights. A sponsor may propose the purpose. A validation authority may judge whether evidence is adequate. A clinical owner may define the permitted action and escalation path. A deployment authority may release, suspend, or retire the service. One person may hold more than one role in a smaller organization, but the rights should remain visible. The clinician’s responsibility for patient-specific judgment does not eliminate the hospital’s responsibility for selecting, configuring, monitoring, and supporting the service. The vendor’s responsibility for product design does not remove the need for local controls.
The registry should distinguish ownership from participation. Information technology may operate an interface without owning the clinical decision. A data-science team may evaluate performance without deciding whether benefit justifies workflow burden. Legal or compliance staff may identify requirements without determining clinical utility. A multidisciplinary review can improve judgment, but participation must not obscure who can approve, reject, pause, and close the use case.
The clinical warranty chain
Every link must be named before the first patient is affected.
Intended Use Defines the Boundary of Safe Reliance
Intended use should be written in operational language. It should state who is included and excluded; the setting and phase of care; required inputs; model version; output; threshold; authorized user; expected response time; permitted action; alternatives; and conditions requiring independent verification or escalation. A description such as “AI for sepsis” is not sufficient. It does not reveal whether the model screens all admissions, predicts an outcome at a particular horizon, interrupts a clinician, initiates a protocol, or merely informs review.
The model’s role also matters. It may inform by presenting a measurement, recommend by suggesting an option, prioritize by changing sequence, constrain by limiting choices, or automate by initiating an action. Each role creates different consequences if the output is late, missing, incorrect, or ignored. Generative systems add a further distinction between drafting and authoritative retrieval. Evidence that a language model can answer questions in isolation does not validate a particular clinical workflow, prompt, interface, source set, or version [32-34].
In a randomized study of physicians completing difficult clinical vignettes, access to a language model did not significantly improve reasoning compared with conventional resources [32]. In a separate simulation using sequential clinical decisions, models missed guideline steps, struggled with laboratory information, varied with prompts, and hallucinated tools [33]. These studies do not establish bedside failure rates. They show why a hospital must validate the exact task and interaction rather than infer reliability from general capability.
Regulatory status follows function and intended use rather than the marketing label. Under federal law, some clinical decision support functions fall outside the device definition only when the statutory criteria are satisfied, including conditions allowing a health care professional to independently review the basis of a recommendation and not rely primarily on it [35]. The FDA’s January 2026 guidance explains this boundary but is nonbinding. Hospitals should not infer that all clinical AI is a device or that all decision support is excluded. For an FDA-authorized function, authorization remains specific to the product and intended use. It does not prove local calibration, workflow fitness, subgroup performance, or net benefit.
| Lifecycle decision | Accountable owner | Required evidence | Stop or escalation trigger |
|---|---|---|---|
| Authorize the use case | Designated clinical sponsor and accountable executive | Clinical purpose, population, setting, intended action, alternatives, risk classification | Unclear purpose, no accountable clinical owner, prohibited or unsupported use |
| Approve data fitness | Designated data steward and validation authority | Source, provenance, completeness, representativeness, label definition, missingness | Unresolved data defect, material population mismatch, unstable label |
| Approve validation | Independent validation authority | External and local performance, calibration, thresholds, subgroup results, workflow and usability tests | Unacceptable error, unbounded uncertainty, inequitable performance, excessive burden |
| Authorize clinical workflow | Clinical service owner and patient-safety authority | User, interface, response expectation, escalation, downtime, documentation, training | Ambiguous decision right, unsafe automation, no fallback, failed usability test |
| Release to production | Deployment authority | Version lock, configuration record, monitoring plan, incident pathway, rollback readiness | Incomplete evidence package, missing monitor, no stop authority |
| Approve material change | Change-control authority | Change description, impact assessment, revalidation, regulatory and contract review | Unapproved model, data, interface, threshold, or workflow change |
| Respond to incident | Patient-safety leader with clinical, technical, legal, and vendor support | Containment, case review, version and logs, harm assessment, reporting analysis | Possible ongoing harm, repeated failure, uncertain scope, lost audit evidence |
| Suspend or retire | Accountable executive and clinical owner | Trigger breach, benefit-risk review, replacement and patient-continuity plan | Persistent harm, lost benefit, unsupported version, failed remediation |
Validation Must Match the Patients, Workflow, Time, and Consequence
Validation is not a certificate that travels unchanged with a model. It is evidence that a specified use is fit for a specified environment at a specified time. The evaluation should begin with the clinical consequence and comparator. A model may discriminate between higher- and lower-risk patients while remaining poorly calibrated, creating too many alerts, missing too many cases, or failing to improve action. Area under the receiver operating characteristic curve is useful, but it does not determine the positive predictive value at local prevalence, workload at the deployed threshold, or patient benefit.
External evidence shows why local evaluation is necessary. A retrospective evaluation of a widely implemented proprietary sepsis model found an area under the curve of 0.63. At the studied threshold, the model missed 67% of sepsis cases and alerted in 18% of hospitalizations [16]. A prospective multicenter evaluation of an updated version found better discrimination, but thresholds, positive predictive value, and alert burden varied across four systems. At a four-hour horizon, positive predictive value was 0.02 to 0.04, requiring review of 24 to 69 patients per true positive across sites [17]. These studies concern particular versions, settings, outcomes, and thresholds. They do not prove that sepsis models generally fail. They show that a vendor average cannot answer a hospital’s local capacity and benefit questions.
Cross-institution testing can expose hidden dependencies. A pneumonia model lost performance across institutions and could identify the source institution with high accuracy, demonstrating reliance on site-specific signals [18]. Regulatory-document studies found that public summaries of many AI-enabled devices did not consistently report multisite, prospective, demographic, sample, or detailed performance evidence [19-21]. Absence from a public summary does not prove that an evaluation was never performed. It does mean procurement teams should obtain the evidence needed for their own decision.
Before clinical release, a hospital should test data mapping, missingness, calibration, sensitivity, specificity, positive and negative predictive values, alert volume, subgroup performance, usability, downtime, and response capacity. High-consequence use should include prospective silent operation when feasible, followed by a controlled deployment with balancing measures. The comparator should be current care, not a theoretical clinician working without information. Validation should also examine the human-AI team, because interface placement, explanation, timing, and workload can change the result.
The validation plan should be consequence-aware. A false negative in screening, a false positive that activates scarce resources, and an inaccurate draft reviewed before use do not carry the same risk. Leaders should predefine acceptable uncertainty and response capacity rather than select a threshold after reviewing favorable results. Confidence intervals, missing subgroup data, and deviations from the intended population belong in the decision record. When benefit is not directly measured, the organization should identify the process measure used as a proxy.
The evidence package must lock the model, threshold, input transformations, exclusions, interface, workflow, and date. If any of those change materially, the prior conclusion may no longer apply. Validation is therefore a lifecycle gate, not a one-time implementation task.
The validation runway
Validation earns a bounded release, not a permanent certificate.
Human Oversight Must Be Designed, Not Assumed
Human oversight is meaningful only when a person has time, information, competence, authority, and an alternative course of action. A clinician who must accept an output to advance the workflow, cannot see its relevant basis, or lacks access to a confirmatory test is not exercising independent control. Likewise, a policy instructing clinicians to “use judgment” does not address predictable automation bias, anchoring, interruption, alert fatigue, or ambiguity about who may stop the system.
Randomized evidence demonstrates the risk. In a vignette study of 457 clinicians, standard AI with explanations modestly improved diagnostic accuracy, while systematically biased AI reduced it by 11.3 percentage points. Explanations did not reliably prevent the loss, and incorrect advice also reduced treatment-selection accuracy [3]. In a preregistered chest-radiograph experiment involving radiologists and internal- or emergency-medicine physicians, inaccurate advice reduced performance whether it was labeled as AI or as human advice, and substantial minorities repeatedly followed wrong recommendations [4]. A larger reader study found heterogeneous benefit across radiologists and tasks; seniority and prior AI experience did not reliably identify who would benefit, and inaccurate assistance degraded aggregate performance [5].
Controls should therefore target reliance, not merely access. Possible safeguards include independent review before revealing a recommendation, confidence gating, selective suppression of outputs in an unsafe range, second review for high-consequence disagreement, and clear escalation. In one laboratory study of anterior cruciate ligament MRI interpretation, automation bias accounted for 45.5% of errors made with AI. A proposed suppression zone was estimated to reduce those errors, but the mitigation was not prospectively demonstrated in care [6]. Hospitals should treat such strategies as testable controls, not proven universal solutions.
Training should use failure scenarios and measure behavior. Leaders should review overrides in both directions: following incorrect advice and rejecting correct advice. The goal is calibrated reliance. Documentation should capture the version and clinically material output without creating a defensive burden that causes new error. Effective oversight requires a functioning fallback and explicit authority to disagree, escalate, pause, or proceed without the model.
Competence should be evaluated at the team level. Yu and colleagues found that years of experience, subspecialty, and prior AI experience did not reliably predict which radiologists benefited from assistance [5]. Hospitals should not exempt senior clinicians from testing or assume that novice users need only more explanation. Simulation should reproduce the deployed interface, time pressure, prevalence, and escalation options. Monitoring should test whether reliance resembles the validated pattern and whether workload or staffing changes have altered it.
The human judgment boundary
Human oversight is meaningful only when the user can independently assess, disagree, and escalate.
Bias Is Produced Across the Whole Care Pathway
Algorithmic bias is not confined to whether a protected characteristic appears among the inputs. It can arise when leaders define the problem, choose a target, construct labels, select data, handle missingness, set thresholds, allocate follow-up resources, design the interface, or respond differently to the same output. A model can reproduce an inequitable system even when its technical performance appears strong. It can also reveal previously overlooked burden when the target better reflects patient experience.
A commercial population-health algorithm illustrates target risk. At the same score, Black patients had greater illness burden because prior health care cost was used as a proxy for need. Replacing that proxy would have increased the proportion of Black patients selected for additional care from 17.7% to 46.5% in the studied setting [9]. The result should not be generalized to every commercial model. It demonstrates that a mathematically accurate prediction of the wrong target can produce an inequitable allocation.
Aggregate performance can hide subgroup error. Chest-radiograph classifiers showed selective underdiagnosis among underserved and intersectional populations [10]. A 2025 evaluation of foundation models similarly found larger underdiagnosis disparities in some intersectional groups [15]. Models have also predicted self-reported race from medical images even after external testing and image degradation, with no complete explanation of the signal [11]. Race detectability alone does not prove discriminatory action or patient harm. It warns that deleting an explicit field does not remove every pathway through which demographic information can influence a model.
Bias is not inevitable or unidirectional. Dermatology models performed worse on dark skin and uncommon diseases in a diverse external evaluation, while fine-tuning with diverse images narrowed the skin-tone gap [12]. A knee-radiograph model linked to patient-reported pain explained more observed racial pain disparity than traditional radiographic grading [13]. That study did not establish improved treatment or outcomes, but it shows that target choice can expose, rather than conceal, unmet burden.
Generative systems require repeated testing. In a prompt audit, four commercial language models produced some responses that repeated debunked race-based ideas, with inconsistency across runs [14]. Because model versions and prompts change, one favorable demonstration cannot establish durable equity.
The remedy must address the pathway that created the disparity. Removing a protected field may leave hidden proxies. Rebalancing a training set may not correct an inequitable label. Equalizing a prediction metric may not change whether patients receive the service. The review team should trace the decision from problem definition through action and outcome before selecting a technical or operational correction.
Fairness Requires Subgroup Evidence, Operational Monitoring, and Remedy
Fairness cannot be reduced to one metric. Equal sensitivity, equal positive predictive value, equal calibration, equal allocation, and equal outcomes may conflict when baseline risk, access, measurement, and treatment differ. The appropriate test depends on the decision and harm. A screening tool may prioritize missed cases, while a scarce-resource allocation tool must also examine who receives the downstream service. The governance record should explain the chosen measures and consequences rather than declare a model “unbiased.”
Before deployment, hospitals should evaluate clinically justified subgroups and intersectional groups for data availability, calibration, false-negative and false-positive rates, positive predictive value, workload, and anticipated action. Small samples require uncertainty intervals and restraint. A subgroup estimate too imprecise to support a conclusion should remain an explicit uncertainty, not be reported as assurance. Where race, ethnicity, sex, age, disability, language, geography, insurance, or another characteristic is relevant and lawful to analyze, the team should examine both model performance and the care pathway that follows.
Deployment can alter fairness even when prediction remains stable. A score may direct patients to a service with limited capacity, require digital access, or be acted upon differently across units. Monitoring should therefore include uptake, time to response, overrides, downstream service receipt, adverse events, complaints, and outcomes. The model and the hospital response form one allocation system.
Federal nondiscrimination rules add a distinct legal lane. For covered entities, 45 CFR 92.210 addresses discrimination through patient-care decision-support tools and requires reasonable efforts to identify uses of variables or factors that measure protected characteristics and to mitigate discrimination risk [40]. The provision should be applied according to its text and current legal status, not converted into a universal technical standard.
A failed equity test requires a remedy owner. Options may include changing the target or threshold, improving data, limiting use, revising workflow, adding support, retraining within authorized controls, or suspending the service. Daneshjou and colleagues showed that diverse fine-tuning could narrow a measured gap in a specific dermatology task [12]. That is evidence that remediation can be possible, not proof that retraining is always appropriate. The hospital must verify the correction and ensure that improvement for one group does not create unacceptable harm elsewhere.
Equity review should close with a documented decision: approve, approve with limits, remediate and retest, or stop. An unresolved disparity should not disappear into a general monitoring plan.
The equity counterweight
Aggregate accuracy cannot cancel unequal error, access, response, or remedy.
Malpractice Risk Follows Authority, Reliance, Causation, and Proof
Clinical-AI malpractice law is not a settled national doctrine. Liability analysis depends on jurisdiction, professional standards, institutional conduct, product characteristics, warnings, contracts, and the facts linking conduct to injury. A peer-reviewed legal review concluded that clinicians, hospitals, and vendors may face different pathways of exposure and that reasonable procurement, validation, supervision, documentation, and incident response can affect defensibility [7]. That analysis is a legal framework, not an empirical estimate of claims or verdicts.
The empirical literature helps explain expectations but does not establish law. In a nationally representative experiment using hypothetical harmed-patient scenarios, participants viewed physicians as more reasonable when they followed AI advice that aligned with standard care. The pattern did not create a simple shield for rejecting nonstandard AI advice and following standard care [1]. In a survey of physicians and the public, both groups frequently assigned responsibility to physicians, while physicians more often attributed responsibility to vendors and health care organizations [2]. These findings concern perceptions. They do not determine duty, breach, causation, damages, or allocation in a court.
Operationally, risk follows the care pathway. Leaders should ask what the model was authorized to do, what evidence supported the deployed version, how the output reached the clinician, what a reasonable user could understand, whether the clinician retained meaningful choice, what action or inaction followed, and how that conduct related to harm. The record should preserve the model and configuration, material inputs and outputs, alert timing, relevant user interaction, clinical reasoning documented in ordinary care, and subsequent investigation. Documentation should support care and reconstruction without forcing clinicians to produce a legal essay for every output.
Deviation from an AI recommendation is not inherently negligent, and agreement is not inherently reasonable. A clinician may have patient-specific information unavailable to the model. Conversely, a recommendation may expose a risk the clinician should investigate. The safest policy avoids blanket instructions to follow or ignore AI. It defines intended use, independent-review expectations, escalation for consequential disagreement, and the circumstances in which the service should not be used.
Hospital responsibility is also broader than bedside choice. The organization selects the vendor, grants data access, configures thresholds, decides placement, allocates response capacity, defines training, monitors performance, and controls suspension. A contract may allocate financial risk between parties, but it cannot guarantee patient safety or erase professional and institutional obligations. Counsel should evaluate applicable state law and the specific product, service, and event.
Proof must be proportionate and clinically usable. For a high-consequence recommendation, the record may need the model version, material output, alert time, and response. For a low-risk draft that is fully reviewed, routine documentation may be adequate. Excessive logging can create privacy, discovery, storage, and workflow burdens, while insufficient logging prevents investigation. The hospital should define what is retained, who may access it, how long it remains available, and how it links to the patient-safety record.
| Care-pathway point | Foreseeable failure mode | Evidence question | Assurance control | Retained proof |
|---|---|---|---|---|
| Input and data quality | Missing, stale, mis-mapped, or nonrepresentative data | Did the deployed input match the validated specification? | Provenance, range and missingness checks, interface testing, downtime rule | Data dictionary, mapping, quality logs, incident timestamp |
| Model output | Incorrect, unstable, out-of-scope, or unsupported output | Which version, threshold, prompt, and configuration produced it? | Version lock, use restrictions, confidence or scope gating | Version record, configuration, material output and time |
| Interface and presentation | Salience, explanation, timing, or default creates unsafe reliance | Could the intended user understand and independently assess the output? | Human-factors test, clear limitations, independent-review pathway | Approved interface, usability results, training version |
| Clinician interpretation | Automation bias, anchoring, misapplication, or unjustified rejection | Was the output used within scope with meaningful professional judgment? | Escalation, second review, fallback, scenario-based training | Clinically material response and ordinary-care documentation |
| Action or inaction | Delayed, omitted, excessive, or misdirected care | What action was expected and what occurred? | Named response owner, service level, capacity plan, balancing measures | Action time, orders, handoff, exception or escalation record |
| Handoff | Output or uncertainty loses meaning across teams | Did the receiving team receive the relevant context and responsibility? | Standardized handoff and unresolved-risk field | Handoff record, owner, acknowledgment |
| Patient outcome | Harm, unequal access, lost benefit, or new burden | Is there a plausible relationship among output, conduct, and outcome? | Patient-safety review with clinical, technical, equity, and legal input | Case chronology, assessment, corrective action |
| Adverse-event review | Evidence lost, recurrence, or incomplete reporting | Was the event contained, reported when required, and closed? | Preservation, reporting analysis, vendor escalation, stop authority | Logs, reports, notices, remediation and closure evidence |
The liability evidence ledger
Defensibility depends on reconstructing authority, reliance, action, causation, and corrective response.
Patient Notice, Consent, and Contestability Require Purpose-Specific Rules
No general United States rule requires special informed consent for every use of clinical AI. A background model that checks an image, a patient-facing conversational agent, an autonomous device function, a research protocol, and an experimental recommendation do not present the same question. Hospitals should distinguish clinical informed consent, research consent, privacy authorization, general notice, explanation of professional judgment, and an opportunity to question or contest a decision.
Rose and Shapiro proposed a peer-reviewed ethical framework that tiers uses among no special notice, notification, and formal consent based on risk, materiality, departure from ordinary practice, and patient-facing operation [8]. It is a framework, not a legal rule or a prospectively validated intervention. Its value is that it rejects both categorical silence and universal consent. The hospital must still evaluate applicable state law, device requirements, research rules, institutional policy, and the clinical context.
Notice should be useful. A generic statement that the hospital “uses AI” does not tell a patient what affected the decision, what the clinician reviewed, or how to raise a concern. For a material patient-facing use, communication should identify the function in plain language, the clinician’s role, important limitations, available alternatives where applicable, and a route for questions or human review. Contestability should not promise that every score can be fully explained or reversed. It should provide a meaningful process to correct wrong inputs, report unexpected outcomes, request review, and avoid an automated dead end.
Privacy and security are separate obligations. When electronic protected health information is involved, HIPAA requires covered entities and business associates to perform risk analysis and risk management, review system activity, respond to incidents, evaluate safeguards, and obtain required business-associate assurances [41]. Those controls do not validate model accuracy, fairness, or clinical benefit.
Vendors Supply Models; Hospitals Retain Governance
Vendor review should begin with evidence, not branding. Public regulatory records may omit clinically important details, and a majority of hospital predictive models in one national survey were supplied by electronic health record vendors [19-21,24]. Procurement should obtain the intended use, version, training and test populations, label definition, exclusions, external validation, calibration, subgroup results, operating thresholds, known failure modes, cybersecurity information, update process, monitoring capability, and supporting publications. Missing evidence should be treated as uncertainty, not as proof of poor performance or as permission to proceed.
The contract should preserve operational control. Required terms depend on the transaction, but leaders should address data rights, permitted use, subcontractors, security, performance documentation, audit access, change notice, version retention, incident notification, regulatory cooperation, service levels, continuity, suspension, termination, data return or destruction, insurance, and allocation of financial risk. An indemnity clause may affect cost allocation. It does not transfer the hospital’s control of workflow or relieve the clinical owner from evaluating patient impact.
Change provisions are especially important. A generative service can change model weights, system prompts, retrieval sources, filtering, or user interface without looking like a traditional software upgrade. A predictive service can change through recalibration, code, input mapping, threshold, or data feed. The hospital needs advance notice, a materiality test, validation rights, and the ability to defer or roll back a change. For an authorized AI-enabled device, a predetermined change control plan may describe planned modifications and their validation. FDA’s 2025 guidance is nonbinding, and it does not authorize a hospital to alter a device outside the authorized plan [36].
Hospital governance should also use, but not overstate, available transparency requirements. The ONC decision-support-intervention criterion applies to developers of certified health IT modules and includes source-attribute transparency and intervention risk management for covered predictive interventions [39]. It does not directly regulate every hospital AI use. The disclosures can support procurement, but they do not replace local validation.
Vendor dependence should appear on the risk register. If the hospital cannot access a prior version, reproduce an output, export monitoring data, or continue care during interruption, it may lack practical control despite formal stop authority.
Monitoring Must Detect Drift, Misuse, Unequal Harm, and Lost Benefit
A model can remain technically available while becoming clinically unreliable. Patient mix, disease prevalence, documentation, laboratory methods, devices, coding, staffing, referral patterns, capacity, or treatment can change. Discrimination may appear stable while calibration deteriorates. A threshold can create a growing workload even if the underlying model is unchanged. Monitoring must therefore examine the data, model, human response, downstream care, and outcomes.
Pandemic-era research illustrates these dimensions. An emergency-department admission model showed a decline in discrimination during COVID-19, while feature-attribution monitoring helped characterize changing inputs [28]. In a study of hospitalized patients with COVID-19, calibration drift occurred even when discrimination changed less; dynamic updating improved retrospective calibration and decision-curve performance [29]. A simulated deployment across seven hospitals found that label-agnostic methods detected demographic, site, admission-source, and laboratory shifts, and that drift-triggered updating improved retrospective performance [30]. These studies do not prove that automatic updating improves patient outcomes. They show that waiting for mature outcome labels can leave a surveillance gap.
The monitoring contract should name measures, owners, frequency, thresholds, and actions. At minimum, it should include use volume, input quality, missingness, population shift, calibration, threshold performance, false-positive and false-negative patterns, subgroup results, alert burden, response time, overrides, downstream service receipt, adverse events, complaints, and benefit measures tied to the original purpose. Large language models also require prompt, retrieval-source, output, and version audits. A source-verification rule is essential because one repeated literature-search experiment found that most citations initially produced by a language model were fabricated or misleading blends [34].
Incident surveillance has limitations. An analysis linking FDA-authorized AI or machine-learning devices with MAUDE reports found strong concentration in a few products and substantial missing context [31]. MAUDE lacks exposure denominators and cannot estimate incidence or causation. Hospitals should not wait for a national signal or interpret report counts as rates. Their own patient-safety system should accept suspected model, interface, data, workflow, and human-AI failures.
Monitoring must lead to decisions. A threshold breach should trigger investigation, restriction, revalidation, rollback, suspension, or retirement according to predefined authority. Retraining is not an automatic remedy. It can introduce overfitting, feedback loops, catastrophic forgetting, and regulatory change concerns [30,36]. The hospital should compare the changed service with the last approved version and current care before release.
Labels often arrive later than the decision. Early surveillance can use input distributions, missingness, service volume, response patterns, complaints, and sentinel-case review, but these are not substitutes for outcome evaluation. Leaders should specify how delayed ground truth will be reconciled and how false reassurance will be avoided when discrimination appears stable. Calibration, threshold burden, and subgroup results should remain visible even when an aggregate dashboard stays green.
The drift signal tower
Monitoring must connect a detectable signal to a named action and decision owner.
Decide.
Act.A green aggregate metric cannot hide calibration loss, unequal harm, or a changed workflow.
Incident Response, Change Control, and Retirement Close the Loop
A clinical-AI incident may begin as a wrong prediction, hallucinated statement, missed alert, biased allocation, interface error, delayed response, unexplained version change, or complaint. The immediate priority is patient care and containment. The organization should identify affected use cases, restrict or suspend the service when ongoing harm is plausible, preserve the relevant model and configuration, secure logs, notify clinical owners, and activate patient-safety review. A broad outage plan is not enough when the model appears available but produces unsafe output.
Investigation should reconstruct the care pathway. The review should distinguish an input defect, model failure, out-of-scope use, display problem, misunderstanding, workflow constraint, capacity failure, or inappropriate reliance. It should examine both the index event and whether similarly exposed patients require review. Clinical, technical, human-factors, equity, privacy, security, vendor, risk, and legal expertise should participate according to the event, while one owner remains accountable for closure.
Reporting analysis must be product- and event-specific. Under 21 CFR Part 803, device user facilities have mandatory reporting duties for certain suspected device-related deaths and serious injuries, generally within the applicable ten-work-day period. The rule applies only when the AI product is a medical device and the event meets the reporting criteria [42]. Privacy, security, research, payer, contract, state, and accreditation obligations may create separate analyses. The article cannot determine those duties for an unspecified incident.
Change control should classify modifications to model weights, code, input mapping, data source, threshold, prompt, retrieval source, interface, user, workflow, and intended population. Material changes return to validation. The FDA’s final predetermined-change-control-plan guidance provides a framework for planned modifications to authorized AI-enabled device software, while its broader January 2025 lifecycle guidance remained draft in the reviewed source set [36,37]. Good Machine Learning Practice principles offer nonbinding direction on representative data, independent testing, human-AI team performance, user information, and postdeployment monitoring [38].
Retirement is a clinical transition. Leaders should identify replacement workflows, patient-continuity risk, record retention, contract closure, data disposition, residual alerts, and the date authority ends. A model that has lost benefit, lacks a supported version, or cannot be monitored should not remain active merely because no single catastrophic event has occurred.
Corrective action should verify effectiveness. Closing an incident after training, a vendor patch, or a threshold change is premature until the organization tests whether recurrence risk fell and whether a new burden emerged.
An Executive Clinical AI Assurance System
Clinical AI requires an operating system rather than a committee that reviews proposals once. Systematic reviews of real-world evaluation and governance frameworks identify recurring needs across workflow integration, risk management, validation, monitoring, accountability, safety, transparency, and long-term impact [26,27]. The frameworks are not a proven universal design, and the underlying governance literature varies in quality. They nevertheless converge with implementation evidence showing that role clarity, infrastructure, frontline design, and workflow ownership matter [22,23].
The governing body should assign one accountable executive and define delegated authorities for clinical ownership, validation, data stewardship, equity, privacy, security, legal review, deployment, monitoring, incidents, and retirement. Routine low-risk uses can move through standardized pathways. Novel or high-consequence uses should escalate when they influence diagnosis or treatment, allocate scarce resources, operate with limited human review, use sensitive data, affect vulnerable populations, or cannot be independently evaluated.
The operating rhythm should include a current use-case registry, predeployment gates, regular performance and equity review, prompt incident escalation, and periodic independent challenge. The registry should not become an inventory that nobody can act upon. Each entry needs a purpose, owner, version, evidence decision, monitoring status, open issue, and stop authority. Metrics should connect to patient care. A count of models or completed reviews does not demonstrate benefit.
Board assurance should answer five questions. Does each clinical use have named authority? Did the hospital validate the deployed version in the relevant environment? Can leaders detect unequal error, unsafe reliance, drift, and lost benefit? Can the organization reconstruct and correct an incident? Can it suspend or retire the service without disrupting care? The board should receive trends, material exceptions, unresolved risk, corrective-action aging, and decisions requiring resources or risk acceptance.
Management should separate three types of evidence. Compliance evidence shows that a required review or control occurred. Performance evidence shows how the human-AI service operated. Benefit evidence shows whether the approved objective improved without unacceptable harm. A complete file can still support a poor service, and a strong retrospective model can still lack clinical benefit. The scorecard should keep those categories distinct.
Independent challenge should be risk based. Internal audit, patient safety, clinical quality, equity, privacy, security, or outside expertise may test whether the registered purpose matches actual use, whether monitoring can reproduce vendor reports, and whether leaders acted on unfavorable results. Reviewers should be able to reach source evidence and escalate without approval from the use-case sponsor. Findings should return to policy, procurement, validation, workflow, and training rather than end as a presentation.
The assurance program also needs capacity. Validation staff and clinical leaders cannot review an unlimited portfolio. The board should see the queue, risk tier, overdue reviews, unsupported uses, and resource constraints.
| Domain | Primary evidence | Balancing measure | Executive escalation condition |
|---|---|---|---|
| Inventory and authority | Active use cases with purpose, owner, model role, version, and approval status | Time from proposal to a defensible decision | Unregistered clinical use, missing owner, or use outside approved scope |
| Validation | Local calibration and threshold performance; comparator; silent or controlled test where appropriate | Evaluation burden and delay to beneficial use | Material performance gap, uncertain benefit, or no local evidence |
| Subgroup performance | Error, calibration, allocation, and response by relevant groups with uncertainty | Avoid unstable inference from small samples | Persistent or unexplained disparity, missing remedy owner |
| Human-factors reliability | Usability results, reliance errors, overrides, escalation, training scenarios | Cognitive load, alert burden, workflow interruption | Unsafe default, ineffective oversight, repeated misinterpretation |
| Clinical benefit | Outcome or process measure tied to the approved purpose | Unintended care, delay, workload, cost, and opportunity effects | Lost benefit, worsening balancing measure, no action on output |
| Drift and monitoring | Input shift, calibration, threshold, subgroup, use, response, and version trends | Monitoring false alarms and review burden | Trigger breach, unexplained change, missing label or audit data |
| Incidents and complaints | Events by failure mode, severity, affected patients, recurrence, and closure | Reporting access and just-culture measures | Possible ongoing harm, repeat event, missing evidence, overdue action |
| Vendor change | Notices, new versions, validation status, service and security events | Continuity and vendor dependence | Unapproved material change, unsupported version, blocked audit access |
| Suspension and retirement | Stop tests, rollback readiness, retired uses, continuity plans | Disruption from premature withdrawal | Stop criterion met without action, no safe fallback |
| Corrective-action closure | Owner, due date, effectiveness review, and residual risk | Workforce capacity and duplicated controls | Overdue high-risk action or ineffective remediation |
The annual assurance docket
Board oversight should review evidence, exceptions, corrective action, and retirement across a recurring cycle.
Strengths and Limitations
This review integrates peer-reviewed evidence on validation, human-AI interaction, bias, hospital adoption, drift, generative systems, and governance with a separately identified set of current United States legal and regulatory authorities. It translates source-specific findings into decision rights and retained evidence while avoiding the claim that a model, clinician, vendor, or hospital always carries liability.
Important limitations remain. The literature is heterogeneous and includes surveys, laboratory and vignette experiments, retrospective external validations, regulatory-document studies, implementation reports, and systematic reviews with different outcomes. Many studies evaluate selected tasks, institutions, model versions, or interfaces and do not measure patient outcomes. Prospective evidence of long-term benefit and harm remains limited. Public regulatory documents may omit information reviewed privately. Generative models change rapidly, making findings version- and prompt-dependent. Fairness definitions and subgroup data are inconsistent, and small samples can create unstable estimates.
The search was purposive rather than systematic. One author conducted screening and extraction, no formal risk-of-bias instrument was used, and no meta-analysis was performed. The legal analysis is limited to selected federal authorities and does not survey state malpractice, informed-consent, corporate-negligence, evidence, privacy, or product-liability law. Guidance and regulation continue to evolve. This article provides an executive governance synthesis, not legal or patient-specific clinical advice.
Conclusions
Clinical AI does not remove accountability; it redistributes decisions across a system the hospital must govern. Every use case needs a defined purpose, named authority, local evidence, safe human interaction, equity testing, monitoring, incident response, change control, and enforceable stop rule. Regulatory authorization and vendor contracts can inform assurance but cannot replace it. The model may advise. The hospital must preserve judgment, proof, and remedy.
The defensible objective is not perfect prediction or a guarantee against liability. It is a reliable process that identifies uncertainty, limits unsupported use, detects deterioration, corrects harm, and preserves an accountable human and organizational path for every material clinical decision.
Acknowledgments
None.
Publication Statements
Reporting Checklist: This article is presented in accordance with a narrative review reporting checklist.
Funding: None.
Conflicts of Interest: The author has completed the ICMJE uniform disclosure form. The author is President and Chief Executive Officer of The Healthcare Executive. No other conflicts of interest are declared.
Ethical Statement: The author is accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This narrative review did not involve human participants or animals; institutional review board approval and informed consent were not applicable.
Data Sharing Statement: No original datasets were generated or analyzed for this narrative review. The completed search strategy is reported in the manuscript and supplementary material.
Disclaimer: The views expressed are those of the author and are intended for executive education. This article does not constitute legal, regulatory, clinical, privacy, cybersecurity, insurance, reimbursement, accounting, or investment advice. Organizations should obtain advice specific to their facts, patients, technologies, and jurisdiction.
References
- Tobia K, Nielsen A, Stremitzer A. When does physician use of AI increase liability? J Nucl Med. 2021;62(1):17-21. doi:10.2967/jnumed.120.256032. PMID:32978285.
- Khullar D, Casalino LP, Qian Y, Lu Y, Chang E, Aneja S. Public vs physician views of liability for artificial intelligence in health care. J Am Med Inform Assoc. 2021;28(7):1574-1577. doi:10.1093/jamia/ocab055. PMID:33871009.
- Jabbour S, Fouhey D, Shepard S, et al. Measuring the impact of AI in the diagnosis of hospitalized patients: a randomized clinical vignette survey study. JAMA. 2023;330(23):2275-2284. doi:10.1001/jama.2023.22295. PMID:38112814.
- Gaube S, Suresh H, Raue M, et al. Do as AI say: susceptibility in deployment of clinical decision-aids. NPJ Digit Med. 2021;4(1):31. doi:10.1038/s41746-021-00385-9. PMID:33608629.
- Yu F, Moehring A, Banerjee O, Salz T, Agarwal N, Rajpurkar P. Heterogeneity and predictors of the effects of AI assistance on radiologists. Nat Med. 2024;30(3):837-849. doi:10.1038/s41591-024-02850-w. PMID:38504016.
- Wang DY, Ding J, Sun AL, et al. Artificial intelligence suppression as a strategy to mitigate artificial intelligence automation bias. J Am Med Inform Assoc. 2023;30(10):1684-1692. doi:10.1093/jamia/ocad118. PMID:37561535.
- Mello MM, Guha N. Understanding liability risk from using health care artificial intelligence tools. N Engl J Med. 2024;390(3):271-278. doi:10.1056/NEJMhle2308901. PMID:38231630.
- Rose SL, Shapiro D. An ethically supported framework for determining patient notification and informed consent practices when using artificial intelligence in health care. Chest. 2024;166(3):572-578. doi:10.1016/j.chest.2024.04.014. PMID:38788895.
- Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453. doi:10.1126/science.aax2342. PMID:31649194.
- Seyyed-Kalantari L, Zhang H, McDermott MBA, Chen IY, Ghassemi M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med. 2021;27(12):2176-2182. doi:10.1038/s41591-021-01595-0. PMID:34893776.
- Gichoya JW, Banerjee I, Bhimireddy AR, et al. AI recognition of patient race in medical imaging: a modelling study. Lancet Digit Health. 2022;4(6):e406-e414. doi:10.1016/S2589-7500(22)00063-2. PMID:35568690.
- Daneshjou R, Vodrahalli K, Novoa RA, et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci Adv. 2022;8(32):eabq6147. doi:10.1126/sciadv.abq6147. PMID:35960806.
- Pierson E, Cutler DM, Leskovec J, Mullainathan S, Obermeyer Z. An algorithmic approach to reducing unexplained pain disparities in underserved populations. Nat Med. 2021;27(1):136-140. doi:10.1038/s41591-020-01192-7. PMID:33442014.
- Omiye JA, Lester JC, Spichak S, Rotemberg V, Daneshjou R. Large language models propagate race-based medicine. NPJ Digit Med. 2023;6(1):195. doi:10.1038/s41746-023-00939-z. PMID:37864012.
- Yang Y, Liu Y, Liu X, et al. Demographic bias of expert-level vision-language foundation models in medical imaging. Sci Adv. 2025;11(13):eadq0305. doi:10.1126/sciadv.adq0305. PMID:40138420.
- Wong A, Otles E, Donnelly JP, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med. 2021;181(8):1065-1070. doi:10.1001/jamainternmed.2021.2626. PMID:34152373.
- Wong A, Currey D, Schwinne M, et al. Multicenter prospective validation of an updated proprietary sepsis prediction model. JAMA Netw Open. 2026;9(2):e260181. doi:10.1001/jamanetworkopen.2026.0181. PMID:41758510.
- Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018;15(11):e1002683. doi:10.1371/journal.pmed.1002683. PMID:30399157.
- Wu E, Wu K, Daneshjou R, Ouyang D, Ho DE, Zou J. How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals. Nat Med. 2021;27(4):582-584. doi:10.1038/s41591-021-01312-x. PMID:33820998.
- Muralidharan V, Adewale BA, Huang CJ, et al. A scoping review of reporting gaps in FDA-approved AI medical devices. NPJ Digit Med. 2024;7(1):273. doi:10.1038/s41746-024-01270-x. PMID:39362934.
- Mehta V, Komanduri A, Bhadouriya RS, et al. Evaluating transparency in AI/ML model characteristics for FDA-reviewed medical devices. NPJ Digit Med. 2025;8(1):673. doi:10.1038/s41746-025-02052-9. PMID:41249460.
- Kashyap S, Morse KE, Patel B, Shah NH. A survey of extant organizational and computational setups for deploying predictive models in health systems. J Am Med Inform Assoc. 2021;28(11):2445-2450. doi:10.1093/jamia/ocab154. PMID:34423364.
- Sendak MP, Ratliff W, Sarro D, et al. Real-world integration of a sepsis deep learning technology into routine clinical care: implementation study. JMIR Med Inform. 2020;8(7):e15182. doi:10.2196/15182. PMID:32673244.
- Nong P, Adler-Milstein J, Apathy NC, Holmgren AJ, Everson J. Current use and evaluation of artificial intelligence and predictive models in US hospitals. Health Aff (Millwood). 2025;44(1):90-98. doi:10.1377/hlthaff.2024.00842. PMID:39761454.
- Everson J, Nong P, Richwine C. Uptake of generative AI integrated with electronic health records in US hospitals. JAMA Netw Open. 2025;8(12):e2549463. doi:10.1001/jamanetworkopen.2025.49463. PMID:41385223.
- Jacob C, Brasier N, Laurenzi E, et al. AI for IMPACTS framework for evaluating the long-term real-world impacts of AI-powered clinician tools: systematic review and narrative synthesis. J Med Internet Res. 2025;27:e67485. doi:10.2196/67485. PMID:39909417.
- Alami H, Sabio RP, Pérez EJ, et al. Artificial intelligence governance in health systems: systematic review of frameworks and integrative model proposal. J Med Internet Res. 2026;28:e87448. doi:10.2196/87448. PMID:42258811.
- Duckworth C, Chmiel FP, Burns DK, et al. Using explainable machine learning to characterise data drift and detect emergent health risks for emergency department admissions during COVID-19. Sci Rep. 2021;11(1):23017. doi:10.1038/s41598-021-02481-y. PMID:34837021.
- Levy TJ, Coppa K, Cang J, et al. Development and validation of self-monitoring auto-updating prognostic models of survival for hospitalized COVID-19 patients. Nat Commun. 2022;13(1):6812. doi:10.1038/s41467-022-34646-2. PMID:36357420.
- Subasri V, Krishnan A, Kore A, et al. Detecting and remediating harmful data shifts for the responsible deployment of clinical AI models. JAMA Netw Open. 2025;8(6):e2513685. doi:10.1001/jamanetworkopen.2025.13685. PMID:40465297.
- Babic B, Cohen IG, Stern AD, Li Y, Ouellet M. A general framework for governing marketed AI/ML medical devices. NPJ Digit Med. 2025;8(1):328. doi:10.1038/s41746-025-01717-9. PMID:40450160.
- Goh E, Gallo R, Hom J, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Netw Open. 2024;7(10):e2440969. doi:10.1001/jamanetworkopen.2024.40969. PMID:39466245.
- Hager P, Jungmann F, Holland R, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med. 2024;30(9):2613-2622. doi:10.1038/s41591-024-03097-1. PMID:38965432.
- McGowan A, Gui Y, Dobbs M, et al. ChatGPT and Bard exhibit spontaneous citation fabrication during psychiatry literature search. Psychiatry Res. 2023;326:115334. doi:10.1016/j.psychres.2023.115334. PMID:37499282.
- United States. Federal Food, Drug, and Cosmetic Act, 21 U.S.C. § 360j(o)(1)(E); U.S. Food and Drug Administration. Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff. January 2026. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software.
- U.S. Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions: Guidance for Industry and Food and Drug Administration Staff. August 2025. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence.
- U.S. Food and Drug Administration. Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations. Draft Guidance. January 2025. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/artificial-intelligence-enabled-device-software-functions-lifecycle-management-and-marketing.
- U.S. Food and Drug Administration. Good Machine Learning Practice for Medical Device Development: Guiding Principles. 2025. https://www.fda.gov/medical-devices/software-medical-device-samd/good-machine-learning-practice-medical-device-development-guiding-principles.
- Office of the National Coordinator for Health Information Technology. Health Data, Technology, and Interoperability: Certification Program Updates, Algorithm Transparency, and Information Sharing Final Rule; 45 C.F.R. § 170.315(b)(11). https://healthit.gov/regulations/hti-rules/hti-1-final-rule/.
- Nondiscrimination in Health Programs and Activities, 45 C.F.R. § 92.210.
- Security Standards: Administrative Safeguards, 45 C.F.R. § 164.308.
- Medical Device Reporting, 21 C.F.R. part 803; U.S. Food and Drug Administration. Mandatory Reporting Requirements: Manufacturers, Importers and Device User Facilities. https://www.fda.gov/medical-devices/postmarket-requirements-devices/mandatory-reporting-requirements-manufacturers-importers-and-device-user-facilities.
Supplementary Table S1. Detailed PubMed/MEDLINE search strategy
The following targeted blocks were rerun through NCBI PubMed ESearch on August 13, 2026. Raw counts are query-specific and nonadditive because records overlap across blocks. Counts will change as PubMed indexing changes.
| Search block | Exact PubMed syntax | Filters and dates | Raw results | Selection use |
|---|---|---|---|---|
| Liability and malpractice | (("artificial intelligence"[ti] OR "machine learning"[ti]) AND (liability[ti] OR malpractice[ti]) AND ("2020/01/01"[dp] : "2026/08/13"[dp])) | English-language peer-reviewed relevance review; January 1, 2020 to August 13, 2026 | 48 | Liability perceptions and peer-reviewed legal analysis |
| Human oversight and automation bias | (("clinical artificial intelligence"[tiab] OR "medical artificial intelligence"[tiab] OR "AI assistance"[tiab] OR "AI-assisted"[tiab]) AND ("automation bias"[tiab] OR "human-AI"[tiab] OR explanation*[tiab] OR reliance[tiab]) AND (clinician*[tiab] OR physician*[tiab] OR radiologist*[tiab]) AND ("2020/01/01"[dp] : "2026/08/13"[dp])) | English-language peer-reviewed relevance review; January 1, 2020 to August 13, 2026 | 135 | Human-AI performance, reliance, explanations, and suppression controls |
| Bias, equity, and disparities | (("artificial intelligence"[ti] OR "machine learning"[ti] OR "deep learning"[ti] OR "foundation model"[ti]) AND ("racial bias"[ti] OR "demographic bias"[ti] OR disparit*[ti] OR underdiagnosis[ti] OR equity[ti] OR fairness[ti]) AND ("2020/01/01"[dp] : "2026/08/13"[dp])) | English-language peer-reviewed relevance review; January 1, 2020 to August 13, 2026 | 406 | Target, proxy, subgroup, intersectional, and remediation evidence |
| External validation and regulatory evidence | (("clinical artificial intelligence"[tiab] OR "medical AI"[tiab] OR "AI/ML"[tiab]) AND ("external validation"[ti] OR generalizab*[ti] OR "local validation"[ti] OR FDA[ti] OR regulator*[ti]) AND ("2020/01/01"[dp] : "2026/08/13"[dp])) | English-language peer-reviewed relevance review; January 1, 2020 to August 13, 2026 | 72 | Generalizability, local thresholds, and public regulatory-evidence gaps |
| Health-system governance, implementation, and adoption | (("artificial intelligence"[ti] OR AI[ti]) AND (governance[ti] OR implementation[ti] OR adoption[ti]) AND ("health system"[tiab] OR "health systems"[tiab] OR hospital*[tiab]) AND ("2020/01/01"[dp] : "2026/08/13"[dp])) | English-language peer-reviewed relevance review; January 1, 2020 to August 13, 2026 | 227 | Organizational models, implementation, adoption, and assurance frameworks |
| Drift, data shift, and updating | (("machine learning"[ti] OR "artificial intelligence"[ti] OR AI[ti]) AND ("data drift"[tiab] OR "model drift"[tiab] OR "dataset shift"[tiab] OR "data shift"[tiab] OR "concept drift"[tiab] OR auto-updat*[tiab]) AND (clinical[tiab] OR hospital*[tiab] OR healthcare[tiab]) AND ("2020/01/01"[dp] : "2026/08/13"[dp])) | English-language peer-reviewed relevance review; January 1, 2020 to August 13, 2026 | 75 | Monitoring, calibration drift, updating, rollback, and change control |
| LLM clinical risk | (("large language model"[ti] OR "large language models"[ti] OR ChatGPT[ti]) AND ("clinical decision"[ti] OR hallucination*[ti] OR fabrication[ti] OR bias[ti] OR safety[ti] OR reasoning[ti] OR citation*[ti]) AND ("2022/01/01"[dp] : "2026/08/13"[dp])) | English-language peer-reviewed relevance review; January 1, 2022 to August 13, 2026 | 642 | Clinical reasoning, hallucination, version, prompt, and citation-verification risk |
| Consent, notification, privacy, and cybersecurity | (("artificial intelligence"[ti] OR AI[ti]) AND (consent[ti] OR "patient notification"[ti] OR privacy[ti] OR cybersecurity[ti]) AND (healthcare[tiab] OR clinical[tiab] OR hospital*[tiab]) AND ("2020/01/01"[dp] : "2026/08/13"[dp])) | English-language peer-reviewed relevance review; January 1, 2020 to August 13, 2026 | 149 | Patient-facing ethical framework and boundary with privacy or security controls |
The final synthesis contains 34 unique peer-reviewed PubMed-indexed publications. This was a targeted narrative review with overlapping thematic searches, known-item verification, publisher lookup, and citation chaining. Per-stage deduplicated counts, title-abstract-screened counts, full-text-screened counts, exclusion reasons, and citation-chain additions were not prospectively logged and therefore are not reported. The raw counts above are insufficient for a PRISMA flow diagram.

