1Introduction

Generative artificial intelligence (AI), including large language and multimodal models, can produce text, images, code, summaries, classifications, and conversational responses from natural-language instructions. In hospitals, the same general-purpose model may be embedded in ambient documentation, inbox management, discharge instructions, call-center support, revenue-cycle workflows, policy search, clinical decision support, staffing analytics, or software development. The World Health Organization (WHO) has emphasized both the potential of AI to improve health and the need to protect autonomy, safety, privacy, transparency, accountability, and equity (1). Its later guidance on large multimodal models highlights risks particular to generative systems, including inaccurate or biased output, automation bias, cybersecurity exposure, and uses that change faster than conventional oversight processes (2).

The governance challenge is not simply whether a model is accurate on a benchmark. A hospital deploys a sociotechnical system: a model and prompt configuration, vendor infrastructure, local data interfaces, user training, workflow rules, escalation paths, and organizational incentives. NIST’s AI Risk Management Framework describes governance, mapping, measurement, and management as linked functions rather than one-time compliance tasks (3). Its generative-AI profile adds risks such as confabulation, information integrity, data privacy, human-AI configuration, and value-chain dependence (4). These concepts are directly relevant to hospitals, where a plausible but wrong sentence can become a clinical instruction, an incorrect code, a denied claim, or a misleading patient message.

Early research demonstrates capability without resolving enterprise reliability. GPT-4 showed substantial medical knowledge, yet experts cautioned that fluent performance does not establish clinical safety or accountability (5). Large language models have performed strongly on medical question-answering benchmarks (6) and generated patient-facing responses rated favorably in a cross-sectional comparison (7). At the same time, benchmark success may not transfer to local populations, live workflows, changing models, or tasks requiring current facts and causal judgment. A randomized clinical-vignette study found that access to a large language model did not significantly improve physicians’ diagnostic reasoning compared with conventional resources, despite strong model-only performance (8). The operational conclusion is not that generative AI should be rejected, but that its use should be bounded, evaluated, and monitored in context.

Many hospitals have pilots but lack a repeatable route from idea to retirement. Fragmented review invites duplicate purchases, unapproved use, inconsistent contracts, weak evidence, and local workarounds. Overcentralized review can be equally ineffective if it delays low-risk uses without adding meaningful protection. This review offers an executive operating model that calibrates scrutiny to consequence while retaining a common lifecycle, control vocabulary, and accountability structure. We present this article in accordance with the narrative review reporting checklist.

2Methods

This narrative review was designed to translate current evidence and authoritative frameworks into an enterprise governance model for US hospitals. Searches were completed on 12 August 2026. PubMed/MEDLINE-indexed records were identified using combinations of terms for generative AI, large language models, hospitals, governance, implementation, safety, evaluation, bias, privacy, and monitoring. Targeted searches of official websites were conducted for WHO, NIST, the Office of the National Coordinator for Health Information Technology (ONC), the US Food and Drug Administration (FDA), and the US Department of Health and Human Services (HHS). Targeted searches also identified implementation and assurance materials from healthcare professional and multistakeholder organizations. Backward citation screening was used for seminal governance, bias, implementation, and reporting sources.

Priority was given to English-language empirical studies, consensus or reporting guidance, government standards, and detailed implementation reports published from 2019 onward. Earlier sources were retained when foundational. Sources were included when they addressed a hospital-relevant use, risk, evaluation method, governance process, human factor, or lifecycle control. Opinion pieces were used sparingly to frame emerging issues and not as proof of effectiveness. Purely technical model-development papers without a clear operational implication, promotional claims lacking methods, and studies whose relevance depended only on a brand name were excluded. Because this was a narrative rather than systematic review, no meta-analysis, formal risk-of-bias score, or exhaustive study-count flow diagram was undertaken. The completed search approach is summarized in Table 1.

The completed evidence-synthesis exhibits are presented in Supplementary Table S1.

3The enterprise problem: one technology, many consequences

3.1Govern the use case, not the marketing category

“Generative AI” is too broad to function as a risk class. Drafting a meeting agenda from nonsensitive inputs, summarizing a patient encounter, generating a response to a portal message, and proposing an oncology treatment plan may use related model architectures but create radically different hazards. Governance should therefore identify the precise intended use, user, decision influenced, data involved, population affected, workflow location, and plausible failure modes. A system that only drafts text for expert review is not automatically low risk: time pressure, interface design, copy-forward behavior, or productivity targets can turn nominal review into approval by default.

Hospital leaders should classify each use by consequence rather than novelty. A practical taxonomy distinguishes assistive administrative uses; operational recommendations affecting resources or access; patient-facing communication; documentation that enters the legal medical record; and diagnostic, prognostic, or treatment-related functions. The risk tier should rise when output can directly change care, when errors are difficult to detect, when vulnerable populations may be affected differently, when protected information leaves the controlled environment, or when the model can act without a meaningful human checkpoint. WHO guidance supports task-specific assessment and stakeholder participation rather than reliance on generalized model claims (2).

Capability evidence must also match the local task. Medical benchmark performance (6) does not validate a model for a hospital’s abbreviations, formularies, policies, patient literacy needs, or EHR configuration. Models can generate polished but false content, a behavior amplified when users equate fluency with knowledge. Discharge-summary generation illustrates the tradeoff: summarization can reduce clerical work, but omission, temporal confusion, or unsupported instruction can cause downstream harm (9). The appropriate question is therefore: under defined conditions, does this configured system improve a specified workflow without unacceptable safety, equity, privacy, or financial consequences?

Executive visual 01

Govern by consequence, not novelty

Scrutiny rises with patient impact, autonomy, data sensitivity, scale, and the difficulty of detecting failure.

Tier 1Assistive

Nonsensitive drafting, search, and administrative support.

Standard controls and accountable review
Tier 2Consequential

Operations, access, patient communication, and legal-record documentation.

Local validation and active monitoring
Tier 3High consequence

Diagnosis, prognosis, treatment, or autonomous action.

Independent review, stronger evidence, pause authority
Editorial synthesis of the article’s risk-tiering framework.

3.2Create a single portfolio view

Shadow use flourishes when approval is slow or unclear. Hospitals need a simple intake route and an enterprise inventory that includes purchased tools, embedded vendor features, locally built applications, research systems transitioning into operations, and material employee use of public models. Each record should identify an executive sponsor, operational owner, technical owner, clinical safety owner when applicable, vendor and model versions, data classes, interfaces, approved users, risk tier, validation status, monitoring plan, contract renewal date, and retirement status.

The inventory should track dependencies, not merely products. A front-end assistant may rely on a cloud platform, foundation-model provider, retrieval database, third-party monitoring service, and EHR integration. Changes anywhere in that chain may alter output. Research on extraction of memorized training data and broader scholarship on large language models demonstrate why data provenance and downstream behavior cannot be assumed from interface appearance alone (10,11). Contracts should therefore require timely notice of material model, hosting, subprocessor, data-use, or security changes and preserve the hospital’s ability to suspend use.

4An accountable governance operating model

4.1Board and executive responsibilities

The board need not approve every model, but it should oversee AI as an enterprise risk and value domain. At minimum, board reporting should cover the portfolio by risk tier, material incidents and near misses, high-risk approvals, control exceptions, cybersecurity and privacy posture, equity findings, value realization, and unresolved vendor dependencies. The chief executive should designate an accountable executive and ensure that incentives do not reward adoption faster than evidence.

A standing AI oversight committee can coordinate work that otherwise fragments across innovation, information technology, clinical informatics, quality, risk, privacy, compliance, legal, cybersecurity, procurement, human resources, finance, and operations. Patient or community representation is particularly important for patient-facing or access-related uses. The committee should set policy, approve risk criteria, adjudicate high-risk cases, review incidents, and publish standard evidence requirements. It should not replace subject-matter owners. Operational leaders remain responsible for whether the use solves a real problem; clinical leaders remain responsible for safe workflow integration; technical teams remain responsible for system integrity; and vendors remain accountable for their contractual representations.

Documented decision rights prevent “governance theater.” One named person should be accountable for each use case. Reviewers may recommend, impose conditions, or reject within defined authority. Emergency pause authority should be explicit and available outside routine meeting cycles. Internal algorithmic auditing frameworks emphasize traceability across the lifecycle and the preservation of artifacts explaining why decisions were made (12). Minutes alone are insufficient; the evidence package should include the intended-use statement, hazard analysis, data-flow map, validation results, usability findings, residual risks, approvals, monitoring thresholds, and change history.

Executive visual 02

Accountability must connect oversight to frontline use

Each layer owns a distinct decision. None substitutes for the others.

01Board oversightRisk appetite, material exposure, performance, and unresolved exceptions
02Accountable executivePortfolio ownership, resources, incentives, and enterprise pause authority
03AI oversight committeeRisk tiers, evidence standards, approvals, incidents, and change control
04Use-case owner and frontline teamsWorkflow fit, human authority, monitoring, escalation, and safe retirement
Required decision recordIntended use → evidence → residual risk → accountable approval → next review date
Operating-model synthesis based on the governance responsibilities described in this review.
People-free executive planning model showing layered AI governance controls protecting a connected hospital system.
Governance in practiceLayered controls connect executive accountability, operational assurance, and local evidence across the hospital enterprise.

4.2Equity is a performance requirement

Bias is not confined to model training. It may arise from labels, missing data, deployment rules, uneven digital access, language quality, differential clinician reliance, or the operational choice of whom to target. A widely used population-health algorithm underestimated the needs of Black patients because cost was used as a proxy for illness (13). That lesson applies to generative AI: seemingly neutral targets can encode structural inequity. Local validation should examine meaningful subgroups where data permit, but subgroup analysis must avoid false assurance from small samples. Qualitative review by users and affected communities can reveal harms that aggregate metrics miss.

Equity review should ask whether the system changes access, waiting, denial, escalation, surveillance, or communication; whether supported languages are genuinely usable; whether disability access is preserved; and whether an alternative pathway exists. A model should not be used to create an inferior service tier for people who decline automation. When performance gaps cannot be mitigated, use should be narrowed or stopped. Recommended decision rights are summarized in Table 2.

5Lifecycle controls: from intake to retirement

Executive visual 03

A governed lifecycle from intake to retirement

Approval is one checkpoint inside a continuous control system.

  1. 01DefineProblem, baseline, intended use, owner, risk tier
  2. 02AssureEvidence, vendor review, privacy, security, contract
  3. 03ValidateLocal cases, human factors, subgroups, error taxonomy
  4. 04DeployControlled gates, training, fallback, escalation
  5. 05Monitor or retireDrift, incidents, value, change control, safe exit
MonitorRestrictPauseRedesignRetire
Lifecycle control model synthesized from the article’s five implementation stages.

5.1Stage 1: problem definition and proportional triage

Every proposal should begin with a workflow problem and baseline, not a product demonstration. The sponsor should specify who experiences the problem, current performance, why non-AI alternatives are inadequate, the expected mechanism of benefit, and possible unintended consequences. This prevents organizations from automating waste or creating a new tool where standardized work, staffing, or interface repair would be more effective.

Risk triage should consider severity, scale, reversibility, detectability, autonomy, data sensitivity, affected populations, and regulatory status. High severity with poor detectability should trigger the strongest review even if exposure is initially small. Low-risk experimentation should still prohibit the entry of protected or confidential information into unapproved systems and should clearly label outputs not intended for clinical use.

5.2Stage 2: evidence and vendor due diligence

Hospitals should request evidence for the configured product and intended use, not a related model or carefully selected demonstration. Due diligence should cover model identity and versioning, training and evaluation data at an appropriate level, known limitations, local configuration, retrieval sources, data retention, secondary data use, subprocessors, hosting, encryption, access control, audit logs, incident notification, business continuity, portability, intellectual-property allocation, and termination assistance. Privacy review should map every input, intermediate representation, log, and output. NIST’s Privacy Framework offers a complementary structure for managing privacy risk across systems and data processing (14).

Regulatory status must be established but should not be mistaken for enterprise readiness. Some AI-enabled functions may meet the definition of a medical device; others may fall within non-device decision support or administrative software. FDA and international partners have articulated good machine-learning-practice principles, including multidisciplinary expertise, representative data, human-AI interaction, clinically relevant testing, clear user information, and monitoring of deployed models (15). ONC’s health IT rules also advance transparency for predictive decision-support interventions distributed through certified health IT (16). Legal counsel should determine applicability to the actual configuration and workflow.

5.3Stage 3: local validation and human-factors evaluation

Validation should use representative local cases, realistic prevalence, current policies, and the actual user interface. The comparator should reflect existing practice. Endpoints should include task performance and downstream consequences: omission, unsupported content, calibration where relevant, correction burden, time, cognitive load, escalation, communication quality, and disparate performance. Evaluators should deliberately test rare but severe hazards, adversarial prompts, conflicting records, ambiguous temporal information, copy-forward errors, and unavailable source material.

Human review is a control only when it is feasible and observable. A reviewer needs appropriate expertise, time, access to source information, a clear responsibility to change output, and an interface that makes uncertainty and provenance visible. “Human in the loop” is otherwise a slogan. DECIDE-AI emphasizes live evaluation, human factors, workflow integration, and safety in early clinical deployment (17). For clinical trials and prediction models, CONSORT-AI, SPIRIT-AI, and TRIPOD+AI provide complementary transparency requirements (18-20). A hospital need not convert every quality-improvement pilot into a clinical trial, but it should borrow the discipline of describing the intervention, users, errors, version, data handling, and analysis.

5.4Stage 4: controlled deployment

Deployment should progress through explicit gates: sandbox; silent or retrospective evaluation; limited users or locations; supervised production; and broader release. Not every use needs every gate, but high-consequence functions should not jump directly from vendor demonstration to enterprise production. Training should explain the intended use, prohibited uses, limitations, verification steps, escalation, downtime behavior, and how to report a suspected incident. Competence checks are more meaningful than attendance records.

Workflow design should retain authoritative sources. If a model summarizes a policy, users should be able to open the controlled policy and see its effective date. If it drafts a patient message, the approving clinician should see relevant chart evidence and unresolved uncertainty. If it proposes codes, the coding professional must retain authority and an audit trail. Real-world implementation of a sepsis model showed that technical performance was insufficient without stakeholder alignment, workflow redesign, trust building, infrastructure, and training (21). Generative AI requires the same implementation discipline.

5.5Stage 5: continuous assurance, change control, and retirement

Monitoring should distinguish model, workflow, safety, equity, security, experience, and value signals. Output review can combine automated checks, stratified sampling, user reports, and targeted audits. Metrics should be interpreted together: faster note closure with more corrections, lower call time with more repeat contacts, or higher message volume with reduced comprehension is not success. Healthcare machine-learning guidance consistently emphasizes post-deployment surveillance and avoidance of preventable harm (22).

Thresholds and responses should be determined before launch. A material increase in unsupported statements, a privacy or security event, unexplained subgroup deterioration, loss of source traceability, vendor model substitution, or inability to perform monitoring may require restriction or pause. Model updates should trigger regression testing proportional to risk. Hospitals should preserve a validated fallback and regularly rehearse downtime or withdrawal for operationally critical uses.

Retirement is a governed stage. Contracts and technical plans should address data return or deletion, removal of credentials and interfaces, retention of required records, transition of users, replacement of generated content that remains operationally active, and evaluation of residual dependencies. Without retirement discipline, hospitals accumulate unsupported models and hidden operational risk. Table 3 consolidates the minimum evidence expected at each lifecycle stage.

6Measuring value without trading away safety

Generative AI business cases often begin with minutes saved. Time is important, but released time is not automatically realized value. A hospital should establish whether time is redeployed to patient care, reduces overtime, improves access, or merely shifts review and correction work. Total cost includes licenses, integration, security, validation, training, monitoring, incident management, workflow redesign, and exit costs. Productivity metrics should be paired with balancing measures. General machine-learning scholarship in medicine reinforces the need to connect technical performance to clinical context and implementation (23). The Coalition for Health AI’s assurance blueprint (24) and the American Medical Association’s augmented-intelligence resources (25) provide additional stakeholder-oriented tools that hospitals can adapt without substituting them for local evidence or accountable decisions.

An executive scorecard can use five domains. First, safety and quality: critical error rate, unsupported assertion rate, correction severity, near misses, and clinically meaningful process measures. Second, equity and access: performance and experience across relevant populations, language quality, accessibility, and escalation availability. Third, workforce and experience: time, cognitive load, perceived control, burnout-related measures, and work shifted across roles. Fourth, operations and finance: cycle time, throughput, rework, denial or appeal outcomes, avoided cost, and total cost of ownership. Fifth, trust and control: user reporting, traceability, update compliance, audit completion, and closure of corrective actions.

Executive visual 04

Measure value with balancing outcomes

Time saved is not enterprise value when risk, rework, inequity, or hidden burden rises.

01Safety and qualityCritical errors, unsupported assertions, near misses
02Equity and accessSubgroup performance, language, accessibility, escalation
03Workforce experienceNet time, cognitive load, control, shifted work
04Operations and financeCycle time, rework, throughput, total cost
05Trust and controlTraceability, audit completion, issue closure
Scale only whenmeasured benefit remains material after safety, equity, workforce, cost, and control are considered together.
Balanced-measure framework derived from the article’s executive scorecard.

Measurement should use a counterfactual where feasible. Phased rollout, matched comparison, interrupted time series, or randomized evaluation may be appropriate depending on risk and operational constraints. Self-reported satisfaction alone is weak evidence. Executives should predefine the minimum benefit needed to justify residual risk and continued expenditure. A system that remains safe but produces no material value should be retired; a system that produces value by shifting hidden burden or inequity should be redesigned. Table 4 translates these principles into an enterprise scorecard.

7A practical 12-month executive roadmap

Executive visual 05

The first 12 months build a repeatable capability

The objective is disciplined enterprise governance, not the largest number of deployments.

  1. Days 0–90Establish controlAccountability, acceptable use, intake, inventory, risk tiers, pause authority
  2. Months 4–6Build assuranceEvidence packet, approved environments, logging, testing, incident pathway
  3. Months 7–9Prove performanceBaseline comparison, subgroup review, usability, contract correction
  4. Months 10–12Govern the portfolioBoard reporting and decisions to scale, modify, pause, or retire
Implementation roadmap synthesized from the article’s first-year executive agenda.
People-free architectural model showing a controlled AI pilot expanding through staged connections across a hospital enterprise.
From pilot to enterpriseResponsible scale is a staged operating journey that expands only as evidence, controls, and management capacity mature.

In the first 90 days, leadership should designate accountability, issue an acceptable-use standard, create one intake channel, inventory current systems, and identify unapproved uses involving sensitive data. The organization should define risk tiers, pause authority, minimum contracting clauses, and a small set of monitoring measures. This initial control layer should enable safe low-risk work rather than freeze all activity.

Between months 4 and 6, the oversight body should review the highest-risk current uses, implement a standard evidence packet, and select a limited portfolio of pilots linked to strategic problems. Technical teams should establish approved environments, logging, identity controls, retrieval-source governance, and a repeatable test harness. Quality and operations teams should build an error taxonomy and incident pathway integrated with existing patient-safety processes.

Between months 7 and 9, the hospital should compare pilot performance with baseline, complete subgroup and usability review, and promote only those uses that meet predefined thresholds. Procurement should amend or replace contracts lacking change notice, auditability, data controls, or exit provisions. Workforce leaders should monitor whether review work and liability are being shifted to clinicians without capacity or training.

By month 12, executives should receive a portfolio report showing risk, outcomes, incidents, cost, and decisions to scale, modify, pause, or retire. The board should review material risks and whether management’s controls are functioning. The cycle then repeats. Generative AI governance is not a launch project; it is a permanent enterprise capability.

The portfolio review should force explicit comparison among competing uses. A high-visibility pilot should not continue merely because a senior sponsor favors it, and a modest administrative use should not be subjected to the same evidentiary burden as a tool that shapes diagnosis, treatment, or access. Reviewers should see the original problem, the current configuration, the population exposed, the evidence available, unresolved hazards, measured benefit, total operating cost, and the next decision date on one page. Exceptions should identify an accountable approver, compensating controls, an expiration date, and the evidence required for renewal. This discipline prevents temporary workarounds from becoming permanent policy.

Leaders should also plan for organizational learning. Incident reports, user corrections, patient complaints, help-desk themes, procurement findings, and monitoring signals should feed a common taxonomy even when no patient harm occurred. Quarterly review can distinguish isolated user error from a recurring interface, data, training, or model problem. Findings should change test cases, training, contracts, and deployment standards across the portfolio. When a use is retired, the organization should document why, preserve required records, revoke access, remove interfaces, confirm vendor data disposition, and communicate the replacement workflow. A safe exit is evidence of governance maturity rather than failure.

Finally, scale should be constrained by management capacity. Every additional use creates obligations for validation, version tracking, monitoring, incident response, retraining, and re-evaluation. Approving more systems than the organization can oversee produces a false appearance of innovation while accumulating unmanaged risk. Portfolio limits, standardized evidence packets, shared assurance services, and reusable technical controls can increase throughput without weakening review. The objective for the first year is therefore not the largest number of deployments, but a repeatable operating model that can make and defend consistent decisions.

8Strengths and limitations

This review integrates current health-specific guidance, cross-sector risk frameworks, implementation research, bias evidence, and reporting standards into a concrete hospital lifecycle. It distinguishes use-case consequence from technology label and connects governance to decision rights, procurement, human factors, monitoring, and retirement. The recommendations are designed to be usable across administrative, operational, and clinical applications.

The evidence base remains limited. Many generative-AI studies use benchmarks, vignettes, short evaluations, or rapidly superseded model versions. Publication bias and vendor opacity may overstate benefits and underdescribe failures. This narrative review was purposive rather than systematic, did not use a formal risk-of-bias instrument, and focused primarily on US hospital operations while drawing on international guidance. Legal and regulatory requirements change and must be confirmed for each use and jurisdiction. The proposed governance model should therefore be adapted, tested, and revised rather than treated as a validated universal standard.

9Conclusions

Hospitals should treat generative AI as a portfolio of evolving sociotechnical systems whose risks depend on the task, workflow, data, users, and consequences. Responsible scale requires named accountability, proportional risk tiers, defensible evidence, data and vendor controls, local validation, meaningful human authority, staged deployment, continuous assurance, and a rehearsed ability to pause or retire. The strongest executive posture is neither indiscriminate acceleration nor blanket prohibition. It is disciplined selection: use generative AI where it solves an important problem, require evidence proportional to consequence, measure value with balancing outcomes, and preserve patient safety, equity, privacy, workforce judgment, and organizational control.

Acknowledgments

None.

Footnote

Reporting Checklist: The author has completed the narrative review reporting checklist.

Funding: None.

Conflicts of Interest: The author has completed the ICMJE uniform disclosure form. The author is President and Chief Executive Officer of The Healthcare Executive. No other conflicts of interest are declared.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This narrative review did not involve human participants or animals; institutional review board approval and informed consent were not applicable.

Data Sharing Statement: No original datasets were generated or analyzed for this narrative review. The completed search strategy is reported in the manuscript and supplementary material.

References

  1. World Health Organization. Ethics and governance of artificial intelligence for health: WHO guidance. Geneva: World Health Organization; 2021.
  2. World Health Organization. Ethics and governance of artificial intelligence for health: guidance on large multi-modal models. Geneva: World Health Organization; 2024.
  3. Tabassi E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). Gaithersburg (MD): National Institute of Standards and Technology; 2023. NIST AI 100-1. doi:10.6028/NIST.AI.100-1.
  4. Autio C, Schwartz R, Dunietz J, et al. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. Gaithersburg (MD): National Institute of Standards and Technology; 2024. NIST AI 600-1. doi:10.6028/NIST.AI.600-1.
  5. Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388:1233-1239. doi:10.1056/NEJMsr2214184.
  6. Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023;620:172-180. doi:10.1038/s41586-023-06291-2.
  7. Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183:589-596. doi:10.1001/jamainternmed.2023.1838.
  8. Goh E, Gallo R, Hom J, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Netw Open. 2024;7:e2440969. doi:10.1001/jamanetworkopen.2024.40969.
  9. Patel SB, Lam K. ChatGPT: the future of discharge summaries? Lancet Digit Health. 2023;5:e107-e108. doi:10.1016/S2589-7500(23)00021-3.
  10. Bender EM, Gebru T, McMillan-Major A, et al. On the dangers of stochastic parrots: can language models be too big? In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. New York: Association for Computing Machinery; 2021. p. 610-623. doi:10.1145/3442188.3445922.
  11. Carlini N, Tramer F, Wallace E, et al. Extracting training data from large language models. In: 30th USENIX Security Symposium. Berkeley (CA): USENIX Association; 2021. p. 2633-2650.
  12. Raji ID, Smart A, White RN, et al. Closing the AI accountability gap: defining an end-to-end framework for internal algorithmic auditing. In: Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. New York: Association for Computing Machinery; 2020. p. 33-44. doi:10.1145/3351095.3372873.
  13. Obermeyer Z, Powers B, Vogeli C, et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366:447-453. doi:10.1126/science.aax2342.
  14. National Institute of Standards and Technology. NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, Version 1.0. Gaithersburg (MD): NIST; 2020.
  15. US Food and Drug Administration, Health Canada, Medicines and Healthcare products Regulatory Agency. Good Machine Learning Practice for Medical Device Development: Guiding Principles. Silver Spring (MD): FDA; 2021.
  16. Office of the National Coordinator for Health Information Technology. Health Data, Technology, and Interoperability: Certification Program Updates, Algorithm Transparency, and Information Sharing. Final rule. Fed Regist. 2024;89:1192-1338.
  17. Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28:924-933. doi:10.1038/s41591-022-01772-9.
  18. Liu X, Rivera SC, Moher D, et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med. 2020;26:1364-1374. doi:10.1038/s41591-020-1034-x.
  19. Cruz Rivera S, Liu X, Chan AW, et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med. 2020;26:1351-1363. doi:10.1038/s41591-020-1037-7.
  20. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378.
  21. Sendak MP, Ratliff W, Sarro D, et al. Real-world integration of a sepsis deep learning technology into routine clinical care: implementation study. JMIR Med Inform. 2020;8:e15182. doi:10.2196/15182.
  22. Wiens J, Saria S, Sendak M, et al. Do no harm: a roadmap for responsible machine learning for health care. Nat Med. 2019;25:1337-1340. doi:10.1038/s41591-019-0548-6.
  23. Rajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380:1347-1358. doi:10.1056/NEJMra1814259.
  24. Coalition for Health AI. Blueprint for Trustworthy AI Implementation Guidance and Assurance for Healthcare. Bedford (MA): Coalition for Health AI; 2023.
  25. American Medical Association. AMA Future of Health: The Emerging Landscape of Augmented Intelligence in Health Care. Chicago: American Medical Association; 2024.

Tables

Table 1. Search strategy summary

Decision Accountable role Required participants Minimum evidence
Approve low-risk pilot Operational executive IT, privacy/security, end users Intended use, data classification, test plan, owner, stop criteria
Approve high-consequence pilot Designated AI committee chair or executive risk authority Clinical safety, quality, legal/compliance, privacy, cybersecurity, informatics, patient representative as appropriate Hazard analysis, external evidence, local validation protocol, human-factors plan, monitoring and incident plan
Promote to production Use-case executive owner Same reviewers proportional to risk Completed validation, workflow acceptance, training, support readiness, signed residual-risk decision
Accept a material vendor/model change Technical owner and use-case owner; committee for high risk Procurement, security, clinical/operational reviewers Change description, regression test, updated risk assessment, contract compliance
Pause or disable Preauthorized safety, security, or operational leader Notify accountable executive and incident command Threshold breach, credible harm signal, data exposure, uncontrolled change, or loss of monitoring
Retire Use-case owner Records, legal, security, IT, finance Transition plan, data disposition, interface removal, archive and lessons learned

Table 3. Minimum lifecycle evidence package

Lifecycle gate Core artifacts Example go/no-go questions
Intake Problem statement, baseline, intended use, owner, initial risk tier Is there a measurable problem and a plausible benefit? Is a simpler intervention preferable?
Due diligence Data-flow diagram, vendor assessment, regulatory analysis, security/privacy review, contract controls Can the hospital identify where data go, who may use them, and what changes require notice?
Validation Protocol, representative test set, error taxonomy, subgroup analysis, usability results, residual-risk register Does evidence match the configured system, local users, and actual workflow?
Deployment Standard operating procedure, training and competence record, access configuration, help desk and escalation, fallback Can users verify outputs and respond safely when the system fails?
Monitoring Dashboard specification, audit sample, incident log, drift/change checks, benefit and burden measures Can the organization detect important deterioration before widespread harm?
Retirement Transition plan, data disposition, interface decommissioning, archive, lessons learned Can use stop without loss of essential operations, records, or accountability?

Table 4. Enterprise generative-AI scorecard

Domain Illustrative measures Escalation signal
Safety and quality Critical errors per 1,000 outputs; unsupported statements; correction severity; incident and near-miss rate Severe preventable error; rising high-severity correction rate; loss of required review
Equity and access Performance by relevant subgroup; language and accessibility testing; escalation completion Material unexplained performance gap or inferior pathway for a protected/vulnerable group
Workforce and experience Net time including review; cognitive load; user control; patient comprehension; repeat contacts Correction burden offsets benefit; users bypass verification; worsening comprehension
Operations and finance Cycle time; backlog; rework; total cost; avoided overtime; revenue-cycle quality Cost growth without benefit; improvement in one unit creates burden elsewhere
Trust and control Traceable sources; version inventory; audit completion; change notification; issue closure Unknown model change; missing logs; monitoring failure; overdue critical corrective action

Supplementary material

Supplementary Table S1. Reproducible search details

Source Search string or route Limits and notes
PubMed/MEDLINE-indexed records (generative artificial intelligence OR large language model OR multimodal model) AND (hospital OR healthcare) AND (governance OR implementation OR safety OR bias OR privacy OR evaluation OR monitoring) English; primarily 2019–12 Aug 2026; foundational earlier records eligible; targeted title/abstract screening
WHO Site search for artificial intelligence ethics governance health and large multimodal models Official publications only
NIST Site search for AI RMF, generative AI profile, and Privacy Framework Final official publications prioritized
ONC/HealthIT.gov Site search for HTI-1 and algorithm transparency Final rule and official explanatory materials
FDA Site search for good machine learning practice and AI-enabled software guidance Final guidance and official principles
Other Backward reference screening and targeted CHAI/AMA searches Included only when directly relevant and verifiable