Artificial intelligence in healthcare refers to computational systems that perform tasks requiring pattern recognition, prediction, or language processing on clinical data. In practice, most deployed clinical AI is machine learning: models trained on large datasets to recognise patterns — a nodule on a scan, a deterioration signature in vital signs, a phrase in a dictated note — and to produce an output a clinician or system then acts on.
The realistic assessment is that AI has demonstrable value in specific, narrow, well-defined tasks, particularly in imaging and in administrative work, and that its broader clinical promise remains substantially unproven. The gap between reported model performance and demonstrated patient benefit is the central issue in the field.
What AI in Healthcare Actually Means
Several distinct technologies sit under the term. Supervised machine learning trains models on labelled examples to predict a defined output. Deep learning uses multi-layer neural networks, and is the basis of most modern medical image analysis. Natural language processing extracts structure and meaning from clinical text. Large language models generate text and are being applied to documentation, summarisation, and information retrieval. Rule-based systems, which encode expert logic explicitly rather than learning from data, remain widespread in clinical decision support and are often more transparent than learned models.
The distinctions matter for evaluation and governance. A deterministic rule-based alert can be inspected and reasoned about; a deep learning model's behaviour must be characterised empirically, and a generative model may produce fluent output that is incorrect.
Where AI Is Used in Healthcare
Medical Imaging
Imaging is the most mature application area, reflecting large volumes of standardised digital data and well-defined tasks. Deployed uses include triage of urgent findings such as intracranial haemorrhage or large vessel occlusion, detection and measurement of nodules and lesions, quantification of volumes and densities, image reconstruction allowing reduced acquisition time or radiation dose, and workflow prioritisation of worklists. Most such tools are regulated devices intended to support rather than replace clinician interpretation.
Clinical Documentation and Administration
Ambient documentation systems transcribe and structure clinical encounters, reducing time spent on notes. Related applications include coding support, correspondence generation, prior authorisation and claims processing, appointment scheduling and capacity forecasting, and summarisation of long records. These uses currently show some of the most consistent practical benefit because they address documented workload burden and because errors, while consequential, are visible to the clinician who reviews the output.
Risk Prediction and Early Warning
Models predicting deterioration, sepsis, readmission, or mortality are widely deployed. Evidence is mixed: some early warning systems have shown benefit, while others have performed substantially worse in external validation than in development, and alert fatigue can blunt any benefit. Prediction is only useful where a specific effective action follows the alert.
Diagnostic Support
Applications include analysis of electrocardiograms, digital pathology slide review, retinal image screening for diabetic retinopathy, dermatological image assessment, and genomic variant interpretation. Diabetic retinopathy screening is among the areas where autonomous or near-autonomous use has been authorised in some jurisdictions under defined conditions.
Drug Discovery and Research
Machine learning is used in target identification, molecular property prediction, protein structure prediction, compound screening, trial site selection, and patient identification for trials. These applications shorten specific steps; they do not remove the requirement for clinical testing, and their downstream impact on approval rates is not yet established.
Operational and System Applications
Demand forecasting, bed and theatre scheduling, staffing optimisation, supply chain management, and population health risk stratification apply established analytics methods to health system operations, often with clearer measurement of benefit than clinical applications.
Potential Benefits
- Consistency. Algorithms do not tire, and can apply the same criteria across large volumes.
- Throughput and triage. Prioritising urgent findings can shorten time to treatment for time-critical conditions.
- Workload reduction. Automating documentation and administrative tasks addresses a well-documented contributor to clinician burnout.
- Extending scarce expertise. Screening support may extend specialist capability to settings without on-site specialists, provided performance in those settings is validated.
- Scale of data. Models can integrate more variables than a clinician can hold in working memory, potentially detecting patterns not otherwise apparent.
Challenges and Risks
The Evidence Gap
Most published clinical AI research reports retrospective performance on curated datasets. Comparatively few studies evaluate prospective deployment against patient outcomes, and fewer still are randomized. Reporting guidelines have been developed specifically for AI trials and prediction models to address this. Performance frequently falls when models are tested on data from institutions, scanners, or populations other than those used in development.
Bias and Equity
Models learn from historical data, and historical data encode existing patterns of access, practice, and measurement. If a group is under-represented in training data, performance for that group may be worse. If a model is trained on a proxy variable that itself reflects unequal access — historical spending as a proxy for need, for example — it can systematically underestimate need in disadvantaged populations. Bias assessment therefore requires performance evaluation stratified by relevant subgroups, not aggregate accuracy alone.
Transparency and Explainability
Many high-performing models are not interpretable by inspection. Explainability methods provide partial insight but can themselves be unreliable. The practical response has centred on transparency about what a model was trained on, in what population it was validated, what its intended use and limitations are, and how it performs across subgroups — information now expected in regulatory submissions and model documentation.
Automation Bias and Workflow Effects
Clinicians may over-rely on algorithmic output, accepting incorrect suggestions they would have caught unaided, or may become desensitised by excessive alerting. Both effects are properties of the deployed system rather than of the model, and can only be detected by studying use in situ.
Drift and Monitoring
Model performance can degrade when populations, clinical practice, equipment, or documentation habits change. Continuous post-deployment monitoring against defined performance metrics is necessary, and frameworks exist for managing planned model updates without a new regulatory submission for each change.
Data Governance and Privacy
Training and deployment involve sensitive personal data, engaging data protection law, requirements for lawful basis and, in some frameworks, consent, and questions about secondary use of clinical data. Techniques including de-identification, federated learning, and synthetic data reduce but do not eliminate risk, and re-identification remains a documented concern with rich health datasets.
Generative AI Specifically
Large language models introduce distinct risks: fluent but incorrect output, sensitivity to how a question is phrased, difficulty in verifying provenance of statements, and non-determinism that complicates validation. Their use in clinical documentation requires clinician review of output, and their use for clinical advice raises regulatory questions that frameworks are still addressing.
Regulation and Governance
Where software is intended for a medical purpose, it is generally regulated as a medical device. In the United States, AI-enabled devices are cleared or approved through established device pathways, and frameworks exist for predetermined change control plans allowing anticipated model updates within agreed boundaries. In the European Union, such software is regulated under the Medical Device Regulation, with the AI Act adding a horizontal risk-based framework in which many healthcare applications fall into a high-risk category carrying obligations on data quality, documentation, human oversight, robustness, and post-market monitoring.
The World Health Organization has published guidance on the ethics and governance of artificial intelligence for health, setting out principles including protecting autonomy, promoting human wellbeing and safety, ensuring transparency and explainability, fostering responsibility and accountability, ensuring inclusiveness and equity, and promoting responsive and sustainable systems.
Institutional governance is equally important: defined approval processes for deploying models, clarity about clinical responsibility for decisions informed by algorithmic output, documented intended use, staff training, incident reporting routes, and monitoring arrangements with predefined thresholds for suspension.
Sources
- World Health Organization — Ethics and governance of artificial intelligence for health; guidance on large multi-modal models
- U.S. Food and Drug Administration — artificial intelligence and machine learning in software as a medical device; predetermined change control plans
- European Commission — Artificial Intelligence Act; Medical Device Regulation
- International Medical Device Regulators Forum — software as a medical device framework
- CONSORT-AI and SPIRIT-AI — reporting guidelines for clinical trials of AI interventions
- TRIPOD — reporting guideline for prediction model studies