Artificial intelligence contributes to medical diagnosis by performing specific pattern-recognition tasks on clinical data and presenting the result to a clinician who interprets it in context. In almost all deployed systems the algorithm produces an output — a probability, a flag, a measurement, a prioritised worklist — and a human makes the decision.
This distinction between producing an output and making a decision is not a technicality. It determines how such systems are regulated, how they must be validated, where their risks lie, and how they should be evaluated in practice.
Where Diagnostic AI Is Actually Used
Radiology and Medical Imaging
Imaging is the most developed application area because images are standardised, digital, and abundant, and because many tasks are well defined. Deployed uses include triage systems that flag studies likely to contain time-critical findings such as intracranial haemorrhage or large vessel occlusion so they are read sooner; detection tools that mark candidate lesions or nodules for clinician review; quantification tools that measure volumes, densities, or dimensions more consistently and quickly than manual assessment; and reconstruction algorithms that permit shorter acquisition times or lower radiation dose.
The most tangible benefit in practice has often come from workflow effects — reducing time to treatment for time-critical conditions by reordering the reading queue — rather than from improvements in diagnostic accuracy.
Digital Pathology
Whole-slide imaging allows algorithmic analysis of tissue sections. Applications include detecting metastases in lymph node sections, quantifying biomarker expression more reproducibly than visual estimation, grading tumours, and prioritising slides for review. Adoption depends on digital scanning infrastructure, which many laboratories do not yet have.
Ophthalmology
Retinal image analysis for diabetic retinopathy screening is the clearest example of authorised near-autonomous diagnostic AI, permitted in some jurisdictions under defined conditions with referral of positive or ungradable results. The task suits automation: a standardised image, a well-defined grading scheme, high volumes, and a screening context where the required action is referral rather than diagnosis.
Cardiology
Automated electrocardiogram interpretation has existed for decades in rule-based form; machine learning has extended it to detecting arrhythmias from continuous monitoring and wearable data, and to identifying structural abnormalities from signals where the association is not visually apparent to a reader. Echocardiographic measurement automation reduces inter-observer variability.
Dermatology
Image classification for skin lesions has been extensively studied. Persistent concerns include under-representation of darker skin tones in training datasets, differences between curated dermatoscopic images and photographs taken by the public, and the risk that consumer-facing applications provide inappropriate reassurance.
Laboratory and Genomic Diagnostics
Applications include flagging abnormal results, supporting microbiology identification, and interpreting genomic variants, where algorithms assist in classifying variants of uncertain significance against reference databases and predicted functional effects.
Risk Prediction and Early Warning
Prediction models estimating deterioration, sepsis, readmission, or specific complications are widely deployed. Evidence is mixed, with several widely implemented systems performing considerably worse in external evaluation than in development. A prediction is clinically useful only where an effective action follows it and where the alert reaches someone able to take that action.
How Diagnostic AI Systems Are Built
A typical development sequence involves defining the clinical task precisely; assembling a labelled dataset, where labels come from expert annotation, from a reference standard such as histopathology, or from outcomes recorded in the record; splitting data into training, tuning, and held-out test sets; training and optimising the model; and evaluating performance on data not used in development.
Two choices determine much of what follows. The first is the reference standard against which the model is trained and judged: a model trained on clinician labels learns to reproduce clinician judgement, including its errors, whereas one trained against histopathology or documented outcomes learns something closer to ground truth. The second is data provenance: models trained on data from a small number of institutions frequently perform worse elsewhere, because they learn features specific to local equipment, protocols, and populations.
Validating Diagnostic AI
Levels of Evidence
- Internal validation on held-out data from the same source. Necessary but weak; it does not test generalisation.
- External validation on data from different institutions, equipment, and populations. This is where performance commonly falls, and it is the minimum credible standard before deployment.
- Prospective evaluation in live clinical use, measuring how the system performs when integrated into workflow with real-time data.
- Outcome evaluation, ideally randomized, measuring whether patient outcomes, time to treatment, or clinician performance actually improve.
Reporting guidelines exist specifically for this literature, including TRIPOD for prediction models and CONSORT-AI and SPIRIT-AI for trials of AI interventions.
Why Reported Accuracy Can Mislead
- Prevalence dependence. Sensitivity and specificity may transfer between settings, but positive predictive value does not: a model validated in a high-prevalence referral population will generate far more false positives in a low-prevalence screening population.
- Dataset shift. Differences in scanner, protocol, population, or documentation practice between development and deployment degrade performance.
- Label noise. If the reference standard is imperfect, measured accuracy reflects agreement with imperfect labels.
- Spectrum bias. Datasets over-representing clear-cut cases overstate performance on the ambiguous cases where help is most needed.
- Shortcut learning. Models can key on incidental features — annotations, positioning, or scanner artefacts correlated with the outcome — producing high test accuracy that does not generalise.
Where the Model Sits in the Workflow
The same model has different risk profiles depending on deployment mode.
- Triage and prioritisation: reorders the reading queue without changing interpretation. Comparatively low risk, since every study is still read, though a missed urgent case may be deprioritised.
- Concurrent assistance: presents findings alongside clinician review. Risk centres on automation bias — accepting incorrect suggestions — and on distraction from findings the model does not address.
- Second reader: reviews after the clinician, flagging potential misses. Adds sensitivity at the cost of additional false positives and follow-up.
- Autonomous operation: produces a result without clinician review. Requires substantially stronger evidence and is authorised only for narrow tasks under defined conditions.
Human Factors and Oversight
Automation bias — over-reliance on algorithmic output — is well documented and can cause clinicians to miss findings they would otherwise have detected. The opposite failure, disregarding a correct alert amid excessive alerting, is equally documented. Both are properties of the deployed system rather than of the model, and both can only be detected by studying use in practice.
Effective oversight requires clinicians to know what the model was designed to do, what population it was validated in, what it does not detect, and how confident to be in its output. Documentation describing intended use, training data, validation populations, subgroup performance, and known limitations supports this, and is increasingly expected by regulators and by institutional governance processes.
Regulation and Governance
Diagnostic AI intended for a medical purpose is regulated as software as a medical device. In the United States, such tools are cleared or approved through device pathways, with frameworks permitting predetermined change control so that anticipated model updates can proceed within agreed limits. In the European Union, the Medical Device Regulation applies, with the AI Act adding obligations for high-risk systems covering data governance, technical documentation, human oversight, robustness, and post-market monitoring.
Institutional governance should establish who approves deployment, what validation in the local population is required, who is clinically accountable for decisions informed by the model, how performance is monitored after deployment, what thresholds trigger review or suspension, and how incidents are reported.
Sources
- U.S. Food and Drug Administration — artificial intelligence and machine learning-enabled medical devices; software as a medical device guidance
- European Commission — Medical Device Regulation; Artificial Intelligence Act
- International Medical Device Regulators Forum — software as a medical device risk framework
- World Health Organization — ethics and governance of artificial intelligence for health
- TRIPOD — reporting guideline for prediction model development and validation
- CONSORT-AI and SPIRIT-AI — reporting guidelines for clinical trials of AI interventions
- STARD — reporting standards for diagnostic accuracy studies