Essay #3: Living With AI in Healthcare series
The Diagnostic Horizon
Essay two ended with a promise: this piece turns from the conversation to the data itself — how AI is beginning to read scans, flag lab results, and catch patterns a tired pair of human eyes might miss. That promise runs up against a much older problem.
For most of the last century, diagnosis has depended on a physician’s own memory — the ability to hold thousands of conditions in mind, recognize a subtle pattern from training, recall a journal article read years earlier, all while standing at a patient’s bedside under real time pressure. That model is being tested now, not because doctors know less, but because there’s simply more to know than one mind can hold. Modern diagnostics produces an overwhelming volume of information daily — high-resolution imaging, continuous vital-sign data, genetic testing, years of EHR notes — far more than any physician could review in real time for every patient. AI is entering that gap not as a replacement for clinical judgment, but as a first pass through the flood.
The High-Speed Filter
In radiology, that first pass is already running at scale. Viz.ai, cleared by the FDA in 2018 as the first AI platform of its kind for stroke triage, is now deployed in more than 1,600 hospitals; the company reports patients reaching treatment roughly 66 minutes faster on average when it flags a large-vessel occlusion the moment a scan comes in. Aidoc, used in hundreds of hospitals for detecting intracranial hemorrhage and other urgent findings, cites a University of Rochester Medical Center study showing a 36.6% reduction in turnaround time for hemorrhage cases. Neither tool replaces the radiologist’s diagnosis — a physician still interprets the scan and makes the clinical decision — but a study that used to sit in a routine queue for hours can now reach a specialist’s screen within minutes of being taken, specifically because the software moved it to the top of the list.

Pathology is following a similar path, just a few years behind. In 2021, Paige.AI’s Paige Prostate Detect became the first FDA-authorized AI application in pathology, scanning digitized biopsy slides to flag areas suspicious for cancer and grade tumor severity, cutting pathologists’ reading time by roughly 20%. A newer tool from the same company, aimed at flagging cancer across many different tissue types at once, received FDA Breakthrough Device status in 2025. Pathology overall remains earlier in this journey than radiology, which already has hundreds of FDA-cleared AI tools in routine use — but the underlying idea is the same one already at work in imaging: a first pass through the slide, flagging what deserves a closer look.
Beyond the Single Image
The next step past triage is connecting the dots across a patient’s whole record, not just one scan in isolation. Early diagnostic AI tools were narrow by necessity — a model trained to read chest X-rays could spot a nodule but had no idea about anything else in that patient’s chart. The current push is toward models that read several kinds of information together: an image alongside recent bloodwork, a genetic marker, and years of clinical notes, weighing them as a single picture rather than separate, disconnected data points. In principle, a shadow on a lung scan that looks ambiguous by itself could read very differently when considered alongside a documented pattern of unexplained weight loss and an elevated blood marker — the same kind of reasoning a physician does by hand today, just done faster and across a much wider set of records than one person could hold in mind at once.
One Model or Many? The Generalist-Specialist Question
There’s a genuine, current debate running underneath all of this, and it’s less about what AI can see than about how it should be built. Do we want one medical AI that knows a little about everything, or many specialist AIs that each know one narrow area extremely well? Some of the strongest results so far — the stroke, hemorrhage, and pathology tools above included — come from narrow specialists, trained for exactly one task and very good at it. But researchers are increasingly testing something broader: large, flexible “generalist” models trained across huge, varied datasets.
One notable answer, from a research team at the Hong Kong University of Science and Technology, doesn’t pick a side. Their framework, GSCo, has a broad generalist model consult narrow specialist models for precise input on a given case, then make the final call informed by both — tested across 32 datasets spanning radiology, pathology, dermatology, and ophthalmology, and outperforming either approach working alone. Newer research since has found a more nuanced split: specialists sharper at picking the single best answer, generalists better at generating a broader range of plausible possibilities — suggesting the two may end up suited to different parts of the same job, rather than one replacing the other. Nothing here is settled. But it’s a live, active area of research, not a dead end.
The Demand for Explainability
None of this works, though, if a doctor can’t tell why the software flagged what it flagged. For years, many of the deep-learning models behind these tools operated as “black boxes” — producing a confident-sounding result with no visible reasoning behind it. That might be a minor annoyance in a shopping app. In a clinical setting, an algorithm that can’t explain itself is a much harder thing to trust, and a much harder thing to stand behind if it turns out to be wrong.
That trust gap is why explainability has become such an active area of its own. Rather than a bare prediction, the better tools now produce something closer to a receipt: a heatmap overlaid directly on a scan, highlighting the cluster of pixels associated with the result — much the way a radiologist might circle a spot with a red pen — or a highlighted passage from a decade of chart notes pointing to what contributed to a risk score. That added transparency gives the physician something concrete to examine alongside the machine’s conclusion, rather than simply trusting or ignoring a number.
The Friction: Automation Bias and Alert Fatigue
Even a well-explained tool creates a new kind of risk: what happens when a busy clinician stops checking. Automation bias is the tendency to accept a confident-looking algorithmic result without the independent verification that’s supposed to happen alongside it. The opposite failure is just as real — alert fatigue, where a system cries wolf so often that clinicians start tuning it out entirely, missing the alerts that actually mattered.
The Epic Sepsis Model is the clearest real example of both problems at once — and its history since is worth following all the way to the present, because it shows how hard this is to actually fix. Built into the electronic health record used by hundreds of U.S. hospitals, the original version was independently validated by University of Michigan researchers in a 2021 study published in JAMA Internal Medicine, and found to miss two-thirds of actual sepsis cases while still generating alerts on 18% of every hospitalized patient. During the early months of COVID, alert volume from the model spiked 43% even as patient census dropped, a burden severe enough that Michigan Medicine paused the alerts entirely for a period in 2020.
Epic responded. A rebuilt version released in 2022 was designed to address that criticism directly and let hospitals fine-tune it on their own local data. It’s now in use at hundreds of hospitals — though many others have stuck with the original, already-criticized version rather than switching. And the most rigorous independent test of the new version yet, a large multicenter study published in JAMA Network Open in early 2026, found real improvement in raw accuracy, but the same underlying problems still showing up: high variability from hospital to hospital, a low rate of alerts actually being correct, and a persistently heavy burden on staff. Five years, one full model rebuild, and hundreds of hospitals later, the core tension between catching real sepsis early and not burying clinicians in false alarms hasn’t been solved. It’s been narrowed — which is real progress, but not the same thing as fixed.
One more limitation worth naming honestly, without expanding on it here: models trained mostly on data from large academic hospitals don’t always perform as well once they reach smaller or lower-resourced clinics, particularly in rural areas — often because the imaging equipment differs, the patient population the model learned from doesn’t match the one it’s now seeing, and smaller hospitals rarely have the budget or staff to locally retest and retrain a model the way a well-funded academic center can. That’s a real gap, and a meaningful one — but it belongs to a bigger conversation about who these tools are built for and who they’re tested on, which is exactly the territory the Guardrails essay later in this series is built to cover properly.
In Plain Terms
Here, in diagnostics, AI isn’t making the call. It’s filtering an unmanageable flood of scans and data down to what a human needs to see first, connecting details a rushed reading might miss, and increasingly showing its work well enough for a physician to check it. What it hasn’t done, and by design isn’t meant to do, is replace the judgment that has to happen after — the physician’s role shifts from gathering every piece of information alone toward critically evaluating what the machine has already assembled, which is a different job, not a smaller one.
The next essay in this series moves from the moment of diagnosis to what happens after a patient walks out the door — how AI is starting to watch, quietly, for what happens next.
State of Play, as of August 2026
Viz.ai’s FDA clearance and hospital deployment figures, along with its reported average treatment-time improvement, come from the company’s own published materials and third-party review coverage; these aren’t independently audited figures. Aidoc’s turnaround-time figure comes from a University of Rochester Medical Center study cited in the company’s own publications. Paige.AI’s Paige Prostate Detect received FDA authorization in 2021 and its reading-time figure is company-reported; Paige PanCancer Detect received FDA Breakthrough Device designation in 2025, with performance data from internal company studies not yet independently confirmed. The GSCo generalist-specialist framework was first published as a research preprint in April 2024 and reached formal peer-reviewed publication in Nature Biomedical Engineering in 2026; the more recent specialist-versus-generalist benchmark finding comes from a September 2025 study not yet through full peer review, included as an emerging data point rather than a settled finding. The original Epic Sepsis Model validation was published in JAMA Internal Medicine in 2021 by University of Michigan researchers, with COVID-era alert-volume figures reported separately and covered by Fierce Healthcare; the updated model’s most recent multicenter validation (95 organizations, more than 730 hospitals using the current version; more than 100 still on the original) was published in JAMA Network Open in early 2026.
