Discover your interests, together

Real deals, honest reviews and shopping stories from people who share your interests — every day on Milik.

Discover your interests, togetherReal deals, honest reviews and shopping stories from people who share your interests — every day on Milik.

AI Can Detect Disease With 90%+ Accuracy—So Why Aren't Doctors Using It?

AI Can Detect Disease With 90%+ Accuracy—So Why Aren't Doctors Using It?
Interest|AI Application Exploration

The uncomfortable truth about high AI accuracy

AI disease detection accuracy refers to how reliably early diagnosis AI tools can spot conditions such as fatty liver disease and Alzheimer’s from medical images or health records, often exceeding 90% sensitivity and specificity on curated datasets but without clear proof that these earlier findings change real-world outcomes for most patients. In liver disease, imaging-based AI for hepatic steatosis reaches pooled sensitivity of 91%, specificity of 92%, and an area under the curve of 0.97, with some convolutional neural networks scoring an AUC of 1.00. Feed one of these models an ultrasound frame and it will reliably say whether there is fat in the liver. Yet that success is misleading. It celebrates detection while dodging the harder question: does labeling millions of people with a condition improve their lives, or only their test reports?

AI Can Detect Disease With 90%+ Accuracy—So Why Aren't Doctors Using It?

When a "successful" test is an answer to the wrong question

Metabolic dysfunction-associated steatotic liver disease affects around 25% of the world’s adults, making it the most common chronic liver disease on the planet. That is exactly the kind of scale that excites AI developers: one test, billions of livers. On paper, AI disease detection accuracy looks like a triumph. In practice, it exposes a clinical validation gap. A screening test earns its keep only if a positive result routes the patient to an intervention that improves the outcome. For MASLD, first-line management is still lifestyle change, and the few new drugs focus on advanced disease, not early fat accumulation. The trials that would show earlier detection changes what happens to patients are the ones that are missing. Until those exist, mass screening for fatty liver with AI is less clinical progress and more diagnostic theater.

Here is the ethical mess: a highly sensitive AI sweeps an at-risk population with 25% baseline prevalence. A positive result raises post-test probability of NAFLD to 79%, flagging enormous numbers of people. For many, that “disease” will never matter. Yet once AI prints the label, clinicians inherit repeat imaging, elastography, referrals, and the anxiety of someone told their liver is diseased—a downstream testing cascade with real cost and no demonstrated benefit for the median patient. The field has built an excellent answer to the wrong question: whether we can find fat in the liver, instead of whether finding it early improves who lives, who dies, and who avoids cirrhosis or cancer.

Alzheimer’s AI: fairer, earlier, still not fully proven

The story in Alzheimer’s disease looks more hopeful—but reveals the same clinical validation gap from a different angle. Researchers developed an artificial intelligence tool that can use electronic health records to identify patients with undiagnosed Alzheimer’s disease, addressing a critical gap in care. Disparities in Alzheimer’s and dementia diagnosis have been longstanding: African Americans are nearly twice as likely to have the disease as non-Hispanic whites but only 1.34 times as likely to receive a diagnosis, while Hispanic and Latino people are 1.5 times more likely to have the disease but only 1.18 times as likely to be diagnosed. This new model uses semi-supervised positive unlabeled learning, explicitly designed to promote fairness while maintaining high accuracy.

On performance, it is a clear upgrade. The model achieved sensitivity rates of 77 to 81% across non-Hispanic white, non-Hispanic African American, Hispanic/Latino, and East Asian groups, compared to 39 to 53% sensitivity for conventional supervised models. Patients predicted to have undiagnosed Alzheimer’s showed significantly higher polygenic risk scores and APOE ε4 allele counts than those predicted not to have it, strengthening the case that the AI is finding real disease. Early identification is crucial as new Alzheimer’s treatments become available and lifestyle interventions can slow disease progression. Yet even here, the team is explicit: they plan to validate the model prospectively in partnering health systems to assess generalizability and clinical utility before routine use. In other words, fairness measures and accuracy gains are not enough; they still need proof that flagging people earlier changes their trajectory.

Diagnostic disparities and the myth of automatic equity

The Alzheimer’s work exposes another myth around early diagnosis AI tools: that higher accuracy automatically fixes diagnostic disparities. Disparities in Alzheimer’s and dementia diagnosis among certain populations have been a longstanding issue. The new model tackles this head-on by baking fairness into its semi-supervised design and using population-specific criteria to reduce diagnostic disparities. "By ensuring equitable predictions across populations, our model can help remedy significant underdiagnosis in underrepresented populations," Chang said. "It has the potential to address disparities in Alzheimer's diagnosis". That is a serious attempt to make diagnostic disparities AI-aware instead of AI-blind.

But equity in detection is still only step one. If underdiagnosed communities get more labels without parallel investment in treatment access, caregiver support, and culturally appropriate follow-up, the gap shifts rather than shrinks. The liver story is a warning here too. For MASLD, the field’s missing piece is risk stratification: separating the minority who will progress to fibrosis and cancer from the majority who will not. One genetics-informed study reached AUROC up to 0.87 for risk prediction, but between “one study at 0.87” and “deployed stratification tool” sits the validation nobody has funded. Without such tools, AI can flood clinics serving marginalized groups with new positives, but not tell clinicians which of those patients most need scarce intervention slots.

Why health systems should say "not yet"—and what needs to change

Healthcare systems face mounting pressure to adopt any AI disease detection accuracy figure that looks impressive on a slide deck. The scoping review on fatty liver AI notes "implications for improved clinical outcomes"—then stops there. Implications are not outcomes. Until there is either a prospective trial showing that AI-detected early MASLD, acted on, produces fewer cirrhosis or liver cancer cases than standard care, or a stratification model validated across ancestries that reliably tells the progressors from the rest, the field has built an excellent answer to the wrong question. The Alzheimer’s team, to their credit, is taking the slower road, planning prospective validation in partner health systems before routine deployment.

The way forward is not to reject early diagnosis AI tools, but to raise the bar. Regulators and payers should demand three things before widespread adoption: proof that earlier detection changes meaningful outcomes; clear pathways for what happens after a positive result; and evidence that benefits, not just labels, are distributed fairly across populations. Until then, every “breakthrough” detector should be treated as experimental. It can tell a billion people they have fatty liver. It cannot yet tell them, or their doctors, what to do about it. The technology is ready to see more disease. Medicine should wait until it is equally ready to do something better with that knowledge.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!