AI-assisted decision support in low-resource primary care: what the Kenyan trial shows and what it doesn't
lpm.org

AI-assisted decision support in low-resource primary care: what the Kenyan trial shows and what it doesn't

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRA randomized trial of GPT-4o as a 'second pair of eyes' in Kenyan primary clinics improved diagnostic quality but didn't show statistically significant patient outcome improvements.

A randomized trial across 16 Kenyan primary clinics tested whether OpenAI's GPT-4o could act as a 'second pair of eyes' for clinicians. The results show improved care quality but no statistically significant patient outcome benefit. For AI builders, the study offers a rare real-world look at how LLMs perform in low-resource clinical settings.

What happened

Researchers tested AI Consult, a tool that uses OpenAI's GPT-4o to review clinicians' electronic notes in nearly 10,000 patient encounters across 16 Kenyan primary clinics operated by Penda Health. Half of the clinicians used the AI tool; the other half took notes without it.

The tool uses a three-color alert system: green means everything is fine, yellow flags a small issue with guidance, and red signals a critical concern requiring prompt action. An independent panel of six Kenyan family physicians graded the notes and found that AI-assisted clinicians produced better diagnoses and treatment plans.

The tool cost about 4 cents per patient. The trial was funded by the Gates Foundation.

Why AI builders should care

This is one of the first randomized trials of an LLM as a clinical decision support tool in a real primary care setting. It demonstrates that a relatively simple notes-review pipeline can improve care quality at very low cost. For builders working on healthcare AI, the study validates the concept of 'information-as-an-intervention' - prompting clinicians with AI-generated alerts can change behavior and catch missed diagnoses.

However, the study also highlights a critical gap: improved process measures did not translate into statistically significant patient outcome improvements. There was a 23% reduction in treatment failures (death or unresolved symptoms), but the result was not statistically significant because such events are rare in primary care. The trial was underpowered to detect a meaningful difference.

Practical implications

For teams building AI-assisted clinical tools, the key takeaway is that proving downstream health outcomes requires massive sample sizes. Co-author Dr. Bilal Mateen noted that a trial would need about 139,000 participants to detect a meaningful difference in treatment failures. That's a high bar for any startup or research group.

On the positive side, the low cost (4 cents per patient) and the fact that most AI recommendations were rated safe and appropriate suggest that such tools can be deployed at scale in resource-constrained settings. Clinician Vyonne Njeri reported that the AI helped her remember protocols and catch issues she would have missed, such as a congenital heart defect in an infant.

Caveats

The study did not find statistically significant improvements in patient outcomes. The non-significant 23% reduction in treatment failures could be due to chance or insufficient power. Additionally, Njeri noted that the AI's recommendations were helpful only about half the time, though they were seldom wrong. Healthcare AI researcher Dr. Nicholas Okumu warned that even approved AI systems can cause harm and require active oversight.

Builders should also note that the trial did not examine whether AI could increase access to care or reduce clinician workload. Those remain open questions for future work.

FAQs

AI Consult uses OpenAI's GPT-4o to review clinicians' electronic notes and flag potential issues via a three-color alert system (green/yellow/red). An independent panel of Kenyan family physicians evaluated the notes for safety and quality, noting improvements in diagnostic and treatment plan quality for AI-assisted cases.

Sources

Latest Tech News