Why Drug Interactions Matter
Many older adults take five or more medicines every day. Doctors call this polypharmacy. In the United States, about 4 in 10 adults over 65 fall into this group.
The more medicines a person takes, the higher the chance that two of them clash. These drug interactions are a leading cause of avoidable side effects, hospital stays, and even deaths.
AI tools are now being promoted as a way to catch these clashes. So the question is simple: can we trust them?
What the Study Tested
Researchers at the University of Colorado tested three AI models: GPT-4-mini, MedGemma-27B, and LLaMA3-70B.
They used 750 drug interaction cases, all checked by clinical pharmacists. The cases came in three levels:
- Two drugs: Is there an interaction?
- Three drugs: Which pair causes the problem?
- Four to six drugs: Which combination is risky?
What They Found
The models did reasonably well with two drugs. As more medicines were added, their accuracy dropped.
They also became less consistent with their answers; in every prompt the responses were never the same. None of the three was reliable across all levels.
Here is why that matters. A tool that is right 60% of the time is wrong 40% of the time. And the doctor cannot tell which answer they are looking at. Right and wrong answers look exactly the same on screen.
In medication safety, inconsistency is not a small flaw. It is a failure.
A Real-World Example
A patient got admitted to the hospital; the models were asked about seven medicines that the patient was prescribed: a blood thinner, a seizure medicine, an acid reducer, an antidepressant, a water pill, a cholesterol medicine, and a new antibiotic.
An AI tool is asked to check for clashes. It easily spots a well-known two-drug interaction. But checking all seven together is a much harder task, and this is exactly where the study found AI slips.
The report comes back with nothing flagged. It looks like a complete check. It is not. The doctor moves on.
The Testing Gap
Most AI tools are tested on simple two-drug questions. Passing those tests is then treated as proof that the tool is ready for real patients.
But real patients rarely take just two medicines. Current tests often do not check:
- How accuracy changes as more drugs are added.
- Whether the AI gives the same answer every time.
- Whether it admits when it is unsure.
- How it handles “monitor closely” versus “never combine” interactions.
The Bottom Line
AI models are not clueless about drug interactions. They are fragile. They work on easy cases and struggle on hard ones.
A tool that handles two medicines but fails on six is only useful for simple cases. Until AI is tested on real, complex medicine lists, we will keep overestimating how ready it is. Consistency is not a side issue. It is the main finding.
Publisher: Source link









