How Accurate Is ChatIBD? An Early Look at Version 2.4
ChatIBD has been answering clinicians' questions for a year now. Here we share how accurate the current version, ChatIBD-2.4, has been since it launched in August, and where it still falls short.
General-purpose AI models are known to invent references, go along with false details slipped into a question, and rarely say "I don't know" when the evidence is unclear. ChatIBD is built to avoid this. Every answer draws on an evidence layer made up of versioned IBD guidelines, the official product labels (SmPCs) for IBD drugs and a structured dosing dataset, and it cites the passages it used so that you can open them. We wanted to measure how well that works.
How we checked
We looked at ChatIBD in two ways.
Real-world monitoring. Since ChatIBD-2.1, two different AI models (GPT-5.6 Sol and Claude Opus 4.8) have each read the previous day's answers every night and flagged anything with a safety problem or a claim its sources don't support. A clinician then checks every flag, and an answer counts if either model flagged it. We ran the same review over the earlier versions too. An answer counts as a strict error if it delivered a clinically material claim that was invented, or wrong when checked against current guidelines or product labelling. Appropriate hedging, correct refusals and formatting slips don't count.
Synthetic stress tests. We also wrote sets of questions designed to find weaknesses, including thin or missing evidence, false premises and hard dosing cases. We ran them through the same underlying AI model twice: once as ChatIBD, and once with the evidence layer switched off.
About 1 in 300 reviewed answers had an error we could find

In its first 25 days as the default, ChatIBD-2.4 gave about 1,140 answers. On busy nights the reviewers read a sample rather than everything, so we estimate they covered about 910 of them (somewhere between 774 and 1,007). Three of those answers (0.33%) contained a strict error, and none involved invented content. Earlier versions ran at 1.54% (2.1), 1.36% (2.2) and 0.96% (2.3). Those estimates overlap, and the underlying model and safety checks changed along the way, so we read this as encouraging rather than as a proven trend.
The rate rises to about 1% (9 answers) if we also count one misread disease name and five answers that stretched evidence beyond its scope, such as leaving out that a child's dose depends on weight. No user reported a clinical error, though that reflects how often people report errors rather than how accurate the answers were.
Without the evidence layer, errors were far more common
To see what the evidence layer contributes, we took the opening question from 200 ChatIBD-2.4 chats and gave it to the same model with the evidence layer switched off. A strict error appeared in 26.5% of those answers, each one confirmed by a clinician, and 13% contained outright fabrication. That is about 80 times ChatIBD-2.4's strict rate, or about 25 times if we use the broader 1% count. The two setups weren't identical (the comparison only used each chat's first question), so treat this as directional, but the gap is large.
The synthetic tests show the same pattern. We gave ChatIBD 27 questions, mostly where the evidence was thin, missing or built on a false premise. It fabricated nothing. Instead it declined to answer, answered only the part it could support, or corrected the premise (one answer still attached the wrong units to a statistic). With the evidence layer switched off, five answers fabricated content, including a trial of duvakitug that doesn't exist and a paediatric UC licence for upadacitinib that doesn't exist either. Two more were wrong without inventing anything: one used an outdated surveillance interval, and one applied a US threshold to a UK question.
On dosing, all 72 questions with a checkable answer matched the structured dosing dataset, across induction and maintenance regimens for 19 drugs in UC and Crohn's disease. Citations linked to a real retrieved passage in 25 of 26 answers; the other had a malformed link. These checks confirm agreement with our own sources, not clinical correctness in every case.
The trade-off is speed and cost. On our hardest question set, answers took a median of 18 seconds with the evidence layer against 10 without it, at about twice the cost. On the everyday specialist questions the gap was smaller: 12.5 seconds against 9.3.
Where it still gets things wrong
The evidence layer reduces errors but doesn't eliminate them. Here are the real-world cases from ChatIBD-2.4:
- It drafted a note saying a patient's hydrocortisone could be stopped, without mentioning the need to taper.
- It quoted a drug's VTE rate correctly but attached it to the wrong group: patients who had stopped the drug because of adverse effects, rather than everyone taking it.
- It conflated hepatitis B vaccination of non-immune patients with reactivation screening of carriers, though it still recommended the right tests.
- It read a misspelling of Bechterew's disease (ankylosing spondylitis) as Behçet's disease, which is managed quite differently. We didn't count this as a strict error, because the information was accurate but about the wrong condition, and the user caught it.
In our 32-question hard set, a safety review flagged 3 ChatIBD answers, compared with 4 with the evidence layer switched off. The clearest failure involved tofacitinib renal dosing. For a UK patient on the immediate-release 10 mg twice daily regimen, ChatIBD retrieved guidance for the prolonged-release formulation and recommended reductions below the SmPC regimen. The model without the evidence layer made the same mistake. The other two flagged answers left out relevant monitoring, one for Pneumocystis prophylaxis and one for ciclosporin rescue.
When ChatIBD goes wrong, it is rarely inventing things. More often it misreads the question or over-extends evidence it did retrieve. Citations let you check the source, but they won't catch a misread question.
What these numbers don't tell you
We built ChatIBD, and we ran this evaluation ourselves, which is a conflict of interest. The review was not blinded and relied mostly on AI reviewers, with a single clinician checking what they flagged. Because the clinician only checked flagged answers, any error both AI reviewers missed isn't counted, so treat the real-world rates as a floor. Our test questions were deliberately hard, so they don't represent everyday use, and each one was answered only once. Until 5 September, users could switch back to the previous version, so a minority of answers in the 2.4 window came from it. Nothing here shows that ChatIBD improves care or is safe to use unsupervised.
The numbers also describe a single point in time. Over the past year ChatIBD has moved from GPT-4.1 to GPT-5.6 Luna and absorbed updates to two guidelines, and either kind of change can shift how it performs. That is why the nightly review keeps running.
Full data, and telling us about errors
The full methods and data have every question set, the scored dosing cases, the per-version counts and the lower-severity cases.
If you spot an answer that looks wrong, please tell us. Those reports are part of how we measure all of this.