Comparisons

AI for Medical Finals: What Actually Helps

Mostafa Ibrahim10 min read
AI for Medical Finals: What Actually Helps

Does AI actually help with finals?

Yes, for recall, explanations and generating practice questions, AI helps and is already part of normal revision habits. For the clinical half of finals, it helps only if the tool makes you speak and then marks what you said against a rubric. Most general study tools don't do that.

Finals are two different exams wearing one name. There's the part you revise by reading and clicking, and the part where you sit in a room with a person, take a timed history, and then say out loud what you think is happening. You're assessed on what you actually say, not what you meant to say.

AI is good for the first half, because it can check recall, explain tricky bits, and spin up fresh practice questions. The second half needs you to talk, be interrupted, and get marked on your responses, which many tools skip. For a quick scan of options, see the full AI tool roundup.

What AI is genuinely good at for finals

  1. Explanation on demand means you can push it to re-explain a point five different ways until it clicks. Ask for a path, a diagram, a viva-style prompt, then a common pitfall, till your phrasing sounds like an OSCE answer.
  2. Paste your cardiology notes and get 6 SBAs that target the gaps you've flagged, not a generic chapter. Make it explain why the distractor you chose is wrong in plain terms, then rewrite the stem to surface the trap.
  3. Give it a 40 page handout and ask for 10 examinable points with a line on how each is tested. You'll catch the framing examiners favour, like red flags, initial investigations, or safety net wording that textbooks bury mid paragraph.
  4. Turn a scrappy page of notes into a clean hierarchy with headings, bullet checklists, and if-then flows you can revise from. Ask for a history template that mirrors a station brief, then map your content so recall matches how you'll be marked.

All four are about handling information you already half know, and that's genuinely useful in the final stretch. It's also what your textbook and question bank already did, so this is a speed up and it's easy to mistake that for progress; if you're planning preparing for finals overall, keep that in mind.

What AI gets wrong, with numbers

One UKMLA study ran a general chatbot across 191 single best answer questions drawn from two public practice papers, three times each. When it was tested on UKMLA practice papers, it answered 67.5% consistently correctly, 129 of 191, and 12.6% consistently incorrectly, 24 of 191, across all three attempts.

It also changed its answer between runs on about 19.9% of items, 38 of 191. Across the three attempts its overall score averaged 76.3%, high enough to pass many papers and low enough to hide traps when you rely on it mid-quiz. That means about one in five stems produced flip flops across repeats.

The danger isn't the average score you see. It's the bit it gets wrong the same way each time, with the same confidence as the bit it gets right, plus a fifth of items where it flips, and you can't tell those apart, so it's safe where you can check, and least safe where you know least.

One study ran a chatbot on the Foundation Programme's 2023 SJT practice paper, 75 questions, and it scored 76% overall. When the same test on the SJT found full marks on only 9% of questions, it signalled answers that are broadly reasonable and rarely exact, a poor shape for an exam that marks exact ranking.

Both studies tested 2023 models, and models have moved on. My sense is the exact scores are stale, but the failure shape looks sticky.

Chatbot on 191 UKMLA questions over three attempts: 67.5% right every time, 12.6% wrong, 19.9% changed

The tools, honestly

Here are the tools UK final-year students actually choose between, with what each is for and where it helps, and how it fits revision on the ward, in the library, or on the bus. Prices are checked in August 2026 so you can pick on function first and cost second, without guessing. No rankings, no fluff. Prices and use cases only.

  • ChatGPT plans and prices What it is for: Explaining, condensing, generating practice questions, and it will role-play a patient if you ask it to · Price as of August 2026: Free tier; Go $8/month; Plus $20/month; Pro from $100/month
  • Anki What it is for: Spaced repetition for the facts that have to be automatic. Not AI, and still the backbone of most finals revision · Price as of August 2026: Free on Windows, Mac, Linux and Android; the official iOS app is paid
  • PassMedicine finals bank What it is for: Over 11,000 single best answer questions for finals and the UKMLA, plus a textbook · Price as of August 2026: £15 for 4 months, £20 for 6, £25 for 9, £30 for 12. Free demo
  • Quesmed What it is for: UK question bank, mark schemes for the practical exams, video courses · Price as of August 2026: Not publicly stated. No price appears on the site; you have to register to see one
  • AMBOSS pricing page What it is for: Medical library plus question bank, with an AI mode · Price as of August 2026: Published in US dollars: $19.99/month, or $12.50/month billed yearly, for students. 5 day free trial. No separate UK price list, so check what your card is actually charged
  • what iatroX gives away free What it is for: UK focused: clinical answers sourced from NICE and CKS with citations, adaptive question banks, and a Socratic tutor · Price as of August 2026: Real free tier at £0 with no card, but the free banks are MRCP Part 1, MRCEM SBA, PSA and PARA. The UKMLA bank is in the paid plan: £29/month or £99/year
  • NICE Clinical Knowledge Summaries What it is for: Not AI at all, and you still need it. The UK reference the answers are written against · Price as of August 2026: Free to NHS clinicians and to individuals in the UK, for personal or educational use
  • Geeky Medics virtual patients What it is for: AI virtual patients from the Geeky Medics team, on their SimChat platform. History taking, counselling, difficult conversations and clinical reasoning, with the patient responding to what you actually say and automated feedback afterwards · Price as of August 2026: Not stated publicly; educators can open a free SimChat account
  • Diagnosica What it is for: An AI patient you take a history from by voice or text, then order investigations, commit to a diagnosis and a management plan, defend it to an AI senior, and get a scorecard · Price as of August 2026: Free tier of one case a week; Standard $29/month or $249/year, £22/month or £189/year on UK cards

What the price column doesn't tell you

AMBOSS lists prices in US dollars and there isn't a separate UK price list, so your statement may not match the page once your bank converts and adds fees. Check the final charge before assuming the dollar figure on the page converts cleanly.

Quesmed doesn't publish a price on its site, not on the pricing page, the FAQ or the homepage, and figures on coupon sites are unreliable. iatroX offers a free tier with no card, but the free banks are MRCP Part 1, MRCEM SBA, PSA and PARA, and the UKMLA bank a finals candidate wants sits in the paid plan, see AI tools for the UKMLA.

Geeky Medics built AI virtual patients on SimChat, and in their words the AI "responds based on what the student actually says, so the conversation unfolds differently depending on the approach taken". Students can practise histories, counselling, difficult conversations and clinical reasoning, with automated feedback against school criteria, and scenarios mapped to the GMC MLA content map, all in the browser.

Diagnosica is live but early with rough edges, it doesn't do physical examination, and it isn't the cheap option, but you can speak or type to an AI patient, take a history, order investigations, commit to a diagnosis and plan, and get scoring calibrated to the published mark sheet. "It doesn't give advice about real patients, and it's not a diagnostic system. Treat every output as educational."

The half a question bank cannot rehearse

Question banks hand you tidy vignettes with every relevant fact already on the page. The practical stations in finals are messier, because a person sits down, offers two or three loose threads, and waits for you to pull the right one. Clicking doesn't rehearse that half.

In our case library the AI patient Ryan Docherty, 24, says he's thirsty and passing large volumes of urine, and he looks wiped out. He won't tell you the thing that matters unless you ask, that he stopped his insulin 2 or 3 days ago because he felt too sick to eat. That one fact is the case.

The AI patient Martin Rowe, 55, opens with chest pain and fear he's having a heart attack. The long haul flight from Australia only appears if you ask about recent travel. The swollen, tender left calf he's had for a week surfaces only if you go hunting with direct questions about leg symptoms.

The AI patient Neil Ashworth, 45, offers back pain after lifting at work and blames his bladder trouble on codeine. You have to ask directly to hear he hasn't passed urine properly for 2 days and that he's numb between the legs. He downplays everything except the leg weakness.

In all three, the marks live in what the patient doesn't volunteer. A page can't test that, because the page has already told you. Eliciting is a spoken skill, so practise speaking, whether that's with a colleague, rehearsing stations solo, or role-playing with a chatbot.

The same goes for reasoning out loud. Recognising the right answer among five isn't the same as producing it, then holding it when someone pushes back. General chat can role-play well, but if you want fixed scripts with withheld findings, ordering, and scoring to a published mark sheet, try talking to an AI patient.

If you're practising UK medical finals by rehearsing a history out loud against a case that holds facts back, Diagnosica does that.Run a free case
The five steps of one spoken-history rehearsal: history, investigations, diagnosis, defence, scorecard

How to actually use these in the last six weeks

  1. Facts first, and not with a chatbot. Anything that must be automatic under pressure belongs in spaced repetition and a question bank. Don't use a chatbot as the source of truth for a fact you can't immediately verify.
  2. Use the chatbot as the thing you argue with, not the thing you learn from. Take a question you got wrong, ask it to unpack the step you missed, then check its explanation against a reference you trust.
  3. Verify anything you'd act on. If a claim would change what you'd do to a patient, look it up properly instead of accepting it.
  4. Then start talking. These weeks need you saying histories out loud to something that answers back, and saying your reasoning out loud to someone who pushes on it. Practise the phrases you actually say when a patient interrupts or the examiner asks you to justify a plan.
  5. Keep one thing you get marked on. Practice with no mark tends to rehearse what you already do, so use a published mark sheet or checklist and get scored against it.

It's easy to spend the last month getting AI to reorganise notes and rebuild flashcard decks, because it feels productive and it's far more comfortable than talking. That's the failure mode here, and it crowds out the one skill you'll actually perform under a clock.

What none of them can do

Physical examination. Nothing on this page examines a patient, and nothing here can tell you whether your hands are in the right place or whether you actually felt what you said you felt. Diagnosica doesn't do examination either and doesn't claim to. That stays with real patients, real tutors and each other.

The real thing. A simulated patient is never frightened in the way a real one is, never cries when you say biopsy, and never drops a comment that reorganises the whole consultation. Placement is still where that happens. These tools are practice between placements, not instead of them.

Knowing what you do not know. Every tool here answers what you ask it, it won't spot the gap you left. None of them tells you which question you failed to ask unless it's marking you against a scheme that expects it. Your examiner catches the missing thread. No bot does that for you.

Where to start

Use AI for the reading half, then check it against your handbook, the BNF, and your school's notes, including local guidance on prescribing and referrals. In the last six weeks, protect time for speaking cases out loud, the way you'll be examined, rather than only clicking through cards.

Diagnosica is one way to do the talking half. You can try the 3 minute demo, it's about three minutes, needs no signup, and marks you when the case ends. Free tier is one case a week, or run a case free, Standard is $29 monthly or $249 yearly, UK cards see £22 monthly or £189 yearly.


Educational use only, not medical advice. AI-generated; verify clinically against primary sources.

Clinical review pending.