Google's AMIE (Video) matched or beat primary care physicians across four clinical dimensions in a randomised trial of 100 video consultation scenarios, according to a preprint posted on arXiv on 10 August . Clinical evaluators rated the Gemini-based AI system on par or better than 30 human doctors in history-taking, diagnosis, management, and physical observation. Yet the same study found patients still preferred humans for something specific, and the entire experiment ran on actors in a simulated clinic, not real patients.

My read: This is the first medical AI system I've seen that claims expert-level performance in real-time video consultations, and the OSCE design is a serious way to test it. But "expert-level" is the authors' own framing, not an independent clinical accreditation, and the study uses a taxonomy and evaluation framework the authors built themselves . I'd want to see this replicated with real patients in actual clinics before drawing any deployment conclusions. The Google affiliation matters too: AMIE is Gemini-based, and the deep research notes show the author list includes Google DeepMind and Harvard Medical School researchers P⁵, so the system's creator and the study's evaluator share institutional ties.

How the trial was built

The study used an Objective Structured Clinical Examination, or OSCE, the standard format medical schools use to test clinical skills with standardised patients. Thirty primary care physicians, 15 patient actors, and 100 clinical scenarios were randomised across three arms: AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video .

AMIE (Video) is a multi-agent system built on Google's Gemini model, combining low-latency dialogue with clinical reasoning and real-time audio-visual perception . In plain terms, it can see the patient on camera, hear them speak, reason through the clinical picture, and talk back in real time. The researchers also built a taxonomy and automated evaluations for clinical audio-visual cues in telehealth, because no standard framework existed for scoring what matters in a video consultation: the visual examination, the way a patient moves, the subtle cues a doctor reads from a face .

Where the AI won and where humans held the line

Clinical evaluators, working from the authors' framework, rated AMIE (Video) on par or better than primary care physicians in history-taking, diagnosis, management, and physical observation and examination . Patient actors preferred AMIE's approach to assessing and explaining conditions .

But primary care physicians were preferred for rapport and partnership building . That gap matters. A patient who feels no human connection may get accurate clinical answers but still not follow through on treatment, return for follow-up, or trust the system when something feels wrong. Whether AMIE's clinical accuracy survives the messiness of real patient behaviour outside a scripted scenario is an open question.

Why video beat text

In a modality ablation, patient actors preferred AMIE (Video)'s interface over text chat, rating it higher for communicative effectiveness and convenience, and saying they felt more understood . This is the part that should worry text-only telehealth platforms. If patients consistently feel better understood by an AI that can see and speak than by one that types, the text chatbot model that many health platforms currently use starts to look like a dead end.

The system has real limitations. The authors flag problems with fine anatomical precision and subtle affective nuances, along with high-frequency movements . A doctor watching a patient walk into the consulting room reads a gait, a limp, a tremor. AMIE (Video) struggles with exactly those fast, subtle physical signals.

What to do about it

A regional telehealth provider running after-hours GP services should watch this closely. If AMIE (Video) can match PCPs on clinical dimensions in a simulated setting, the pressure on after-hours staffing models is real, even before the system is cleared for deployment. A pathology lab that uses video for remote specimen review faces a different question: the system's weakness in fine anatomical precision means it is not ready to replace a trained eye for reading slides or examining tissue on camera.

For now, the practical move is to read the preprint itself, available at arxiv.org/abs/2608.09861v1, and check the authors' affiliations and conflicts of interest in the full text. The "expert-level" claim is the authors' framing based on their own study metrics, not an independent clinical accreditation .

What we don't know yet

The study is a preprint, not peer-reviewed . The sample of 30 physicians and 15 actors may be too small to generalise, and the OSCE format with patient actors is not a real clinical environment . The comparative performance claims depend on a taxonomy and evaluation framework the authors developed themselves . The authors state that further research is needed before real-world translation , and the system cannot yet handle the full sensory range of clinical practice.

The next signal: whether this preprint survives peer review and independent replication with real patients in live clinics. We'll check both against the claims here when the published version appears. Subscribe and we'll track whether AMIE's simulated wins survive contact with real patients.


Sources: S1 — Towards Expert-level Medical AI for Real-time Video Consultations · P2 — Towards Conversational Medical AI with Eyes, Ears and a Voice · P3 — joeljang/ELM · P4 — UCSC-VLAA/MedReason · P5 — Towards Conversational Medical AI with Eyes, Ears and a Voice

More from Not A Tech Guy


Generated from an audited evidence pack with primary-source research. Social-media items are discussion signals, not verified facts. Nothing here is financial, legal or medical advice.