Google AMIE Matches Doctors On Key Measures In Controlled Video Test
AI News reported that Google tested AMIE in synchronous video consultations with trained patient actors, where physician evaluators rated the system against primary care doctors in controlled scenarios.

Google's AMIE video research moved medical AI from text chat into a live consultation test, with AI News reporting that professional patient actors and clinical evaluators rated the system alongside primary care physicians in controlled scenarios.
The system did not enter real patient care.
Google used 15 trained actors, five body-system categories and prepared consultation cases, while the company said real-patient studies are still needed before clinical use can be judged.
Multi-Agent Design Keeps The Visit Moving
AMIE's video version separates the consultation into three cooperating agents.
A talker agent handles the spoken exchange, a planner agent updates possible diagnoses and management goals, and a perception agent reads video and audio signals for physical or non-verbal clues.
That split addresses a practical telehealth problem rather than a model-size claim.
Detailed reasoning can slow a conversation, so the patient-facing agent can continue speaking while planning and perception continue in the background.
Google reported that automated evaluations showed gains from each agent across history-taking, clinical reasoning, treatment recommendations, patient-centred communication and response latency.
Physician Review Set The Controlled Benchmark
For the human benchmark, AMIE video, text AMIE and physician visits were judged under a common consultation setup.
The physician comparator included 10 board-certified primary care doctors, and a separate 20-member primary-care panel scored each encounter with established rubrics and case-specific criteria.
Those scores placed the system near physicians on taking histories, reaching diagnoses, selecting management steps and communicating with patients.
The video version also matched or exceeded the text-only system on those measures.
The video interface gave AMIE its strongest reported advantage when consultations required visual or physical-examination guidance.
Evaluators rated it higher than both comparison groups for eliciting physical signs and guiding actors through virtual examination manoeuvres.
Patient actors also preferred the video format over text chat, with Google reporting higher ratings for ease of use, communication of health concerns, empathy, rapport and confidence in care.
Controlled Evidence Leaves A Deployment Gap
The development path still relies on simulation before clinical deployment.
Google's automated suite tested visual cues, auditory signals and physical examination tasks, while some multi-turn simulations used text descriptions of visual input rather than an end-to-end live video stream.
Prepared actors also limit the evidence.
They cannot reproduce the variability of real patients, and some presentations were excluded because actors could not portray them authentically.
Google also noted occasional perception and reasoning errors and intermittent technical issues that affected conversational naturalness.
The related clinical work is still earlier and narrower than a commercial rollout.
Text-based AMIE has been studied with Beth Israel Deaconess Medical Center for feasibility evidence on safety and utility, and Google is also running a nationwide randomised study with Included Health.
Those projects matter because the video result depends on whether the same reasoning, perception and conversational timing can survive outside rehearsed cases.
For hospitals, insurers and virtual-care platforms, the operating question is not only whether a model can answer correctly.
A live clinical assistant must decide when to ask follow-up questions, when to use visual information, when to slow down for safety and when to hand control back to a clinician.
The practical result is a stronger research case for AI-assisted video consultation, not proof of production readiness.
AMIE has shown controlled performance on telehealth behaviour, examination guidance and physician scoring; the next clinical test is whether those results hold with real patients and real care settings.




















