When you visit a doctor, the consultation extends far beyond words: a physician notices a cough, observes gait, registers visible signs of discomfort. Google Research and Google DeepMind have now moved their research medical AI system, AMIE, into exactly that dimension.

What it can do

Built on Gemini and Project Astra using a multi-agent architecture, AMIE now does three things at once:

  • Interprets visual and auditory cues
  • Guides virtual physical exams
  • Reasons diagnostically in real time

The company describes this as a first-of-its-kind demonstration: expert-level capability in a real-time clinical video consultation setting.

How the study was set up

The evaluation used a randomised study. Simulated consultations were conducted with patient actors, with a group of primary care physicians serving as the comparison.

Clinical evaluators assessed AMIE favourably across core clinical competencies. The measured categories were:

  • Thoroughness of history-taking
  • Diagnostic accuracy
  • Appropriateness of management
  • Communication quality

One of the more striking findings concerns patient preference: the patient actors preferred the video experience over text chat.

Why video makes a difference

The logic behind that finding is about what a text interface loses in a medical consultation. A physician assesses not only what is said but what is not: pallor, the rhythm of breathing, how much effort a movement takes. Text chat erases that entire layer.

Guiding a virtual physical exam is a separate threshold. In a remote consultation the examination is actually performed by the patient; the system's job is to ask for the right thing in the right order and interpret the result.

The limits are stated plainly

Google emphasises that AMIE remains a research system and that more research is needed before responsible real-world clinical deployment. The study used simulated consultations and patient actors; what happens with real patients, under real uncertainty and real risk, is a separate question.

The direction is nonetheless clear: medical AI is moving out of the text box and into the video call, and that changes how such systems must be evaluated.

Evaluation itself gets harder

Evaluating a medical AI system through text is comparatively easy: questions and answers are written down and can be compared. In a video consultation there is far more to assess — not only what the clinician asks but what they notice, what they ask the patient to do, and how they interpret what they see.

That is why the study's design matters: patient actors on one side, a group of primary care physicians as the comparison on the other. Three of the four measured categories are technical competence; the fourth is communication quality — that is, how the patient felt.

The patient actors' preference for video over text may belong to that last category. Describing a health problem is not only information transfer; seeing that the other side is listening is part of the process.