According to a paper by Google researchers Anil Palepu, Senior Research Scientist, and Mike Schaekermann, Research Lead at Google, the company is presenting a significant advancement in its research medical AI system, AMIE (Articulate Medical Intelligence Explorer). The updated system is now designed to conduct real-time video medical consultations, demonstrating expert-level performance for the first time in a randomized controlled study based on clinical simulations.
When a physician meets a patient, the clinical encounter extends far beyond the words exchanged. The doctor observes the patient’s gait, identifies visible signs of physical discomfort, notes their breathing rhythm, and guides them through physical examinations. This continuous stream of visual and auditory information integrates with the patient's spoken medical history. These non-verbal cues, both visual and auditory, are central to effective diagnosis, building patient trust, and quality clinical communication. AI systems capable of clinical reasoning and dialogue could dramatically expand access to medical expertise and care, allowing doctors to focus on the most meaningful aspects of patient interaction.
In previous work, AMIE demonstrated expert-level performance in text-based diagnostic dialogue and proved to be an effective aid for clinicians in formulating a differential diagnosis. Recently, researchers extended AMIE's capabilities beyond diagnosis toward long-term disease treatment and management. Its capabilities were also expanded to expert-level evaluations in oncology, cardiology, and ophthalmology, as well as multimodal diagnostic reasoning integrating images and clinical documents in simulated environments with actors portraying patients. At the same time, translation to clinical practice has begun through a physician-led oversight framework and initial real-world clinical studies, including a feasibility study with Beth Israel Deaconess Medical Center and an ongoing nationwide randomized study in partnership with Included Health.
The Limitations of Text-Based Interfaces
Despite this progress, a fundamental limitation of the original research remained: text-based interfaces completely omit the visual and auditory dimensions of clinical practice. Patients must translate complex symptoms into written descriptions—a process that discards vital diagnostic information and can pose barriers for patients with limited digital or health literacy. Text-only systems cannot independently observe visual and auditory cues that inform clinical reasoning, nor can they guide patients through physical examinations that shape a differential diagnosis.
To overcome these limitations, the researchers present the system's new real-time video configuration, known as AMIE (Video). Built on Gemini and Project Astra models, AMIE (Video) conducts synchronous clinical video consultations, perceives non-verbal clinical cues, guides patient-portraying actors through virtual physical examinations, and performs diagnostic reasoning—all in real time.
Asynchronous Multi-Agent Architecture
Conducting an effective clinical conversation over video requires balancing competing demands: the system must respond to patients at natural conversational speed, while simultaneously performing deep clinical reasoning and continuously processing visual and auditory streams. Currently, a single AI agent cannot satisfy all these requirements at once; deep reasoning requires processing time, but prolonged conversational pauses erode patient trust and rapport.
To address this challenge, AMIE (Video) utilizes an asynchronous multi-agent architecture that splits the workload among three specialized agents operating continuously in parallel:
- Talker Agent: This patient-facing agent drives fast, ultra-low-latency spoken interaction. Its role is to maintain a natural conversational flow while incorporating guidance and instructions from other agents in the system.
- Planner Agent: Operating in the background, this agent continuously refines the system's clinical reasoning, updating differential diagnoses and treatment plans, identifying information gaps, and re-prioritizing clinical goals.
- Perception Agent: This agent continuously reviews the audio and video streams, identifying clinically relevant non-verbal cues (such as visible signs of distress, physical findings, or auditory signals) and contextualizing these observations within the ongoing conversation.
This structural separation allows AMIE (Video) to maintain a natural conversational pace without abnormal delays, while performing diagnostic reasoning and perceiving visual and auditory cues that would otherwise cause unacceptable latency. Automated evaluations confirm that each of these three agents makes a significant contribution to improving clinical metrics, such as competency in history-taking, clinical reasoning, and treatment recommendations, as well as metrics related to dialogue quality, including patient-centric communication skills and rapid response times.
Automated Evaluation-Guided Development
A central challenge in building audio-visual medical AI is characterizing the system’s perception and reasoning capabilities at scale. To guide development, a taxonomy of clinical audio-visual competencies relevant to telehealth was compiled from the medical literature, covering non-verbal visual cues, auditory signals, and physical examination maneuvers.
Following this, an automated evaluation system was constructed based on this taxonomy. The evaluation infrastructure combines targeted, single-turn audio-visual tests with multi-turn conversational simulations. The single-turn tests evaluate specific instances of clinical perception and reasoning (such as correctly identifying anatomical laterality or recognizing signs of respiratory distress). The multi-turn conversational simulations evaluate end-to-end dialogue performance while injecting visual cues as textual descriptions into the simulation (for example, an AI patient simulator in a Parkinson's disease scenario, prompted to present their handwriting, might inject a structured verbal description such as "[holding up paper to camera showing cramped, tiny script]"). These complementary evaluations allowed for rapid system design iterations and richly characterized the capabilities and failure modes of AMIE (Video) prior to human evaluation.
Randomized Video Study (OSCE)
To evaluate clinical competence in a more challenging and realistic setting of a full medical consultation encounter, a large-scale study based on the OSCE (Objective Structured Clinical Examination) methodology was conducted using a synchronous video interface. To ensure coverage of a wide range of medical conditions, the study encompassed 100 clinical scenarios across five different body systems: cardiopulmonary, abdominal, head/eyes/ears/throat (HEENT), neurology/psychiatry, and musculoskeletal.
Within the study, 15 trained professional actors conducted 300 standardized consultation sessions, which were randomly assigned across three research arms:
- AMIE (Video): The video configuration of the system conducting real-time video consultations.
- AMIE (Text): A text-only baseline version used to isolate the specific contribution of audio-visual capabilities.
- PCP (Video): Ten board-certified Primary Care Physicians (PCPs) who conducted consultations via the same video interface.
An independent evaluation team of 20 experienced primary care physicians evaluated all consultation sessions using established clinical rubrics, which included both general clinical competence scales and detailed case-specific scoring criteria tailored to each scenario.
Key Results and Patient Preferences
The analysis of the study results yielded significant findings:
- Expert-Level Clinical Performance: Clinical evaluators rated AMIE (Video)'s performance on par with that of human primary care physicians across core clinical competence, history-taking thoroughness, diagnostic accuracy, treatment plan appropriateness, and communication quality. Additionally, AMIE (Video)'s performance matched or exceeded AMIE (Text) in these dimensions.
- Superiority in Physical Observation and Examination: The AMIE (Video) system was rated significantly higher, on average, compared to both human physicians and the text-based version AMIE (Text) in terms of extracting physical signs and actively guiding patient actors through virtual examination maneuvers. This advantage was also reflected in the case-specific perception and examination rubric scores.
- Clear Actor Preference for the Video Interface: The actors portraying the patients expressed a strong preference for the synchronous video interface over text-based chat, rating it as significantly easier to use and more effective for communicating their health concerns. Furthermore, they gave AMIE (Video) higher scores on measures of empathy, rapport-building, and confidence in care compared to both human physicians and the text-based version.
Limitations and Responsible Development
There are important limitations to this study that must be taken into account when interpreting the results. The experiment was conducted entirely with professional actors in simulated clinical environments, rather than real patients facing actual health challenges. Actors, no matter how skilled, cannot fully replicate the complexity and unpredictability of real clinical encounters. Additionally, the scenarios were limited to medical conditions that can be reliably portrayed through acting, omitting important clinical presentations where visual and auditory perception carries critical diagnostic implications.
Beyond the scope of the main study, targeted automated evaluations revealed occasional perception and reasoning errors by the system, despite overall high-quality dialogue and diagnostic accuracy. The system still exhibits intermittent technical issues that can disrupt conversational naturalness. Given the prototype nature of Project Astra, future system-level developments may address these technological aspects, beyond the specific medical application evaluated in this study. The researchers emphasize that evaluating these findings in studies involving real patients and under real clinical conditions is a vital and necessary step before any conclusions can be drawn regarding the system's practical real-world utility.
Looking Ahead
The current study demonstrates that transitioning from text-based medical AI to an audio-visual system is achievable at expert-level quality. The AMIE (Video) system engages with the perceptual richness of clinical practice—observing non-verbal cues, guiding physical examinations, and conversing naturally through spoken dialogue—capabilities that bring the system closer to the realistic experience of a telehealth video encounter.
However, important questions remain open on the path to establishing responsible real-world evidence. The findings must be validated among real patients, trials must be expanded to clinical presentations that cannot be portrayed through acting, and the system must be supported by robust safety frameworks. Initial steps have already been taken in this direction: a real-world feasibility study in collaboration with Beth Israel Deaconess Medical Center provided early evidence of the safety and utility of the text version of AMIE in clinical practice, and the ongoing nationwide randomized study with Included Health is evaluating and assessing AI in real-world virtual care. Together, these research experiences will help inform how audio-visual capabilities can be responsibly integrated into daily clinical practice.