Guide

Building a Voice Travel Concierge with Bedrock AgentCore & Nova Sonic

An AWS guide presents an architecture based on Bedrock AgentCore, the Nova 2.5 Sonic model, and the MCP protocol

4 min read
עב
Building a Voice Travel Concierge with Bedrock AgentCore & Nova Sonic
Based on original reporting byAWS Machine Learning ↗Translated and summarized by our AI-assisted news systemHow we work

Executive summary

What to know

  1. The solution is based on Amazon Bedrock AgentCore, the Nova 2.5 Sonic model, and Amazon Bedrock Knowledge Bases.

  2. The connection between the agent and backend services is handled via the Model Context Protocol (MCP) to maintain loose coupling.

  3. The architecture separates the front end, AI agent, and backend services into three distinct layers, and is deployed using AWS CDK.

  4. Voice interaction occurs in real time over a direct WebSocket connection using 16 kHz PCM format.

In a technical post published on the AWS blog, Ravi Kumar presented an architecture for building a voice travel concierge for airline applications. The solution is built on three managed services: Amazon Bedrock AgentCore, which serves as a platform for building, deploying, and operating AI agents in a secure environment; Amazon Nova Sonic on Bedrock, a speech-to-speech model for real-time voice; and Amazon Bedrock Knowledge Bases, a managed RAG service for grounding responses in policy documents. The concierge allows travelers to carry out tasks such as changing a seat, checking delays, updating meal preferences, inquiring about policies, and escalating to a human agent upon request, while integrating alongside existing screens in the application.

System Architecture and Layer Separation

The presented architecture separates the front end, the AI agent, and the backend services into distinct layers, enabling each component to be developed and scaled independently. To link the agent to backend services, the architecture uses the Model Context Protocol (MCP), an open standard for standardized message passing that maintains loose coupling.

The solution deploys several AWS services:

  1. Amazon Cognito – Handles user authentication and provides temporary AWS credentials for signed API access.
  2. Amazon Bedrock AgentCore runtime – Hosts the agent with microVM-level isolation for each session.
  3. Amazon Bedrock AgentCore Gateway – Exposes backend endpoints as discoverable MCP tools.
  4. Amazon API Gateway – Publishes the backend as REST endpoints with AWS Identity and Access Management (IAM) authorization.
  5. AWS Lambda – Runs business logic for itineraries, seat maps, passenger updates, flight status, loyalty programs, policy lookups, and escalations.
  6. Amazon DynamoDB – Stores customer profiles, bookings, passengers, seat maps, purchase history, preferences, conversation transcripts, and flight data.
  7. Amazon Bedrock Knowledge Bases – Answers policy questions by grounding responses in airline documents.
  8. Amazon Simple Email Service (SES) – Sends email notifications.
  9. AWS Amplify – Hosts the React front end of the application.

The infrastructure is split into four sections in AWS CDK: Section A includes five CDK stacks for the backend infrastructure; Section B includes one stack to create the AgentCore Gateway with the MCP protocol; Section C includes two stacks to provision the runtime infrastructure, including the use of Amazon ECR for the container image, Amazon S3 for source code uploads, and AWS CodeBuild to produce an ARM64 Docker image; and Section D includes a stack for deploying the React application on AWS Amplify.

User Request Flow and WebSocket Connection

The interaction flow begins when the user opens the web application on AWS Amplify and enters login credentials. Amazon Cognito authenticates the request and returns JWT tokens and temporary AWS credentials. Subsequently, the user interface opens a SigV4-signed WebSocket connection directly to Amazon Bedrock AgentCore.

The runtime validates the token against Cognito and initializes Amazon Nova 2.5 Sonic. When the user speaks, audio is streamed in 16 kHz PCM format over the WebSocket. The Nova 2.5 Sonic model processes the speech and triggers tool calls. The agent invokes the AgentCore Gateway using MCP to retrieve flight data or perform actions, and the Gateway translates the calls into REST API requests forwarded to API Gateway and Lambda functions. Data is retrieved from DynamoDB, and Nova 2.5 Sonic generates a voice response streamed back to the user over the WebSocket connection.

Voice Processing Capabilities with Amazon Nova 2.5 Sonic

The solution utilizes Amazon Nova 2.5 Sonic, a speech-to-speech model featuring real-time reasoning capabilities. The model provides speech recognition across a variety of accents, robustness against background noise, spoken responses that adapt to the passenger's tone of voice, and low-latency bidirectional streaming.

In addition, the system includes asynchronous tool calling in parallel without pausing the conversation, latency masking through the generation of interim spoken responses while awaiting results, barge-in capability, natural turn-taking management, and context retention across multiple turns. The agent is configured to follow a confirm-before-write pattern, requiring traveler confirmation prior to executing changes, and reads flight numbers and confirmation codes character by character.

Answering Policy Questions and Escalation to a Live Agent

For inquiries regarding baggage, change fees, pet travel, and loyalty terms, the agent connects to Amazon Bedrock Knowledge Bases through a dedicated connector on the AgentCore Gateway. Policy documents are uploaded to Amazon S3, and the service manages embedding, chunking, indexing, storage in Amazon S3 Vectors, and retrieval. A Smart Parsing mechanism prepares PDF files so that tables and complex layouts are retrieved accurately.

When a user requests to speak with a human agent, or when the agent cannot fulfill the request, the EscalateToAgent tool is triggered after receiving user confirmation. A Lambda function logs the escalation in DynamoDB and returns a reference number and estimated wait time. The Amplify interface then initiates dialing to the configured support number from the user's device.

Prerequisites and Deployment

Deployment is performed using a single AWS CDK script. Prerequisites include an AWS account, access to the Amazon Nova 2.5 Sonic model in the deployment region, Node.js version 20.x or later, Python version 3.12 or later, AWS CLI version 2.x, and AWS CDK CLI version 2.x. System monitoring is handled via Amazon CloudWatch, and encryption of data at rest is managed using AWS KMS. Resource teardown is executed in reverse order of deployment using a dedicated script.

Was this useful for your business?

Questions & Answers

FAQ

This article was produced by our AI-assisted system through translation, summarization, and automated quality controls based on original reporting by AWS Machine Learning. Read about our editorial process. Link to the original source.

Get useful AI updates by email

A concise digest from our news desk.

More from AWS Machine Learning

All articles from AWS Machine Learning
ארכיטקטורת סוכני AI לניתוח חוזים עם Amazon Bedrock ו-Quick
ניתוח
4 דקות
מ־AWS Machine Learning

ארכיטקטורת סוכני AI לניתוח חוזים עם Amazon Bedrock ו-Quick

בפוסט שפורסם בבלוג של AWS הציגו מהנדסי החברה ארכיטקטורה לפלטפורמת ניתוח חוזים, המשלבת סוכני בינה מלאכותית מבוססי Amazon Bedrock AgentCore וכלי תשאול וניתוח ב-Amazon Quick. הפתרון מתמודד עם מגבלות כלי RAG בעת ביצוע חישובי אגרגציה על מאות מסמכים, באמצעות חילוץ שדות מפתח למסד נתונים מובנה ב-Amazon Aurora PostgreSQL. המערכת משתמשת בסוכן חילוץ מבוסס Claude Sonnet ובסוכן אימות מבוסס Claude Haiku, לצד Amazon Textract כגורם מכריע לזיהוי חתימות בעזרת ראייה ממוחשבת. הגישה מאפשרת לבצע הן שאילתות רוחביות והן איתור מקטעים מתוך מסמך יחיד בממשק מאוחד.

קרא עוד
כיצד HEMA בנתה שכבת ידע ארגונית עם Bedrock ו-MCP
ניתוח
4 דקות
מ־AWS Machine Learning

כיצד HEMA בנתה שכבת ידע ארגונית עם Bedrock ו-MCP

רשת הקמעונאות ההולנדית HEMA בנתה שכבת ידע פנימית המבוססת על Amazon Bedrock AgentCore ו-Model Context Protocol (MCP) במטרה לאחד מידע מבוזר ולמנוע מעבר ידני בין פורטלים ומערכות ויקי שונות. העוזר הפנימי HAL, שפותח תחילה ככלי עצמאי מבוסס Next.js ו-Strands, הורחב לשימוש ישיר מתוך כלי העבודה של המהנדסים (כגון Kiro ו-Claude) באמצעות שער Entra MCP ייעודי ופרוקסי אימות. המערכת משרתת כיום מפתחים, מנהלי מוצר ומנתחי מערכות, כאשר השלב הבא מתוכנן להרחיב את יכולות העוזר ממענה לשאלות לביצוע פעולות תפעוליות ישירות מתוך ממשקי השיחה.

קרא עוד
אוסף מיומנויות סוכן פתוח מבית AWS לשיפור הסקת מסקנות בבריאות
מחקר
5 דקות
מ־AWS Machine Learning

אוסף מיומנויות סוכן פתוח מבית AWS לשיפור הסקת מסקנות בבריאות

בפוסט שפורסם ב-AWS הוצג אוסף של 38 מיומנויות סוכן (Agent Skills) בקוד פתוח ב-11 תחומי בריאות ומדעי החיים (HCLS) תחת רישיון MIT-0. המיומנויות בנויות כקובצי Markdown מובנים ומסווגות למיומנויות הסקה ולמיומנויות צינור, הניתנות להרצה על יותר מ-20 שירותים, כולל Amazon Bedrock AgentCore, AWS Strands SDK ו-Kiro CLI. הערכה השוואתית שבוצעה על 410 פרומפטים הראתה כי סוכנים המצוידים במיומנויות השיגו שיעור ניצחון של 69.5% עד 85.9% מול סוכני בסיס ללא מיומנויות, כאשר השיפור המשמעותי ביותר נמדד בממד החשיבה הביקורתית (שיעור ניצחון של 78% עד 85.1%). בנוסף, המיומנויות הפחיתו את שונות הציונים בעד 61.9%.

קרא עוד
אופטימיזציית עלויות וזמני תגובה עם Prompt Caching ב-Bedrock
מדריך
3 דקות
מ־AWS Machine Learning

אופטימיזציית עלויות וזמני תגובה עם Prompt Caching ב-Bedrock

בפוסט של ארכיטקט הפתרונות דניאל אביב מ-AWS, מוסבר כיצד מנגנון ה-Prompt Caching ב-Amazon Bedrock מפחית עד 90% מעלויות טוקני הקלט על פגיעות במטמון ומקצר את זמן התגובה לטוקן הראשון (TTFT). המאמר סוקר שישה תרחישי יישום באמצעות ה-Converse API: שמירת מסמכים, שמירת פרומפט מערכת, שמירת הגדרות כלים לסוכנים, שילוב זמני חיים שונים (Mixed TTL), בידוד דיירים במערכות מרובות משתמשים באמצעות תחילית SHA-256, ואינטגרציה עם ספריית LangChain. מודלי Anthropic Claude Sonnet 4.5 ו-4.6 דורשים סף מינימלי של 1,024 טוקנים להפעלת המטמון.

קרא עוד

More articles you might like

All articles
מדריך כלי ה-AI ללא קוד לעסקים קטנים לפי Salesforce
מדריך
4 דקות
מ־Salesforce Blog

מדריך כלי ה-AI ללא קוד לעסקים קטנים לפי Salesforce

מאמר שפורסם על ידי חברת Salesforce סוקר את תחום כלי ה-AI ללא קוד (No-Code AI) עבור עסקים קטנים ובינוניים. לפי נתוני דוח המגמות המצוטט במאמר, 75% מהעסקים הקטנים משקיעים בבינה מלאכותית, אך 88% מתוכם עדיין נמצאים בשלב הבחינה. המאמר מפרט קריטריונים לבחירת כלים, סוקר פתרונות בתחומי המכירות, השיווק, האוטומציה, השירות והדוחות, ומדגיש את היתרון של פלטפורמה מחוברת שבה 91% ממשתמשי ה-AI מדווחים על גידול בהכנסות לעומת שימוש בכלים נקודתיים מבודדים.

קרא עוד
בדיקת פרומפטים ליישומי LLM: מדריך n8n לזיהוי רגרסיות
מדריך
4 דקות
מ־n8n

בדיקת פרומפטים ליישומי LLM: מדריך n8n לזיהוי רגרסיות

מדריך שפורסם על ידי n8n מפרט כיצד מסגרות עבודה לבדיקת פרומפטים מאפשרות לאתר רגרסיות ביישומי LLM לפני עלייתם לסביבת הייצור. בשל האופי הבלתי-דטרמיניסטי של מודלי שפה, בדיקות התאמה מדויקת מסורתיות אינן מספקות. המדריך סוקר כלים נפוצים בתחום, מבחין בין שיטות הערכה דטרמיניסטיות לבין שימוש ב-LLM כשופט, ומציג כיצד לבצע בדיקות והשוואות מול קו בסיס ישירות בתוך פלטפורמת n8n.

קרא עוד
אופטימיזציית עלויות וזמני תגובה עם Prompt Caching ב-Bedrock
מדריך
3 דקות
מ־AWS Machine Learning

אופטימיזציית עלויות וזמני תגובה עם Prompt Caching ב-Bedrock

בפוסט של ארכיטקט הפתרונות דניאל אביב מ-AWS, מוסבר כיצד מנגנון ה-Prompt Caching ב-Amazon Bedrock מפחית עד 90% מעלויות טוקני הקלט על פגיעות במטמון ומקצר את זמן התגובה לטוקן הראשון (TTFT). המאמר סוקר שישה תרחישי יישום באמצעות ה-Converse API: שמירת מסמכים, שמירת פרומפט מערכת, שמירת הגדרות כלים לסוכנים, שילוב זמני חיים שונים (Mixed TTL), בידוד דיירים במערכות מרובות משתמשים באמצעות תחילית SHA-256, ואינטגרציה עם ספריית LangChain. מודלי Anthropic Claude Sonnet 4.5 ו-4.6 דורשים סף מינימלי של 1,024 טוקנים להפעלת המטמון.

קרא עוד
15 דרכים לשימוש בסוכני AI לניהול רשתות חברתיות לפי Salesforce
מדריך
4 דקות
מ־Salesforce Blog

15 דרכים לשימוש בסוכני AI לניהול רשתות חברתיות לפי Salesforce

מדריך של חברת Salesforce מפרט 15 דרכים שבהן סוכני בינה מלאכותית לרשתות חברתיות מסייעים לעסקים קטנים ובינוניים. הכלים האוטונומיים מאפשרים יצירת תוכן בקול המותג, תזמון פוסטים בזמנים מותאמים אישית, מענה אוטומטי לשאלות נפוצות 24/7, ניתוב פניות מורכבות לנציגים אנושיים, ניטור אזכורים וסנטימנט, וחיבור מעורבות ישירות למערכות ה-CRM לצורך יצירת לידים. בנוסף מובאת דוגמת חברת reMarkable, שטיפלה ביותר מ-18,000 שיחות שירות באמצעות סוכני AI.

קרא עוד