Guide

Building an AI- and Avatar-Based Knowledge System on AWS

An AWS guide presents a cloud RAG architecture with an avatar interface and smart caching for institutional knowledge

4 min read
עב
Building an AI- and Avatar-Based Knowledge System on AWS
Based on original reporting byAWS Machine Learning ↗Translated and summarized by our AI-assisted news systemHow we work

Executive summary

What to know

  1. The solution combines Amazon Bedrock Knowledge Bases, Amazon S3, and OpenSearch Serverless for institutional data retrieval.

  2. The user interface supports voice and text interactions through an avatar powered by DeepBrain AI technology and WebRTC.

  3. In testing, for workloads where 50 to 70 percent of questions were repeats, using caching reduced inference costs by a comparable rate.

  4. S3 event notifications trigger automated Lambda synchronization to update the knowledge base when new documents are uploaded.

In a post published on the AWS blog, a cloud-based solution featuring a smart caching mechanism was presented, designed to capture, maintain, and deliver accumulated institutional knowledge through an intelligent avatar system running on AWS services. According to the post's authors, organizations across diverse industries struggle with managing institutional knowledge and experience accumulated over years of operations, which is often lost when key employees leave. Traditional documentation methods frequently prove inadequate, leaving outdated or inaccessible information.

Target Audiences and System Use Cases

According to the post, organizations across diverse industries can use the system to preserve critical knowledge. In manufacturing organizations, for example, production procedures and maintenance protocols can be captured before experienced technicians retire. Additionally, healthcare facilities, financial services firms, energy companies, and government agencies can tailor the solution to their needs. Knowledge workers can access procedures and policies through natural language queries instead of searching multiple repositories, while subject matter experts or retiring employees can upload documentation to preserve their expertise for future generations.

The system can be deployed with desktop browser access for detailed research, voice interaction for hands-free operation, and text-based queries for quick reference depending on the operational context of each industry.

Solution Architecture and Cloud Components

At the foundation of the solution is a browser-based interface supporting both voice and text interactions, connecting to a configurable avatar system that works with third-party avatar solutions (implemented here based on DeepBrain AI).

Behind the scenes, the following AWS services operate:

  • Amazon Cognito secures access management and user controls.
  • Amazon API Gateway provides managed, monitored endpoints for communication between system components.
  • Amazon Bedrock Knowledge Bases serves as the knowledge-processing core for managed Retrieval Augmented Generation (RAG): data stored in Amazon S3 serves as the data source, and Bedrock performs chunking, vector generation (embedding) using the Amazon Titan Text Embeddings model, and document-grounded retrieval.
  • Amazon OpenSearch Serverless serves as a vector store for the Bedrock knowledge base.
  • Amazon DynamoDB provides a response caching layer.
  • AWS Lambda functions orchestrate request and response workflows.
  • Amazon Transcribe converts voice input to text.
  • Amazon Polly converts text responses to natural speech.
  • Avatar video streaming operates over WebRTC alongside WebSocket connections for response control.

Regarding costs, the post noted that the OpenSearch Serverless vector store is billed per OpenSearch Compute Unit (OCU) with an always-on minimum threshold independent of query volume, representing a fixed baseline cost of a few hundred dollars per month in the default configuration, on top of which are variable language model costs that the caching mechanism is designed to reduce.

Implementation Phases of the Solution

The setup process is described in three main phases:

  1. Knowledge Foundation Setup: Documents are uploaded to a repository in Amazon S3 in supported formats such as Word, PDF, plain text, Markdown, or JSON (with Markdown and JSON recommended for optimal retrieval performance). An AWS Glue ETL job can optionally be used to convert documents from other formats before uploading.
  2. Infrastructure Deployment: Infrastructure deployment using AWS CloudFormation templates, which includes setting up Amazon Cognito for user authentication and access controls, configuring API Gateway endpoints, establishing S3 repositories for knowledge storage, and implementing DynamoDB for response caching.
  3. AI Integration: Configuring Amazon Bedrock, connecting Lambda functions for orchestration, setting up the audio processing pipeline with Transcribe and Polly, and integrating the avatar system.

Initial deployment requires appropriate AWS IAM permissions for all involved services. The foundation models used in the implementation are Amazon Titan Text Embeddings for document encoding, as well as Amazon Nova Pro and Anthropic Claude 3 Sonnet as selectable models for generating responses. According to the post, Amazon Bedrock provides automatic access to Amazon-owned serverless models, so Amazon Nova Pro and Amazon Titan Text Embeddings are available by default with no manual enablement; conversely, Anthropic Claude models are not covered by automatic access in all Regions, and before the first Claude model request, a one-time model-access request must be completed in the Bedrock console and appropriate AWS Marketplace subscription permissions ensured.

Performance Optimization and the Caching Mechanism

The system implements a dual-layer caching strategy:

  • A browser-side cache using an LRU algorithm in browser memory for instant responses to repeated questions.
  • A server-side cache in DynamoDB with intelligent Time-To-Live (TTL) settings tailored to content type: longer retention for fundamental knowledge, medium-term for operational procedures, and short-term for time-sensitive information.

In the current implementation, cache matching is performed based on exact query text matching. In AWS testing, for workloads with 50 to 70 percent repeated questions, the cache hit rate resulted in a comparable reduction in variable language model inference costs.

Knowledge Updates, Accuracy Considerations, and Limitations

To keep the system current, Amazon S3 event notifications are configured: adding or removing a document triggers a Lambda function that initiates synchronization and re-indexing in Bedrock Knowledge Bases without manual intervention.

In terms of accuracy, Bedrock grounds answers in passages retrieved from the organization's verified documents and can return source citations for verification. However, the post emphasizes that this grounding reduces but does not completely eliminate the risk of incorrect answers, so for high-consequence or safety-related decisions, a human should be kept in the loop and the system should be treated as decision support rather than an exclusive authoritative source.

Additionally, this architecture depends on continuous cloud connectivity and includes no offline capability in this prototype, making it suitable for connected environments (such as training rooms, quality labs, and control rooms) rather than network-disconnected plant floors. In testing, the default configuration was estimated to support roughly 50 to 100 concurrent users, with the figures representing indicative estimates only and not guarantees, depending on Region, model choice, document size, cache hit rate, and service quotas; enterprise-scale expansion to over 5,000 concurrent users is achievable with the same architecture but requires proactive service-quota increases.

Was this useful for your business?

Questions & Answers

FAQ

This article was produced by our AI-assisted system through translation, summarization, and automated quality controls based on original reporting by AWS Machine Learning. Read about our editorial process. Link to the original source.

Get useful AI updates by email

A concise digest from our news desk.

More from AWS Machine Learning

All articles from AWS Machine Learning
תשלום לפי קריאה עבור סוכני AI: שילוב AgentCore payments ב-AWS
מוצר חדש
4 דקות
מ־AWS Machine Learning

תשלום לפי קריאה עבור סוכני AI: שילוב AgentCore payments ב-AWS

בפוסט שפרסמו מהנדסי AWS הוצגה היכולת המנוהלת AgentCore payments ב-Amazon Bedrock, המאפשרת לסוכני בינה מלאכותית לשלם לפי דרישה עבור שירותים דוגמת הסקת מודלים. השירות מנהל ארנקים, מטפל בפרוטוקול התשלום x402 ואוכף מגבלות הוצאה ברמת התשתית במקום ברמת המודל. הפוסט מפרט כיצד חברת Incarna שילבה את המערכת כדי לאפשר לסוכנים לשלם לנתב ההסקה BlockRun עבור כל קריאה בנפרד ברשת Base ב-USDC. האינטגרציה הושלמה בשלושה ימים ובכ-200 שורות קוד, ובשלב הבטא עובדו מעל 1,000 תשלומים בסכומים של 0.001 עד 0.05 דולר לקריאה. היכולת זמינה כעת באופן כללי.

קרא עוד
Claude Haiku 5.5 זמין כעת ב-Amazon Bedrock וב-AWS
מוצר חדש
4 דקות
מ־AWS Machine Learning

Claude Haiku 5.5 זמין כעת ב-Amazon Bedrock וב-AWS

חברת AWS הודיעה על זמינות המודל Claude Haiku 5.5 ב-Amazon Bedrock וב-Claude Platform on AWS. לפי Anthropic, המודל הוא המהיר והיעיל ביותר במשפחת Claude 5.5, נבנה עבור תתי-סוכנים ומשימות בנפח גבוה, ועולה כ-75 אחוז פחות מ-Claude Haiku 4.5 במרבית המשימות. השירות ב-Bedrock שומר על תושבות נתונים אזורית ומשתלב עם מנגנוני IAM, CloudTrail, CloudWatch ו-Guardrails. המודל כולל בקרות מאמץ לכוונון עלות מול תבונה, ומיועד להשתלב לצד Claude Opus 5.5 שמבצע את התכנון והשיפוט המורכב. המודל זמין בפרופילי הסקה גלובליים ואזוריים וב-AWS GovCloud.

קרא עוד
בניית סוכן נסיעות קולי עם Bedrock AgentCore ו-Nova Sonic
מדריך
4 דקות
מ־AWS Machine Learning

בניית סוכן נסיעות קולי עם Bedrock AgentCore ו-Nova Sonic

בפוסט טכני של AWS הציגו מומחי החברה ארכיטקטורה לפריסת סוכן נסיעות קולי לחברות תעופה. הפתרון משלב את Amazon Bedrock AgentCore לניהול והרצת סוכנים, מודל הדיבור Amazon Nova 2.5 Sonic לקול בזמן אמת, ו-Amazon Bedrock Knowledge Bases למענה על שאלות מדיניות מתוך מסמכים. המערכת מקשרת בין הסוכן לשירותי ה-backend באמצעות פרוטוקול Model Context Protocol (MCP) ומאפשרת טיפול בהזמנות, החלפת מושבים והסלמה לנציג אנושי.

קרא עוד
ארכיטקטורת סוכני AI לניתוח חוזים עם Amazon Bedrock ו-Quick
ניתוח
4 דקות
מ־AWS Machine Learning

ארכיטקטורת סוכני AI לניתוח חוזים עם Amazon Bedrock ו-Quick

בפוסט שפורסם בבלוג של AWS הציגו מהנדסי החברה ארכיטקטורה לפלטפורמת ניתוח חוזים, המשלבת סוכני בינה מלאכותית מבוססי Amazon Bedrock AgentCore וכלי תשאול וניתוח ב-Amazon Quick. הפתרון מתמודד עם מגבלות כלי RAG בעת ביצוע חישובי אגרגציה על מאות מסמכים, באמצעות חילוץ שדות מפתח למסד נתונים מובנה ב-Amazon Aurora PostgreSQL. המערכת משתמשת בסוכן חילוץ מבוסס Claude Sonnet ובסוכן אימות מבוסס Claude Haiku, לצד Amazon Textract כגורם מכריע לזיהוי חתימות בעזרת ראייה ממוחשבת. הגישה מאפשרת לבצע הן שאילתות רוחביות והן איתור מקטעים מתוך מסמך יחיד בממשק מאוחד.

קרא עוד

More articles you might like

All articles
בניית סוכן נסיעות קולי עם Bedrock AgentCore ו-Nova Sonic
מדריך
4 דקות
מ־AWS Machine Learning

בניית סוכן נסיעות קולי עם Bedrock AgentCore ו-Nova Sonic

בפוסט טכני של AWS הציגו מומחי החברה ארכיטקטורה לפריסת סוכן נסיעות קולי לחברות תעופה. הפתרון משלב את Amazon Bedrock AgentCore לניהול והרצת סוכנים, מודל הדיבור Amazon Nova 2.5 Sonic לקול בזמן אמת, ו-Amazon Bedrock Knowledge Bases למענה על שאלות מדיניות מתוך מסמכים. המערכת מקשרת בין הסוכן לשירותי ה-backend באמצעות פרוטוקול Model Context Protocol (MCP) ומאפשרת טיפול בהזמנות, החלפת מושבים והסלמה לנציג אנושי.

קרא עוד
מדריך כלי ה-AI ללא קוד לעסקים קטנים לפי Salesforce
מדריך
4 דקות
מ־Salesforce Blog

מדריך כלי ה-AI ללא קוד לעסקים קטנים לפי Salesforce

מאמר שפורסם על ידי חברת Salesforce סוקר את תחום כלי ה-AI ללא קוד (No-Code AI) עבור עסקים קטנים ובינוניים. לפי נתוני דוח המגמות המצוטט במאמר, 75% מהעסקים הקטנים משקיעים בבינה מלאכותית, אך 88% מתוכם עדיין נמצאים בשלב הבחינה. המאמר מפרט קריטריונים לבחירת כלים, סוקר פתרונות בתחומי המכירות, השיווק, האוטומציה, השירות והדוחות, ומדגיש את היתרון של פלטפורמה מחוברת שבה 91% ממשתמשי ה-AI מדווחים על גידול בהכנסות לעומת שימוש בכלים נקודתיים מבודדים.

קרא עוד
בדיקת פרומפטים ליישומי LLM: מדריך n8n לזיהוי רגרסיות
מדריך
4 דקות
מ־n8n

בדיקת פרומפטים ליישומי LLM: מדריך n8n לזיהוי רגרסיות

מדריך שפורסם על ידי n8n מפרט כיצד מסגרות עבודה לבדיקת פרומפטים מאפשרות לאתר רגרסיות ביישומי LLM לפני עלייתם לסביבת הייצור. בשל האופי הבלתי-דטרמיניסטי של מודלי שפה, בדיקות התאמה מדויקת מסורתיות אינן מספקות. המדריך סוקר כלים נפוצים בתחום, מבחין בין שיטות הערכה דטרמיניסטיות לבין שימוש ב-LLM כשופט, ומציג כיצד לבצע בדיקות והשוואות מול קו בסיס ישירות בתוך פלטפורמת n8n.

קרא עוד
אופטימיזציית עלויות וזמני תגובה עם Prompt Caching ב-Bedrock
מדריך
3 דקות
מ־AWS Machine Learning

אופטימיזציית עלויות וזמני תגובה עם Prompt Caching ב-Bedrock

בפוסט של ארכיטקט הפתרונות דניאל אביב מ-AWS, מוסבר כיצד מנגנון ה-Prompt Caching ב-Amazon Bedrock מפחית עד 90% מעלויות טוקני הקלט על פגיעות במטמון ומקצר את זמן התגובה לטוקן הראשון (TTFT). המאמר סוקר שישה תרחישי יישום באמצעות ה-Converse API: שמירת מסמכים, שמירת פרומפט מערכת, שמירת הגדרות כלים לסוכנים, שילוב זמני חיים שונים (Mixed TTL), בידוד דיירים במערכות מרובות משתמשים באמצעות תחילית SHA-256, ואינטגרציה עם ספריית LangChain. מודלי Anthropic Claude Sonnet 4.5 ו-4.6 דורשים סף מינימלי של 1,024 טוקנים להפעלת המטמון.

קרא עוד