Voice AI Design Principles and the Voice Quality Framework
Analysis

Voice AI Design Principles and the Voice Quality Framework

Principles for human-adapted voice agents, 15 heuristics, and a three-layer implementation model

4 min read
Based on original reporting bySalesforce BlogTranslated and summarized by our AI-assisted news systemHow we work

Executive summary

Key Takeaways

  • Voice interface design is based on real-time conversation dynamics such as pacing, silence, and interruptions.

  • The Voice Quality Framework defines 3 failure tiers and 15 heuristics to evaluate agent behavior.

  • Design implementation in Agentforce comprises prompt instructions, deterministic logic, and voice channel settings.

Voice AI Design Principles and the Voice Quality Framework

  • Voice interface design is based on real-time conversation dynamics such as pacing, silence, and interruptions.
  • The Voice Quality Framework defines 3 failure tiers and 15 heuristics to evaluate agent behavior.
  • Design implementation in Agentforce comprises prompt instructions, deterministic logic, and voice channel settings.

In an article published on the Salesforce blog on Voice AI design, it is explained that designing voice interfaces is a distinct discipline shaped by real-time conversation dynamics such as pacing, silence, and interruptions. Voice-only interfaces do not function like graphical user interfaces (GUIs), where people can scan a screen, reread an answer, or skip ahead at their own pace. Voice interactions are inherently ephemeral, requiring active listening and holding information in memory, which can create cognitive load and damage trust if the conversation is not designed around how humans communicate. Unlike familiar and predictable interactive voice response (IVR) systems, voice AI systems generate responses based on each caller's situation, requiring an understanding of caller needs, detection of emotional state and context, and an appropriate response.

Differences Between GUI and VUI and Learning from Real Conversations

In a graphical user interface, users control the pace of progress, whereas a voice user interface (VUI) unfolds in real time through spoken conversation and communication patterns. A customer might not explicitly say they are stressed, but other signals indicate this:

  • Pauses and hesitations during speech.
  • Interruptions and repetitions of spoken statements.
  • Changes in speech pacing and word choice.
  • Changes in tone and emotional language.

At the same time, AI agent behaviors can create friction. A three-second pause by the agent can feel much longer when there is no indication of what is happening in the system; repeating a question can create the feeling that a task is stuck; and an unnecessary confirmation can add seconds to an extended interaction.

Analyzing and evaluating call recordings and annotated transcripts between human agents and callers serve as a central pillar for understanding how people communicate. Customers tend to report symptoms rather than root causes, such as describing the agent as "robotic," "slow," or "frustrating," or noting uncertainty about whether the agent is still listening. Technical success is not necessarily equivalent to user success.

The article presents two examples from insurance claims intake illustrating the difference between rigid design and design tailored to real conversation:

  1. Rigid slot filling versus adapted response: When the agent asks "Is this loss related to you, the caller, or a third-party claimant?", and the caller answers "Both", a rigid agent not designed for this response will repeatedly ask to choose between the options until the caller gives up and selects an inaccurate answer. In contrast, an adapted agent recognizes that the response refers to both parties, confirms that the claim will include both, and continues collecting details.
  2. Lack of empathy versus improved empathy: When a caller reports an accident and loss of life, a brief response such as "I'm sorry to hear that, do you have a policy number?" does not provide empathy. An adapted response acknowledges the difficulty, does not rush the caller, and offers multiple options for moving forward: immediately connecting to a human representative, starting the process with the agent, or receiving a callback at a later time.

The Voice Quality Framework: Three Tiers and 15 Heuristics

The Voice Quality Framework was developed based on conversation analysis and IVR design, dividing design failures and rules into three tiers:

  • Tier 1 (Foundational Layer): Safety and accuracy questions—whether the agent is safe and accurate, whether users can be understood, and whether the agent can be understood. If the agent fails at this tier, it is recommended not to launch it.
  • Tier 2 (Functionality and Invisible Churn Risks): Checking whether users trust that tasks will be completed efficiently, or whether they get lost and turn to human escalation when the conversation departs from the happy path.
  • Tier 3 (Customer Experience): Examining whether the agent is easy to use, natural, consistent, and adaptable—factors that distinguish an agent that merely functions from one users want to continue using.

Within this framework, 15 heuristics (design principles) are defined for evaluating conversational experience:

  1. Truthfulness: Avoiding hallucinations, contradictions, and factual errors.
  2. Trust: Adhering to defined permissions, protecting private information, and honestly admitting an inability to help.
  3. Recovery-Awareness: Identifying when the conversation goes off track and changing tactics instead of repeating the same things.
  4. Intelligibility: Good audio quality, natural speech pacing, and correct pronunciation.
  5. Effectiveness: Actually executing requests, including navigation requests and exits to a human representative.
  6. Responsiveness: Avoiding awkward silence and letting the user know the system is processing information.
  7. Confirmation Behavior: Appropriate calibration of confirmations for critical actions without confirming every minor detail.
  8. Memory and Context: Remembering provided information or data existing in the system without needing to ask for it again.
  9. Adaptability: Smoothly adjusting when the user corrects the agent or changes their mind, without being argumentative.
  10. Decisiveness: Moving the conversation forward with clear answers and next steps.
  11. Interruption Tolerance: Stopping and listening when the user interrupts the agent, and resuming smoothly afterward.
  12. Conversation Flow: Natural speech without unnecessary verbiage, and providing a plain explanation during a malfunction without feigning sympathy or assigning blame.
  13. Consistency: Maintaining uniform vocabulary, tone, and personality throughout the conversation, avoiding a generic, clichéd LLM voice.
  14. Approachability: Understanding everyday speech, slang, and accents, and adapting to language switching by the user.
  15. Thoroughness: Completing the task during the call, confirming execution, and offering relevant help before ending.

This framework bridges design, product, and engineering, enabling automated conversation evaluation processes at scale, identifying patterns, and mapping failures to specific conversational turns.

Applying Voice Design in Agentforce: The Three Layers

According to a forthcoming paper by Philipp Hänggi, design implementation in Agentforce is divided into three layers:

  1. Prose prompt instructions: Controlling conversational behavior (such as turn shape, end focus, and repair phrasing) by writing defined turn rules instead of vague tone instructions.
  2. Deterministic logic & variables: Ensuring a fixed sequence of actions, such as variables for slot filling, retry counters, mandatory summary confirmations, and automated escalation logic (for example, escalation after three failed repair attempts).
  3. Voice channel configuration: Managing vocal aspects and timing, including text-to-speech (TTS) voice selection, custom pronunciations, filler turns, acoustic barge-in, and telephony routing.

This hybrid approach combines the linguistic fluency of a language model with deterministic logic that ensures reliability and an accurate sequence of actions. Voice agents built on this design address real conversation characteristics including hesitations, emotions, interruptions, and unexpected answers.

Questions & Answers

FAQ

This article was produced by our AI-assisted system through translation, summarization, and automated quality controls based on original reporting by Salesforce Blog. Read about our editorial process. Link to the original source.

Get useful AI updates by email

A concise digest from our news desk.

More from Salesforce Blog

All articles from Salesforce Blog
סיילספורס מציגה את סוכן התזמון של Agentforce לשירות שטח
מוצר חדש
4 דקות
מ־Salesforce Blog

סיילספורס מציגה את סוכן התזמון של Agentforce לשירות שטח

סיילספורס הציגה את Scheduling Agent במסגרת Agentforce Field Service, סוכן בינה מלאכותית הפועל 24/7 לתיאום, שינוי וביטול פגישות שירות שטח. הסוכן מחובר לנתוני הלקוחות, ללוחות הזמנים של הטכנאים ולמנוע האופטימיזציה של הארגון, ופועל בערוצי תקשורת מגוונים בהם וואטסאפ, דוא"ל, SMS, iMessage ושיחות קוליות. המערכת מבוססת על Agent Script לקבלת החלטות דטרמיניסטית ואכיפת כללים עסקיים ללא ניחושים של מודלי שפה, ומספקת מענה לפניות לקוחות, סדרנים, טכנאים וטריגרים מנכסים. המערכת תוצג בכנס Dreamforce וב-Salesforce+.

קרא עוד
שלושה סימנים לכך שסוכן AI יחיד אינו מספיק בארגון
מוצר חדש
4 דקות
מ־Salesforce Blog

שלושה סימנים לכך שסוכן AI יחיד אינו מספיק בארגון

לפי פרסום של צוות Agentforce מבית Salesforce, תזמור מרובה-סוכנים (Multi-agent orchestration) זמין כעת באופן כללי. הפרסום מציג שלושה סימנים לכך שסוכן AI בודד הגיע לקצה גבול היכולת שלו: הסוכן מנסה לבצע משימות רבות מדי וסובל מהתנגשות כוונות (Intent collision), הנתונים הנדרשים נמצאים מחוץ לגבולות ה-Salesforce org, או שצוותים שונים נדרשים לנהל חלקים שונים של הסוכן. המאמר מציג ארבעה נתיבי ארכיטקטורה ושאלות מוכנות לבחינת המעבר, ומציין כי בבדיקות פנימיות בעיות הסקה מתרחשות לרוב מעבר לשבעה תת-סוכנים.

קרא עוד
15 דרכים לשימוש בסוכני AI לניהול רשתות חברתיות לפי Salesforce
מדריך
4 דקות
מ־Salesforce Blog

15 דרכים לשימוש בסוכני AI לניהול רשתות חברתיות לפי Salesforce

מדריך של חברת Salesforce מפרט 15 דרכים שבהן סוכני בינה מלאכותית לרשתות חברתיות מסייעים לעסקים קטנים ובינוניים. הכלים האוטונומיים מאפשרים יצירת תוכן בקול המותג, תזמון פוסטים בזמנים מותאמים אישית, מענה אוטומטי לשאלות נפוצות 24/7, ניתוב פניות מורכבות לנציגים אנושיים, ניטור אזכורים וסנטימנט, וחיבור מעורבות ישירות למערכות ה-CRM לצורך יצירת לידים. בנוסף מובאת דוגמת חברת reMarkable, שטיפלה ביותר מ-18,000 שיחות שירות באמצעות סוכני AI.

קרא עוד
מדריך Salesforce: כיצד להרחיב צוות מכירות ברבעון אחד
מדריך
4 דקות
מ־Salesforce Blog

מדריך Salesforce: כיצד להרחיב צוות מכירות ברבעון אחד

מדריך של Salesforce מציג תוכנית רבעונית להרחבת צוות מכירות ללא שחיקה, באמצעות הגדרת תהליך מכירות ברור, אוטומציה של מעקבים ושימוש בבינה מלאכותית. לפי המדריך, 76% מעסקי ה-SMB פועלים מתצוגת CRM משותפת, ו-88% כבר משתמשים ב-AI לניהול לידים ותובנות עסקה. המדריך מפרט צעדים חודשיים הכוללים הגדרת יעדים, קליטת עובדים מבוססת מערכת והדרכה שוטפת.

קרא עוד

More articles you might like

All articles
תזמור תהליכים: מודלי ביצוע, אתגרי ייצור ותזמור מול כוריאוגרפיה
ניתוח
4 דקות
מ־n8n

תזמור תהליכים: מודלי ביצוע, אתגרי ייצור ותזמור מול כוריאוגרפיה

בפוסט שפורסם בבלוג של n8n, נסקרים מודלי הביצוע המרכזיים בתזמור תהליכים (Process Orchestration): דטרמיניסטי, דינמי וסוכני (Agentic). המאמר מנתח את הפשרות בין יכולת ניבוי, הסתגלות ואוטונומיה, מציג את המאפיינים של תהליכים המתאימים לתזמור מרכזי, וסוקר אתגרי ייצור נפוצים כגון צווארי בקבוק, השחתת מצב, נדידת סכמות וניפוי שגיאות במערכות מבוזרות. כמו כן, מוסברים ההבדלים בין תזמור לכוריאוגרפיה ואוטומציית משימות בודדות.

קרא עוד
הרחבת השימוש בסוכני פיתוח ב-Salesforce ל-15,000 מהנדסים
ניתוח
4 דקות
מ־Salesforce News

הרחבת השימוש בסוכני פיתוח ב-Salesforce ל-15,000 מהנדסים

בפוסט הנדסי שפורסם מטעם Salesforce מפורט כיצד הורחב השימוש בסוכני פיתוח בינה מלאכותית ל-15,000 מהנדסים בחברה. לפי הדיווח, המהלך לווה בעלייה של 90.5% בהשלמת משימות למפתח ועלייה של 200.3% במדד הפרודוקטיביות Effective Output שפותח עם אוניברסיטת סטנפורד. התהליך כלל פיילוט של 30 ימים, הגדרת מודל בשלות בן תשעה שלבים, ומשמעת ניהול הקשר וטוקנים שהביאה לחסכון כספי ולשיפור איכות הקוד.

קרא עוד
מילון מונחי AI מקיף: המושגים המרכזיים שצריך להכיר
ניתוח
4 דקות
מ־TechCrunch

מילון מונחי AI מקיף: המושגים המרכזיים שצריך להכיר

במדריך מושגים מקיף שפורסם ב-TechCrunch, מציגים כתבי האתר מילון מונחים מרכזי בעולם הבינה המלאכותית. המילון כולל הגדרות ברורות למונחים כמו AGI, סוכני AI, סוכני תכנות, ארכיטקטורת תערובת מומחים (MoE), פרוטוקול MCP לחיבור מקורות מידע, וטכניקת הישנות עמומה (Opaque recurrence) המייעלת עיבוד אך מעלה שאלות בטיחות ומעקב. בנוסף מפורטים תהליכי אימון, זיקוק, הסקה, מטמון זיכרון והשפעות המחסור בחומרת זיכרון המכונה RAMageddon.

קרא עוד
אבטחת תהליכי עבודה: בקרות לענפים מוסדרים לפי n8n
ניתוח
4 דקות
מ־n8n

אבטחת תהליכי עבודה: בקרות לענפים מוסדרים לפי n8n

בפוסט שפרסמה חברת n8n נסקרות שש בקרות אבטחה מרכזיות לתהליכי עבודה אוטומטיים בענפים מוסדרים כגון בריאות ופיננסים: בקרת גישה מבוססת תפקידים (RBAC), ניהול סודות, רישום יומני ביקורת, תושבות נתונים, בידוד סביבות ומערכות ניטור. המאמר מסביר כיצד כלי אוטומציה סגורים במודל SaaS עלולים להקשות על ביצוע הערכות אבטחה עצמאיות בשל היעדר שקיפות בקוד, ומנגד כיצד פלטפורמות עם קוד מקור זמין בהתקנה עצמית מאפשרות שליטה בהגדרות ובהרצה לצורך עמידה בתקני רגולציה כמו GDPR, HIPAA ו-SOC 2.

קרא עוד