Salesforce Scales Agentic Coding to 15,000 Engineers
Analysis

Salesforce Scales Agentic Coding to 15,000 Engineers

An overview of the rollout to 15,000 engineers, productivity data, and token management principles at Salesforce.

4 min read
Based on original reporting bySalesforce NewsTranslated and summarized by our AI-assisted news systemHow we work

Executive summary

Key Takeaways

  • Work items completed per developer increased by 90.5% year-over-year in July, and pull requests merged per developer rose by 88.1%.

  • The Effective Output score developed in collaboration with Stanford University recorded a 200.3% increase.

  • The 30-day pilot ran in March across roughly 44 teams from ten product clouds and over 200 engineers.

  • Auto-compaction where context windows automatically compact at 200,000 tokens reduced engineering spend by 24.8% over three weeks.

Salesforce Scales Agentic Coding to 15,000 Engineers

  • Work items completed per developer increased by 90.5% year-over-year in July, and pull requests merged...
  • The Effective Output score developed in collaboration with Stanford University recorded a 200.3% increase.
  • The 30-day pilot ran in March across roughly 44 teams from ten product clouds and...
  • Auto-compaction where context windows automatically compact at 200,000 tokens reduced engineering spend by 24.8% over...

In a post published on behalf of Salesforce, the company describes the process of scaling and adopting AI agent-based coding (agentic coding) across 15,000 engineers in the organization. According to the report, the organization runs a $40 billion business operation and serves hundreds of thousands of customers, with the rollout of the tools tested across active production systems where every deployment directly affects revenue and customer trust.

The figures presented in the post show that in July, a 90.5% year-over-year increase was recorded in the number of work items completed per developer. In addition, pull requests merged per developer increased by 88.1%. The Effective Output metric—a machine learning–based productivity score developed in collaboration with Stanford University—rose by 200.3%. This score is assigned to every commit after a machine learning model automatically reviews the code similarly to a panel of senior engineers, evaluating quality, complexity, and effort. However, the post emphasizes that the operating model and organizational culture behind the numbers represent the primary learning focus.

An Organizational Culture of Building and Sharing

According to the description in the post, even before a single tool was rolled out, the organization maintained a culture where the default response to a lack of technological capability was to build it internally and then share it. When engineers needed to manage fleets of agents—including tracking progress, maintaining the flow of autonomous sessions, and catching issues without constant monitoring—no external vendor offered an exact solution. A group of engineers built an orchestration layer and distributed it across the entire organization without top-down direction or funding under a dedicated program.

The combination of bottom-up innovation from the ground and managerial top-down curation and standardization enabled the rapid creation of tools and skills. For this reason, teams did not standardize on a single tool, and options like AI Expert Suite and Dev Bar exist because engineers needed the flexibility to move between tools and models according to work needs.

The Pilot Framework and Organization-Wide Expansion

The rollout began with a targeted 30-day pilot in March, rather than an immediate deployment for all 15,000 engineers. The pilot included roughly 44 teams across ten different product clouds and over 200 engineers, selected to represent the full complexity of the company's systems: greenfield projects alongside interconnected legacy systems, high-velocity development teams alongside maintenance teams, and early adopters alongside skeptics.

The objective of the pilot was to test whether tools like Claude Code could be trusted in mission-critical systems in an enterprise environment, as well as to identify friction in workflows and human habits. Following the pilot, joint training and enablement activities were conducted with Anthropic. Leadership presented the use as an ongoing operational expectation, and at the same time, a structure of "champions" was established to transfer knowledge laterally across teams.

At this stage, a goal was set for all engineering teams: to achieve exponential productivity within 90 days. The goal was intended to compel a rethinking of how software is built, and served as a diagnostic tool that identified bottlenecks in cross-team dependencies and approval processes.

The Agent Coding Maturity Model

To establish a shared language for the depth of adoption, a nine-stage maturity model (Agent Coding Maturity Curve) was developed, ranging from basic code generation to trusted autonomous operation. The organizational objective was defined as moving the entire engineering population to stage 6 and beyond.

According to the post, binary adoption metrics only show who has access to the tool, but do not reflect whether the usage is for writing a single function or orchestrating an autonomous agent across a multi-service migration. The maturity model also changed the nature of managerial conversations, which focused on where the engineer is positioned on the model and the actions required to advance to the next stage, such as improving context quality, task decomposition, and delegating full ownership of workflows to agents.

Token Discipline as an Engineering Methodology

At Salesforce, token optimization was treated as an engineering discipline. The core insight was that over-contexting degrades quality and performance, not just budget: it creates unnecessary financial costs, increases latency, and dilutes the model's accuracy and focus.

The organization defined six levers for context management:

  1. Context hygiene: Using commands such as clear, compact, or branch depending on the relevance of the session history, and treating context like working memory rather than a transcript.
  2. Skill libraries: Building global, local, and team-specific skills to avoid redefining instructions and context from scratch every time.
  3. Model selection: Setting cost-efficient models as the default for routine tasks, and reserving the most capable models for complex challenges.
  4. Effort calibration: Reserving high-effort modes for ambiguity, complexity, and critical decisions.
  5. Tool selection: Using focused retrieval tools like CodeSearch MCP instead of feeding extensive amounts of data that create noise.
  6. Task decomposition: Using subagents and execution plans to split complex work into parallel threads with defined scoped context.

Auto-compaction alone—where context windows automatically compact at 200,000 tokens—drove a 24.8% reduction in total engineering spend, sustained over three weeks. Smart model defaults saved $864,000 in the first week. In selective cases, optimized prompting modes reduced tokens by 70% or more with no measurable quality drop.

In addition, it was noted that model diversification became a first-class engineering concern, based on building routing infrastructure, skills, and context management that allows models to be swapped as the market evolves.

Core Principles for Scaling

The post concludes with a series of continuous principles for the process:

  • Build the foundation before scaling: An agent-ready codebase, CLAUDE.md files, accessible institutional knowledge, observable workflows, and an initial reusable skill library.
  • Prove capability in the organization's real systems through a complex pilot.
  • Set an exponential target to force rethinking of processes and identify gaps.
  • Create a shared language for depth of adoption using a maturity model.
  • Treat cost discipline as an engineering discipline.
  • Continuously invest in an organizational culture that encourages internal building and knowledge sharing.

Questions & Answers

FAQ

This article was produced by our AI-assisted system through translation, summarization, and automated quality controls based on original reporting by Salesforce News. Read about our editorial process. Link to the original source.

Get useful AI updates by email

A concise digest from our news desk.

More from Salesforce News

All articles from Salesforce News
העקרונות להטמעת סוכני בינה מלאכותית בשירות לקוחות לפי סיילספורס
ניתוח
4 דקות
מ־Salesforce News

העקרונות להטמעת סוכני בינה מלאכותית בשירות לקוחות לפי סיילספורס

במאמר שפורסם מטעם סיילספורס, נותחו הגורמים להצלחת הטמעת סוכני בינה מלאכותית בשירות לקוחות על בסיס נתוני תוכנית פרסי הלקוחות של החברה. הניתוח מציג שלושה עקרונות מרכזיים: התמקדות בבעיה תפעולית מוגדרת, בניית תשתית נתונים מוצקה ושיתוף העובדים בתהליך. המאמר מדגים עקרונות אלה באמצעות שלושה מקרים: מועדון הכדורגל טוטנהאם הוטספור שאיחד נתוני 4.6 מיליון אוהדים וקיצר את זמני המענה; רשת The Grout Guy שקיצרה את זמן הפקת הצעות המחיר מ-3–5 ימים ל-20 דקות; וחברת Sammons Financial Group שטיפלה ביותר מ-16,000 שיחות פוליסה באמצעות סוכן בינה מלאכותית ופיקוח אנושי.

קרא עוד

More articles you might like

All articles
מילון מונחי AI מקיף: המושגים המרכזיים שצריך להכיר
ניתוח
4 דקות
מ־TechCrunch

מילון מונחי AI מקיף: המושגים המרכזיים שצריך להכיר

במדריך מושגים מקיף שפורסם ב-TechCrunch, מציגים כתבי האתר מילון מונחים מרכזי בעולם הבינה המלאכותית. המילון כולל הגדרות ברורות למונחים כמו AGI, סוכני AI, סוכני תכנות, ארכיטקטורת תערובת מומחים (MoE), פרוטוקול MCP לחיבור מקורות מידע, וטכניקת הישנות עמומה (Opaque recurrence) המייעלת עיבוד אך מעלה שאלות בטיחות ומעקב. בנוסף מפורטים תהליכי אימון, זיקוק, הסקה, מטמון זיכרון והשפעות המחסור בחומרת זיכרון המכונה RAMageddon.

קרא עוד
אבטחת תהליכי עבודה: בקרות לענפים מוסדרים לפי n8n
ניתוח
4 דקות
מ־n8n

אבטחת תהליכי עבודה: בקרות לענפים מוסדרים לפי n8n

בפוסט שפרסמה חברת n8n נסקרות שש בקרות אבטחה מרכזיות לתהליכי עבודה אוטומטיים בענפים מוסדרים כגון בריאות ופיננסים: בקרת גישה מבוססת תפקידים (RBAC), ניהול סודות, רישום יומני ביקורת, תושבות נתונים, בידוד סביבות ומערכות ניטור. המאמר מסביר כיצד כלי אוטומציה סגורים במודל SaaS עלולים להקשות על ביצוע הערכות אבטחה עצמאיות בשל היעדר שקיפות בקוד, ומנגד כיצד פלטפורמות עם קוד מקור זמין בהתקנה עצמית מאפשרות שליטה בהגדרות ובהרצה לצורך עמידה בתקני רגולציה כמו GDPR, HIPAA ו-SOC 2.

קרא עוד
העקרונות להטמעת סוכני בינה מלאכותית בשירות לקוחות לפי סיילספורס
ניתוח
4 דקות
מ־Salesforce News

העקרונות להטמעת סוכני בינה מלאכותית בשירות לקוחות לפי סיילספורס

במאמר שפורסם מטעם סיילספורס, נותחו הגורמים להצלחת הטמעת סוכני בינה מלאכותית בשירות לקוחות על בסיס נתוני תוכנית פרסי הלקוחות של החברה. הניתוח מציג שלושה עקרונות מרכזיים: התמקדות בבעיה תפעולית מוגדרת, בניית תשתית נתונים מוצקה ושיתוף העובדים בתהליך. המאמר מדגים עקרונות אלה באמצעות שלושה מקרים: מועדון הכדורגל טוטנהאם הוטספור שאיחד נתוני 4.6 מיליון אוהדים וקיצר את זמני המענה; רשת The Grout Guy שקיצרה את זמן הפקת הצעות המחיר מ-3–5 ימים ל-20 דקות; וחברת Sammons Financial Group שטיפלה ביותר מ-16,000 שיחות פוליסה באמצעות סוכן בינה מלאכותית ופיקוח אנושי.

קרא עוד
חמש החלופות המובילות ל-Intercom לשנת 2026 לפי Zapier
ניתוח
4 דקות
מ־Zapier

חמש החלופות המובילות ל-Intercom לשנת 2026 לפי Zapier

סקירה שפורסמה בבלוג של Zapier מציגה חמש חלופות עיקריות לפלטפורמת Intercom לשנת 2026, על רקע המורכבות ומודל התמחור של Intercom המבוסס על תשלום לפי פתרון של סוכן AI. הסקירה מחלקת את החלופות לפי צרכים: Zendesk עבור תמיכה רב-ערוצית בקנה מידה רחב; HubSpot עבור מערכת אחודה המשלבת מכירות, שיווק ושירות סביב CRM משותף; Freshdesk עבור ניהול פניות ונגישות לסוכני AI במחיר התחלתי נמוך; Customer.io עבור אוטומציה של מסעות לקוח והודעות מחזור חיים; ו-LiveChat עבור התמקדות בצ'אט חי והודעות בזמן אמת. כל הכלים נבחנו לפי יכולות AI, ערוצי תמיכה, תקשורת יזומה, חיזוי עלויות ואינטגרציות.

קרא עוד