Product launch

Pay-Per-Inference for AI Agents: AWS AgentCore Payments

Amazon Bedrock AgentCore payments lets AI agents pay for model inference over x402 with infrastructure spending limits

4 min read
עב
Pay-Per-Inference for AI Agents: AWS AgentCore Payments
Based on original reporting byAWS Machine Learning ↗Translated and summarized by our AI-assisted news systemHow we work

Executive summary

What to know

  1. The AgentCore payments capability in Amazon Bedrock is generally available (GA), allowing agents to pay for services on demand.

  2. Incarna integrated the service to enable agents to pay the BlockRun router, which serves over 90 models from more than 15 providers over x402.

  3. Spending limits are enforced at the infrastructure level outside the model, so prompt manipulation cannot bypass defined budgets.

  4. Incarna's integration took 3 days with roughly 200 lines of code, processing over 1,000 payments between $0.001 and $0.05 in beta.

In a post published by Peter Jiang from AWS, the author presented the use of the AgentCore payments capability of Amazon Bedrock, which allows artificial intelligence agents to pay directly on demand for services such as model inference. The post describes how Incarna integrated the service to enable its agents to pay the BlockRun inference router for each call individually over the x402 protocol, with spending limits enforced at the infrastructure layer rather than by the model itself.

The Challenge: Paying for Model Inference per Single Request

According to the post's authors, when an AI agent runs, it often needs to purchase resources to finish a task: a model inference, an API response, access to web content, or a call to another agent. These purchases are frequent and low in value, sometimes a fraction of a cent per operation, and they take place inside the agent's runtime loop without an available human to approve them.

Credit card networks were not built for sub-cent payments. The post explains that building an independent payments infrastructure requires solving several complex challenges at once: determining where the money is held, how each payment is signed, supporting emerging payment protocols such as x402, and preventing an autonomous agent from exceeding the spending framework defined for it.

Capabilities Provided by the AgentCore Payments Service

Amazon Bedrock AgentCore is defined in the post as a platform to build, connect, and optimize agents at scale, with any framework or model. The managed AgentCore payments capability allows developers to add payments to agents in a few lines of code. The system handles the payment protocol, connects to a wallet, signs the transaction, and enforces spending limits, so developers use a single managed service instead of assembling the parts on their own.

The service includes several key features:

  1. Managed wallets: Incarna provisions a wallet for each agent using the Coinbase CDP connector. The customer owns the wallet and grants Incarna delegated authorization to use it.
  2. Native protocol handling: When a paid endpoint responds with HTTP 402 ("Payment Required"), the agent uses AgentCore payments to execute the payment over x402, signs the transaction using the configured wallet, and returns cryptographic proof to the seller. The service works with x402-compatible endpoints, including Amazon Bedrock inference endpoints.
  3. Spending governance at the infrastructure layer: Spending limits are enforced outside the model, so the agent cannot exceed them even if its prompt is manipulated.
  4. Auditable settlement: Payments settle in a stablecoin. Incarna uses USDC on the Base network, and each transaction is verifiable on the blockchain.

Prerequisites and Component Setup

According to the post, teams can set up the components through a guided conversation with the AgentCore payments skill in the Agent Toolkit for AWS (in Claude Code, Kiro, or Codex), or create each component independently using the AgentCore CLI, the AgentCore SDK, or the AWS SDK.

The process begins by storing Coinbase CDP or Stripe Privy credentials as a payment credential provider, which keeps secrets in AWS Secrets Manager instead of application code. Next, developers create a Payment Manager and connector to coordinate payments against those credentials, and establish a default spending limit. In the next step, a payment instrument is created, which is the embedded wallet from which the agent pays. The end user funds the wallet and grants signing permission via a redirect URL, and in a test environment it can be funded with testnet USDC.

The Flow of Buying and Selling a Single Inference

In the described integration, BlockRun serves as the seller. BlockRun is a pay-as-you-go inference router serving more than 90 models from more than 15 providers over x402, with each call quoted and settled independently.

The flow for a single inference includes several stages:

  • The agent needs a model call and interfaces with BlockRun, which handles provider selection and delivery without requiring a separate subscription for each provider.
  • BlockRun responds with a PaymentRequired challenge and a price for that specific call.
  • A payment session is opened and a call is made to ProcessPayment. The system checks the quote against the spending limits defined for that session and signs the authorization from the agent's wallet address.
  • The seller verifies the payment signature.
  • BlockRun delivers the inference and records the charge, a small amount per call.

The service supports two x402 payment mechanisms: the exact mechanism, typically used when the price is known in advance, and the upto mechanism, suitable for dynamically priced resources where the agent authorizes a ceiling amount and the inference provider settles actual usage up to that ceiling.

Budget Control and Integration Results

The component enabling spending control is the payment session, which establishes a maximum ceiling the agent can spend. AgentCore payments enforces this ceiling at the infrastructure layer, so the agent's code and prompt cannot alter it. Each session also carries an expiration time, and Incarna configured the sessions according to a daily budget.

Incarna deployed the process into production on the Base network. According to the post's figures, the Incarna team completed the full integration in three days (one day for development and two days for testing) with approximately 200 lines of application code, compared to an original estimate of two to three months. Throughout the beta phase, agents processed more than 1,000 payments ranging from $0.001 to $0.05 per call, with each payment settled individually on-chain.

Vicky Fu, founder of BlockRun, noted that the company's open-source router gives developers control over models alongside benchmark-driven routing improvements, and that AgentCore payments and x402 provide the necessary spending controls. Justin Zhou, founder of Incarna, stated that the service covered everything needed for an agent to pay over x402: a customer-owned wallet, a funding and revocation flow, a platform-enforced spending limit, and signing supporting both x402 versions.

About BlockRun and Incarna

BlockRun serves model inference over the x402 protocol through a single metered endpoint, with each call priced, authorized, paid, and settled on the Base network in USDC without per-provider subscriptions or billing relationships. Incarna, built by SpreadX and operating on Base mainnet, provides AI agents with a consistent identity including its own wallet, email address, social accounts, and an action history associated with the same ID, so the payer in the transaction is the agent's identity rather than a shared platform key. The Amazon Bedrock AgentCore payments feature is now generally available (GA).

Was this useful for your business?

Questions & Answers

FAQ

This article was produced by our AI-assisted system through translation, summarization, and automated quality controls based on original reporting by AWS Machine Learning. Read about our editorial process. Link to the original source.

Get useful AI updates by email

A concise digest from our news desk.

More from AWS Machine Learning

All articles from AWS Machine Learning
Claude Haiku 5.5 זמין כעת ב-Amazon Bedrock וב-AWS
מוצר חדש
4 דקות
מ־AWS Machine Learning

Claude Haiku 5.5 זמין כעת ב-Amazon Bedrock וב-AWS

חברת AWS הודיעה על זמינות המודל Claude Haiku 5.5 ב-Amazon Bedrock וב-Claude Platform on AWS. לפי Anthropic, המודל הוא המהיר והיעיל ביותר במשפחת Claude 5.5, נבנה עבור תתי-סוכנים ומשימות בנפח גבוה, ועולה כ-75 אחוז פחות מ-Claude Haiku 4.5 במרבית המשימות. השירות ב-Bedrock שומר על תושבות נתונים אזורית ומשתלב עם מנגנוני IAM, CloudTrail, CloudWatch ו-Guardrails. המודל כולל בקרות מאמץ לכוונון עלות מול תבונה, ומיועד להשתלב לצד Claude Opus 5.5 שמבצע את התכנון והשיפוט המורכב. המודל זמין בפרופילי הסקה גלובליים ואזוריים וב-AWS GovCloud.

קרא עוד
בניית סוכן נסיעות קולי עם Bedrock AgentCore ו-Nova Sonic
מדריך
4 דקות
מ־AWS Machine Learning

בניית סוכן נסיעות קולי עם Bedrock AgentCore ו-Nova Sonic

בפוסט טכני של AWS הציגו מומחי החברה ארכיטקטורה לפריסת סוכן נסיעות קולי לחברות תעופה. הפתרון משלב את Amazon Bedrock AgentCore לניהול והרצת סוכנים, מודל הדיבור Amazon Nova 2.5 Sonic לקול בזמן אמת, ו-Amazon Bedrock Knowledge Bases למענה על שאלות מדיניות מתוך מסמכים. המערכת מקשרת בין הסוכן לשירותי ה-backend באמצעות פרוטוקול Model Context Protocol (MCP) ומאפשרת טיפול בהזמנות, החלפת מושבים והסלמה לנציג אנושי.

קרא עוד
ארכיטקטורת סוכני AI לניתוח חוזים עם Amazon Bedrock ו-Quick
ניתוח
4 דקות
מ־AWS Machine Learning

ארכיטקטורת סוכני AI לניתוח חוזים עם Amazon Bedrock ו-Quick

בפוסט שפורסם בבלוג של AWS הציגו מהנדסי החברה ארכיטקטורה לפלטפורמת ניתוח חוזים, המשלבת סוכני בינה מלאכותית מבוססי Amazon Bedrock AgentCore וכלי תשאול וניתוח ב-Amazon Quick. הפתרון מתמודד עם מגבלות כלי RAG בעת ביצוע חישובי אגרגציה על מאות מסמכים, באמצעות חילוץ שדות מפתח למסד נתונים מובנה ב-Amazon Aurora PostgreSQL. המערכת משתמשת בסוכן חילוץ מבוסס Claude Sonnet ובסוכן אימות מבוסס Claude Haiku, לצד Amazon Textract כגורם מכריע לזיהוי חתימות בעזרת ראייה ממוחשבת. הגישה מאפשרת לבצע הן שאילתות רוחביות והן איתור מקטעים מתוך מסמך יחיד בממשק מאוחד.

קרא עוד
כיצד HEMA בנתה שכבת ידע ארגונית עם Bedrock ו-MCP
ניתוח
4 דקות
מ־AWS Machine Learning

כיצד HEMA בנתה שכבת ידע ארגונית עם Bedrock ו-MCP

רשת הקמעונאות ההולנדית HEMA בנתה שכבת ידע פנימית המבוססת על Amazon Bedrock AgentCore ו-Model Context Protocol (MCP) במטרה לאחד מידע מבוזר ולמנוע מעבר ידני בין פורטלים ומערכות ויקי שונות. העוזר הפנימי HAL, שפותח תחילה ככלי עצמאי מבוסס Next.js ו-Strands, הורחב לשימוש ישיר מתוך כלי העבודה של המהנדסים (כגון Kiro ו-Claude) באמצעות שער Entra MCP ייעודי ופרוקסי אימות. המערכת משרתת כיום מפתחים, מנהלי מוצר ומנתחי מערכות, כאשר השלב הבא מתוכנן להרחיב את יכולות העוזר ממענה לשאלות לביצוע פעולות תפעוליות ישירות מתוך ממשקי השיחה.

קרא עוד

More articles you might like

All articles
Claude Haiku 5.5 זמין כעת ב-Amazon Bedrock וב-AWS
מוצר חדש
4 דקות
מ־AWS Machine Learning

Claude Haiku 5.5 זמין כעת ב-Amazon Bedrock וב-AWS

חברת AWS הודיעה על זמינות המודל Claude Haiku 5.5 ב-Amazon Bedrock וב-Claude Platform on AWS. לפי Anthropic, המודל הוא המהיר והיעיל ביותר במשפחת Claude 5.5, נבנה עבור תתי-סוכנים ומשימות בנפח גבוה, ועולה כ-75 אחוז פחות מ-Claude Haiku 4.5 במרבית המשימות. השירות ב-Bedrock שומר על תושבות נתונים אזורית ומשתלב עם מנגנוני IAM, CloudTrail, CloudWatch ו-Guardrails. המודל כולל בקרות מאמץ לכוונון עלות מול תבונה, ומיועד להשתלב לצד Claude Opus 5.5 שמבצע את התכנון והשיפוט המורכב. המודל זמין בפרופילי הסקה גלובליים ואזוריים וב-AWS GovCloud.

קרא עוד
גוגל משיקה שרת MCP מרוחק עבור Google Cloud CLI בתצוגה מקדימה
מוצר חדש
4 דקות
מ־Google Cloud AI

גוגל משיקה שרת MCP מרוחק עבור Google Cloud CLI בתצוגה מקדימה

גוגל הכריזה על השקת שרת MCP מרוחק עבור Google Cloud CLI בגרסת Preview. השרת החדש מאפשר לסוכני AI גישה ישירה להרצת פקודות gcloud ו-bq (BigQuery) בסנדבוקס ביצוע מבודד ומאובטח על גבי תשתית הענן של גוגל, ללא צורך בהתקנה או בתחזוקה של סביבות ריצה מקומיות. השרת כולל מנגנוני אבטחה ארגוניים המבוססים על אימות IAM ו-OAuth 2.0, מניעת אישורים סביבתיים (zero ambient credentials), הגנה מפני הזרקת פרומפטים באמצעות Model Armor, ותיעוד פעילות ביומני ביקורת בענן. השרת חושף שני כלים מרכזיים, run_gcloud_command ו-run_bq_command, ומאפשר אוטומציה של ניהול תשתיות ושאילתות BigQuery ללא עלות נוספת על שרת ה-MCP עצמו.

קרא עוד
הכרזת n8n Agents: שילוב סוכני AI עצמאיים לצד תהליכי עבודה
מוצר חדש
4 דקות
מ־n8n

הכרזת n8n Agents: שילוב סוכני AI עצמאיים לצד תהליכי עבודה

פלטפורמת n8n הכריזה על השקת Agents (סוכנים), המאפשרים למשתמשים להגדיר מטרות בשפה חופשית ולהשאיר לסוכן לקבוע את שלבי הביצוע בעזרת מודלים, כלים ותהליכי עבודה קיימים. הסוכנים יכולים לפעול מתוך Slack, Telegram, Discord, לפי תזמון מוגדר או מתוך תהליכי עבודה באמצעות הצומת החדש Message an Agent. כל סוכן כולל ניהול זיכרון, הפעלות, כלים, מיומנויות ומנגנוני אישור אנושי לפעולות רגישות. התכונה זמינה כעת ב-Preview למשתמשי n8n Cloud ובהתקנה עצמאית.

קרא עוד
מודל Jev של TypeSafe AI: קבלת החלטות מהירה לאוטומציה ללא הזיות
מוצר חדש
4 דקות
מ־TechCrunch

מודל Jev של TypeSafe AI: קבלת החלטות מהירה לאוטומציה ללא הזיות

חברת TypeSafe AI, שהוקמה על ידי חוקר OpenAI לשעבר דיוגו אלמיידה, השיקה את Jev — מודל טרנספורמר חדש שאינו מפיק טקסט אלא הסתברויות והחלטות מכוילות. המודל מאפשר קבלת החלטות מהירה וזולה לאוטומציית תוכנה ללא סכנת הזיות, הודות להגדרת הפלטים מראש על ידי המשתמש ואימונו הבלעדי על נתונים סינתטיים. מפתחים מדווחים על שיפורי מהירות משמעותיים ועלויות נמוכות בהשוואה למודלי שפה מסורתיים.

קרא עוד