In an opinion column published on SiliconANGLE, M. Touheed, a growth specialist at Imagine Art, writes that most traditional enterprise systems were built around three foundational assumptions: jobs finish quickly, retrying one costs no money, and the same input always produces the same output. Agentic AI workloads break all three of these expectations, which is why pilots that demonstrate strong performance turn into operational problems once they run unattended without human supervision. According to Touheed, the difficulty rarely stems from the model itself; instead, the issue lies in the surrounding infrastructure and management practices that assume properties these workloads simply no longer possess.
Three Differences Between Agentic Work and Traditional Software
According to the article, agents operate under different rules than conventional software, with three significant differences standing out:
First, work takes minutes rather than milliseconds. An agentic task can run long enough to exceed timeout thresholds that no component in the technology stack has encountered before. Systems that previously appeared stable begin failing in ways that look mysterious, until someone checks the execution duration of the task.
Second, retries now cost money. Traditionally, retry logic was almost free, which led teams to retry generously. However, every attempt against a metered model consumes compute resources whether the result is usable or not. The combination of loose quality thresholds and automated retries creates a budget event. Because cloud and model application programming interface (API) usage is billed asynchronously on monthly cycles, these compounded retry costs accumulate quietly and become visible only when the invoice arrives weeks later.
Third, failures cannot be reproduced. Touheed notes that this represents the biggest shift. When an engineer investigates a defective result, standard practice is to rerun the process and watch it fail again. This approach does not work when agents are involved. Without a step-by-step trace of agent decisions, tool calls, and API runs, there is no data to investigate, and outcome reviews turn into mere speculation.
Why Pilot Environments Conceal Operational Problems
Touheed explains that experienced teams are caught off guard because the evaluation environment conceals all of these operational challenges. When work is performed interactively, humans function as error handlers. They read each result, notice issues, and try again. Costs remain visible because attempts are counted and fixes are applied manually, and there is no need for an audit record or a mechanism to precisely reproduce the failure.
In contrast, automated agent workflows run headlessly, carrying the potential for unrecorded failures to disrupt downstream systems. A workflow that behaved reliably when a person chatted with the AI and reviewed each response behaves entirely differently when a scheduler triggers it 400 times overnight with nobody watching. The error was always present, but it was invisible because a human operator absorbed and corrected it response by response. Therefore, before moving to automated production with agents, engineering leaders must identify every task the human performed manually and specify which automated check or system will take over that responsibility.
Three Risks Worth Attention
The column outlines three primary operational risks:
-
Invisible costs: Expenses are no longer a function of usage volume alone. Autonomous AI loops automatically retry failed tasks, regenerate responses, and call APIs repeatedly without human intervention or approval. A threshold set by a developer can influence the monthly bill more than a procurement negotiation. Touheed recommends asking for the cost per completed unit rather than the cost per API call, since the latter metric excludes discarded attempts.
-
Failures that report success: The most expensive defects are not crash errors, but results that appear structurally valid while being substantively wrong. Probabilistic agents cannot recognize their own logical errors, leading them to generate properly structured outputs containing bad data. Downstream automated systems accept and process them without triggering any alerts, and conventional monitoring misses these errors because technically nothing failed. AI output quality must be measured using predefined programmatic criteria (such as assertion checks, LLM-as-a-judge rules, or semantic benchmarks) rather than relying on standard uptime or error logs.
-
Incidents nobody can explain: If a team cannot state which model version, inputs, and settings produced a specific output, they cannot investigate the incident, nor can an auditor. This information is inexpensive to capture while the work is running, but nearly impossible to reconstruct later.
Five Questions to Ask Development Teams
According to Touheed, five questions address most of the scenarios described:
- What defines acceptable output, and is it written down before the work runs? Automation requires clear, reproducible acceptance criteria. If success criteria shift based on individual human opinion during review, software cannot validate the output automatically.
- How many attempts does a typical completed unit require, and is that capped? An unbounded retry loop against a vague standard is the most common source of overspending.
- What is recorded and documented for every run, and could the team produce that record on request six months from now? Job data should include the exact input payload, prompt and model version, timestamp, execution latency, retry counts, token cost, and the final output artifact.
- Who approves various classes of output, by name? This refers to a single individual rather than a general team, because when something goes wrong, that is the first question that will be asked.
- What happens when the provider updates the model? Model updates can subtly alter response formatting, accuracy, or reasoning logic. These downstream shifts can break automated pipelines, which is why model version changes must be tested in a staging environment before being deployed to production.
Consolidating the Workspace and the Skills Gap
Touheed points out that agentic work tends to be scattered because teams manage agentic tools using traditional software practices: prompts reside in one team's repository, artifacts are stored across different buckets, approvals take place in chat threads, and no one can produce an execution record without hours of investigation. The desired AI workspace is a single place where runs, inputs, outputs, quality results, and approvals are recorded together as an operational surface that enables execution management, re-executing inputs, log tracing, and workflow control in real time. Reporting and logging must be continuous across every execution run rather than periodic. Platforms serving production pipelines, such as ImagineArt, increasingly expose these capabilities, emphasizing that a polished product demo showcases best-case results but fails to reveal how the model handles edge cases, unexpected errors, or long-running tasks.
Finally, a skills gap exists: teams building agentic capabilities often come from application development, where synchronous request-and-response patterns are the norm. In contrast, agentic workloads behave like data pipelines: they are long-running, experience partial failures, and are expensive to re-execute. Agentic AI did not create a new class of operational problems; rather, it removed assumptions that existed for decades, and most incidents stem from these process gaps rather than model quality. Addressing them requires defining correctness in advance, measuring cost per completed unit, maintaining sufficient documentation for investigation, and designating named personal accountability.