What a running AI agent actually costs
Build an operating-cost model for an AI agent from inference, tools, hosting, monitoring, retries, human review, maintenance, and the completed work it produces.
A running agent has several bills and a human operating load. Count every call and review needed to complete the process, then compare the result with the value of the completed work.
Syntalith
The cost of an AI agent starts after the build is finished. Every case can create model calls, retrieval, tool calls, retries, logs, and human review. The service also needs hosting, access management, maintenance, and a plan for changes.
List the cost centres
Track each item separately:
- input and output inference;
- cached and uncached context;
- embeddings or search requests;
- browser, database, messaging, and other tools;
- hosting, queues, storage, and backups;
- monitoring, traces, alerts, and incident work;
- human review, correction, and approvals;
- maintenance, model changes, security updates, and support;
- integration and provider account fees.
Keep one line for one responsibility. A platform invoice can cover several items, so reconcile it with the process log and the account that paid it.
Use a cost-per-case model
For a completed case, record every call and human step:
case cost = inference + retrieval + tools + hosting share + observability + retries + review
Then calculate:
monthly operating cost = case cost x completed cases + fixed service costs
Use the actual provider rate card and measured tokens. Separate input and output rates, cache rules, minimum charges, and currency. If a case can fail and restart, include the retry path. If a person reviews only exceptions, measure the exception share and average review time.
The denominator should be accepted completed cases. A draft rewritten from scratch remains a review cost and does not count as a completed case.
Measure the process signals
The agent log should expose:
- case identifier and process version;
- model and provider for each call;
- input and output tokens or provider usage units;
- cache status;
- tool, retry, and fallback events;
- review decision and correction time;
- final accepted or rejected outcome.
Avoid storing sensitive source content in an analytics table when a reference and retention policy will do. Protect usage logs because they can reveal customer and company activity.
Understand the fixed work
Usage can be small while the service still consumes engineering time. Budget for access reviews, dependency updates, prompt and schema tests, data-source changes, model selection, on-call response, backup restores, and incident investigation. Name the person who owns each task and record expected response time.
The AI agent maintenance guide can help structure that service. The Syntalith pricing page describes how an implementation and maintenance proposal is scoped. The proposal should keep build, usage, and ongoing support as separate lines.
Three levers that usually matter
- Route simple classification and extraction to a suitable model, while reserving a stronger model for cases that need it.
- Keep context focused and retrieve only the source needed for the current case.
- Validate early so a malformed result does not trigger extra tools or a long retry chain.
Measure each change against accepted-case cost and review time. A lower token bill can be a bad trade when corrections or escalations increase.
Decide if the process earns its place
Compare accepted-case cost with the current cost of completing the same case. Include delay, correction, support, and the risk of a wrong action. If the agent cannot reach the target after a controlled evaluation, keep the process manual or narrow the task before expanding it.
A practical cost brief
Bring these values to a review:
- completed cases per period;
- average calls and tokens per case;
- tool and retrieval calls;
- retry and escalation rates;
- review minutes per accepted case;
- fixed hosting and support work;
- provider rate cards and account fees;
- target quality and availability.
A process scan can turn those inputs into a cost-per-accepted-case model and an operating plan. The result should make the volume assumptions visible and show which line changes when the workflow changes. The AI cost-management guide explains how to reconcile that model with provider reports and budgets.
Free process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary of what to automate first and the likely cost range.
The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form