AI Software Development: How We Cut Agent Token Use

Webdelo's CTO opens the logs of AI coding agents building an algorithmic trading platform: where billions of tokens went, how scripts and larger tasks cut usage per merged PR from 50.3M to 22.6M, and which checks stayed mandatory.
— Estimated reading time: 17 minutes
cover

Why a quick prototype is not a ready system

Since 2006, we at Webdelo have built complex B2B enterprise systems, including FinTech, trading platforms, and ERP. I am Andrey Popov, CTO of Webdelo. For us, AI software development means agents write code while engineers manage the work and accept the results. A quick prototype does not mean a system can be trusted with money.

On our algorithmic trading platform, an error could cause a loss of funds or a duplicate order - an instruction to buy or sell. That makes independent code review and testing essential. But our logs revealed another problem: agents spent 59% of input tokens waiting.

When a client discusses Web Development with us or needs a web design agency for a complex product, I first ask what operations the system will perform. In this project, the main risk was inside the system: tracking funds and processing trading orders.

An AI agent is a program that calls a model and uses tools to complete a task. In our setup, a coordinator agent assigned work to three or four coding agents. Reviewer agents checked the changes, while people made architectural decisions and accepted the work.

Three terms are enough to understand the figures:

  • Token - a small piece of text used to measure the amount of information a model processes.
  • PR, or pull request - a set of changes proposed for inclusion in the main version of the code. A merged PR means those changes have been included.
  • Review - a check of the changes by another participant in the process.

How many tokens AI coding agents used and what they read

From October 5-9, 2026, we counted 276 Codex sessions, 25,827 model requests, and 3.25 billion input tokens. During that period, 94 PRs were merged into the main codebase. Cached context accounted for 97.4% of the input - information already sent to the model that it processed again.

Context is the work history available to the model: instructions, messages, and command results. In our process, a request contained an average of 126,000 input tokens, while a response contained about 600. Generated code and model reasoning together amounted to less than 1% of the input volume.

Input token breakdown for October 5-9
Repeated, cached context 97.4%
New information 2.6%

Source: Webdelo logs. The chart shows only model input, not its responses.

For me, this came down to a simple formula: input token usage = number of requests × average context size. Optimizing token usage means reducing unnecessary calls and the amount of history sent with each one.

Why agents waited instead of working

About 1.9 billion input tokens, or 59% of all input, went toward waiting and checking status. An agent would ask: "Have the tests passed yet?" or "Has the worker replied?" Each question sent the accumulated context to the model again.

In our configuration, the tool returned control to the model within 300 seconds. A full check took 15-23 minutes. During a single check, the model woke up three to five times. The coordinator even reacted to a worker's routine "I'm alive" message.

At the same time, there were long gaps between useful actions. The coordinator could wait half an hour for a reminder, even though the reviewer finished in 5-15 minutes. During one 24-hour measurement period, we counted 11.6 hours of this idle time.

Share of waiting requests by day and role
Coordinator Workers
October 5 69% / 57%
October 6 61% / 59%
October 7 77% / 60%
October 8 58% / 45%
October 9 71% / 51%

Source: Webdelo logs. This chart compares requests. The 59% figure above refers to input tokens.

Waiting accounted for a large share of calls in both roles. This category included necessary waits for the agents' own tests. The problem was calling the model repeatedly just to check status.

We found other waste in how tasks were organized. The plan had grown to 550 small items, of which three were completed in a day. There were 43 review sessions versus 32 development sessions, and a new coding agent was often started to make fixes.

181 million tokens and not a single completed task

One session ran for eight hours, made 643 requests, and used 98 million tokens. A related task went through five rounds of "author - review - fix" across ten sessions. Together, they used 181 million tokens without completing a single task.

An engineer stopped this loop. We saved the branches containing the changes and postponed the work until the main path was ready. The agent did not make that decision on its own.

What we changed in AI agent orchestration

We moved repeatable actions into scripts and grouped tasks into testable modules. One author now owned a module through to merge, while one reviewer checked it and any subsequent fixes. We also began explicitly setting the model's reasoning level when assigning work.

The dividing line is simple: if an action follows fixed rules and its result fits on one line, a script handles it. The model handles task analysis, coding, and finding the cause of an error. We moved the coordinator's memory from a long conversation into a compact table in a file.

Waiting became a single script call that filters out routine status signals. At intermediate steps, we check the changed files. We run the full check on the finished module alongside the review, removing 15-23 minutes of sequential waiting.

AI agent workflows before and after optimization: status polling replaced by scripts and targeted fixes
The diagram shows a process change, not measured usage or speed results.

Before and after

Area of workBeforeAfter
CoordinationThe model polls workers and keeps state in the conversationA script waits for a meaningful message, and state is stored in a file
TasksSmall items, with a new author for the next fixA testable module, with one author through to merge
ScriptsAssignment, submission, acceptance, and checks require chains of model callsEach operation starts with a single call
TestsFull code style checks at every step, with a shared queue for checksChanges checked at intermediate steps, full checks before submission, stop at the first error
ReviewRepeated checks of the entire moduleOne full review, then the same reviewer checks only the fixes
Model settingHigh reasoning level for all authorsLevel explicitly assigned by task risk, without the maximum setting

Instead of 550 small tasks, we defined 20 modules, each of which could be run and tested. For example, in a crm project, one such task could cover the full request approval workflow. For an ecommerce website, it could cover checkout with payment verification.

On one actual change, the author's check script completed its work in 49 seconds with a single call. Previously, starting and managing those checks took 10-30 model calls. This was the result of one specific run, not a promised time for every task.

Automation has its own cost. The scripts required a separate PR and 45 tests. They need maintenance, and the message-waiting logic must not fail: a script could lose an important worker response.

How we assign model reasoning levels

Medium and High are model reasoning levels. We assign Medium to build tasks and environment setup. For logic involving money and recovery after failures, we use High for both the author and the full review.

In our logs, some usage and step-duration metrics were about 30-60% higher on High. That is not an increase by orders of magnitude. Unnecessary calls and repeated review rounds consumed far more tokens.

Median input token usage per review session
Full review on High 4.1 million
Full review on Medium 1.4 million
Review of fixes only 0.5 million

Source: Webdelo logs. These figures show token usage, not monetary cost. Tasks assigned to High and Medium differed.

Median review session duration
Full review on High 29 min
Full review on Medium 3 min
Review of fixes only 2 min

Source: Webdelo logs. The median is the middle observation: half the sessions were shorter and half were longer.

The charts show the resources used by different types of review in our process. They do not prove that Medium would do the same work faster. Reviewing only the fixes naturally takes less time than reviewing an entire module.

High proved useful for financial logic: in two days, reviews found three serious defects. One would have caused a duplicate order when the system was running. We kept those checks.

An instruction about choosing the reasoning level was not enough. The rule "use Medium for ordinary code" already existed, but the coordinator started all 12 authors on High. An engineer spotted this in the logs: the agent had no usage tracking of its own.

What optimizing agentic coding showed

Over five days, usage per merged PR fell from 50.3 to 22.6 million input tokens. In the comparison windows, the merge rate rose from 0.75 to 1.6 PRs per hour. These are process metrics, not the cost of an equal amount of work: PRs became larger after the changes.

Million input tokens per merged PR by day
October 5 50.3
October 6 37.6
October 7 37.4
October 8 35.1
October 9 22.6

Source: Webdelo logs and PR history. October 9 includes data through 11:40 UTC. PR sizes varied.

The chart shows a steady decline in usage per PR. In a separate comparison of windows before and after the changes, the figure fell from 44 to 25 million tokens, or by 43%. That percentage applies specifically to those comparison windows, not to the first and last points on the chart.

After replacing the coordinator, 45 tasks from the original breakdown were completed in the first eight hours, compared with three in the previous 24 hours. Measured coordinator idle time fell from 11.6 hours to zero. The new windows contained no pauses longer than ten minutes, which was our threshold for counting idle time.

The mix of new lines of code changed too. The share of stubs - automatically generated placeholders - fell sharply. Tests and product code that performs the system's operations accounted for a larger share.

Breakdown of new lines of code: October 7-8
Stubs 39%
Tests 34%
Product code 19.5%
Other 7.5%

Source: changes in the Webdelo repository. Shares are based on 176,000 added lines.

Breakdown of new lines of code: overnight, October 8-9
Stubs 3%
Tests 47%
Product code 42%
Documentation and other 8%

Source: changes in the Webdelo repository. Shares are based on 28,700 added lines.

The charts show that work shifted from placeholders to system behavior and testing. The number of lines alone does not measure code quality.

We did not fully solve the waiting problem. On October 9, it still accounted for 71% of coordinator requests because of the tool's 300-second limit. The coordinator's share of total usage remained around 30%: the process produced more results, but the usage breakdown barely changed.

Limits of our measurements

  • PRs varied in size. Before the changes, a PR usually completed a small task. Afterward, it covered a module of 1,500-3,000 lines.
  • Medium and High received different tasks. We did not compare the same work at both levels.
  • Post-merge quality was not measured. We did not count later defects or establish equal quality across reasoning levels.
  • Monetary costs were not calculated. We worked within subscription usage limits.
  • The project is not finished. By the end of the observation period, about 65% of the sprint criteria had been met. Financial reconciliation and failure recovery scenarios were still ahead.

How token costs affect AI development cost

Tokens translate into actual expenses through a subscription and its limits, or through usage-based model billing. We worked within subscription limits, so we did not measure dollar savings. Unnecessary requests bring forward the point when work must stop until the quota resets or requires additional spending.

Depending on the service, billing takes the form of:

  • a paid subscription with usage limits
  • additional usage packages or quotas
  • an API - programmatic access billed by usage

With token-based billing, cached input may cost less than new input, but it is not free. In our records, it also consumed usage limits. For Claude Code, the terms are described in Anthropic's official guide to models and limits.

In our process, cutting unnecessary calls mattered more than lowering the model's reasoning level. I focus on the resources needed to reach an accepted result, including retries and failed checks. This approach to assessing a completed task is also described in a Habr tutorial on reducing token usage.

For product budgeting, it helps to separate development from customer acquisition. If a corporate website connects to the platform, its updates should be budgeted separately from digital marketing.

Seo and geo have their own outcomes and measurement timelines. Using fewer tokens to write code does not mean reducing the entire digital product budget.

On our side, optimization required three log analyses over three days and three instruction documents of 30-40 lines each. An agent performed the analyses at an engineer's request. The team also wrote and tested scripts, so those documents do not represent all the work involved in making the changes.

What must not be optimized at the expense of quality

Developing algorithmic trading platforms requires checks of financial logic and recovery after failures. We reduced the repetitive work around those checks. Scenarios that could reveal a loss of funds remained mandatory.

The coordinator itself proposed three constraints in response to our instructions. We agreed with them:

  • A confirmed defect must not be hidden to stay within a limit on fix cycles.
  • Checks involving a real database, emergency shutdown, and financial reconciliation are mandatory.
  • The first failed check result is archived. It must not be replaced by a successful rerun.

Which decisions stayed with the engineer

I define a task as a module that can be run and tested. At the previous pace, the remaining 343 small tasks would have meant about 114 days of work. This was a linear estimate based on the old process. It exposed a problem in the task breakdown itself.

We rejected a proposal to temporarily connect only one exchange to the new gateway. We chose the full workflow for three exchanges: a temporary setup in a financial system can easily become permanent. This decision increased the current scope of work.

We migrated the strategy with minimal changes and a verification run to compare results. Rewriting it to make it "cleaner" would have added the risk of changing its behavior. Preserving proven logic mattered more here than making the new code look tidy.

An engineer needs to read the logs regularly and adjust the rules. In our case, daily checks helped us spot unnecessary actions returning. An instruction to an agent did not, by itself, ensure that the process was followed.

When discussing AI development optimization, an enterprise platform, or support for complex software at Webdelo, I start with the current process and acceptance criteria.

What I learned from this sprint

In our case study, the main waste came from how agent work was organized. Logs helped us find unnecessary waiting. Scripts and larger tasks reduced repeated model calls. We kept the financial risk checks.

The same principle applies to B2B platforms, FinTech, and ERP: measure the process first, then automate routine work and make responsibility for the result clear. Enterprise software development with agents requires engineering decisions at every one of these steps.

Frequently Asked Questions

What is AI-first software development in simple terms?

It is an approach where AI agents write the code while engineers manage the work and accept the results. At Webdelo, a coordinator agent split the work between three or four coding agents, reviewer agents checked the changes, and people made the architecture decisions. A quick prototype is not a reliable system: where software handles money, independent code review and tests stay mandatory.

Why do AI coding agents use so many tokens?

Most of the usage comes not from writing code but from resending the accumulated work history. In the Webdelo case, agents used 3.25 billion input tokens in five days, and 97.4% of that was cached context, meaning information already sent before. About 59% of input tokens went to waiting: an agent asked whether the tests had finished, and each such question sent the whole context to the model again. The formula is simple: the number of requests multiplied by the average context size.

How can you reduce token usage in development with AI agents?

Webdelo moved repeatable actions to scripts: waiting, handing out and submitting tasks, and running checks. Instead of 550 small tasks, the team defined 20 modules that can be run and verified, each with one author and one reviewer. The model level is now assigned by task risk. In five days, usage per merged PR fell from 50.3 to 22.6 million input tokens and the pace rose from 0.75 to 1.6 PRs per hour, but PRs also became larger, so these are process figures and not the price of equal work.

How do tokens affect AI development cost?

Tokens turn into real spending through a paid subscription with usage limits, the purchase of extra packages or quotas, or pay-as-you-go API access. Unneeded requests bring closer the moment when work stops until the quota resets or requires extra payment. Cached input may cost less than new input, but it is never free. Webdelo worked within subscription limits, so the team did not measure savings in dollars.

What must not be cut when optimizing AI software development?

Checks of financial logic and recovery after failures must stay. Webdelo reduced the repeatable work around these checks, but the checks with a real database, an emergency stop, and financial reconciliation remained mandatory. A confirmed defect cannot be hidden to limit the number of fix rounds, and the first failed check result is kept in the archive. In two days, this kind of review found three serious defects, one of which would have caused a duplicate order.

Why is an engineer still needed when working with AI agents?

An agent does not track its own usage and does not always follow instructions. In the Webdelo case, two tasks used 181 million tokens without being completed, and it was an engineer, not the agent, who stopped that loop. The coordinator started all 12 authors on the high model level although the rule said otherwise, and a person noticed this in the logs too. So an engineer needs to read the logs regularly, adjust the rules, and decide how tasks are split and how the architecture is built.

How reliable are these figures and what are their limits?

The figures come from agent logs and PR history over five days, but the measurements have limits. PRs differed in size: before the changes one PR closed a small task, afterwards a module of 1,500-3,000 lines. The Medium and High levels were not compared on the same task, and defects after merging and money spent were not counted. The project is not finished: about 65% of the sprint criteria were met by the end of the observation period.

cookies We use Cookies

We use cookies to improve website performance, personalize content, and analyze traffic. You can choose which categories of cookies to allow. For more information, please see our Cookie Policy. You can change your preferences at any time.

Essential (Required)

Ensure the website functions properly (navigation, access to secure areas). Always enabled and can only be changed in your browser settings.

Analytics

Help us understand how you use the website so we can improve our services. Do not collect personal data. We use several analytics tools for this purpose.

Advertising

Used to deliver personalized ads and measure the effectiveness of advertising campaigns.

On this site we also use the OpenAI measurement pixel for ChatGPT Ads. It stores a first-party identifier named __oppref in your browser. When you send us an enquiry, a hashed form of your email address and phone number may be transmitted to OpenAI in the USA so that the enquiry can be attributed to the ad you clicked; we never transmit them in readable form.