Blog
Agentic Operations
2 September 2026

A well-functioning agent is a feature. It isn't a capability.

Blog
Abstract network illustration: pale interchangeable nodes on the left gaining weight and colour toward a solid red cluster on the right

Every quarter, agents get cheaper, more reliable and easier to stand up. That trend has one direction and it isn't going to reverse.

Which means a well-functioning agent is a technological feature. It is not a business capability. It sits in the same category as a message queue, a database replica or a container runtime — necessary, unremarkable, and priced by a market that keeps driving it down.

Nobody builds their own database engine to gain competitive advantage in logistics. But teams are building their own agent infrastructure right now, and calling it capability.

The shift happened fast enough that most operating plans have not caught up with it. On short, bounded tasks, the best agent systems now match or beat human experts and return the answer sooner. Tool use, structured output and orchestration have hardened from brittle scripts into production patterns. And what needed a bespoke build eighteen months ago is now assembled from pre-built models, frameworks and connectors in weeks. Three separate things had to become true for agents to be deployable rather than demonstrable, and all three now are.

Infrastructure gets cheaper. That's what infrastructure does.

If agents are becoming infrastructure, then every euro you spend building that layer buys an asset that depreciates faster than you can amortise it. The rate is not a matter of opinion. Holding capability fixed, the price of running a model at GPT-3.5 level fell from $20.00 to $0.07 per million tokens in about eighteen months — a factor of 280 — and the underlying rate for equivalent capability runs at roughly tenfold a year. Whatever you spend on the harness, the orchestration plumbing and the eval scaffolding this year, the market will have shipped the same thing, better, for a fraction of the price before you have finished amortising it.

Log-scale line chart: the cost of querying a model at fixed GPT-3.5-level capability falls from $20.00 per million tokens in November 2022 to $0.07 by October 2024, a factor of 280, with each gridline marking a tenfold drop

You will still be maintaining yours. Your engineers will still be defending it, because it was expensive and because they built it.

Two layers moving in opposite directions: the operating model, governance, business logic and operational knowledge appreciate, while the agent harness, orchestration plumbing, eval scaffolding, integrations and model serving depreciate

Both answers create the wrong kind of capability

Build it yourself and you own a maintenance obligation. That is the whole of it. You set out to acquire capability and you acquired a standing cost, on a depreciating asset, staffed by people who now cannot be redeployed.

Buy a pre-packaged solution and you can run fast but you inherit someone else's definition of correct. Their thresholds. Their exception logic. Their assumptions about how an order should be allocated and what counts as an acceptable delivery promise.

And customising it is worse than it looks. Every rule you configure, every tolerance you tune, every piece of operational knowledge you encode — you are writing your intellectual property into their platform, in their format, behind their interface. You will not get it out. The more you invest in making it yours, the less able you are to leave, which is precisely when you discover you no longer control it.

One option makes you a maintainer of depreciating infrastructure. The other makes you a tenant in someone else's operating model. Neither leaves you owning the thing that matters.

Is there a third option that both allows you build rapidly and retain your competitive advantage?

The third option

Buy customisable, pre-trained agent pods that give you the full mechanism, then configure them against your operating model.

You rent the infrastructure at whatever rate the market sets, and you re-rate it as the market falls. You keep the operating model, the governance, the business logic and the knowledge — in your own terms, where you can change them, take them elsewhere, or hand them to a different provider.

Three options compared: building it yourself leaves a depreciating asset, buying it packaged leaves tenancy rather than ownership, configuring pods leaves the asset that appreciates

The value is in the symphony, not the instruments

As agents commoditise, the hard part moves. It stops being how to build one and becomes how to manage what they produce: harnessing the knowledge, integrating with the process, guardrailing with business rules, and enforcing governance and accountability across all of it.

Programmes are not failing because the models are weak. They fail on integration, business logic, rules and policy, governance and accountability — every layer that sits between a working agent and a governed operation, and the layer nobody budgeted for.

This is why two or three agents rarely move the needle. They prove the parts work — which was never in doubt. But money doesn't leak from one place. It leaks across the whole operation, at every seam between stages, and through the long tail of cases nobody has the capacity to handle. Until the coverage is end to end, most of the value stays on the table.

Neither building your own nor buying packaged automation gets you there. One gives you a few excellent agents and no orchestration. The other gives you orchestration you don't control.

What this looks like in order-to-delivery

Take one operation end to end. No single leak in it is fatal, and no single agent fixes it. The money drains at every seam and through the tail, all at once.

StageWhere the money leaksPod that closes it
Order intakeOrders with errors, missing data or non-standard terms — fixed by hand, or let through to create downstream rework and promises you can't keep.Order Intake
Availability & allocationStock committed first-come-first-served rather than by customer value, so a tail account consumes what a strategic account needed.Availability & Allocation
Delivery schedulingDates promised without a real feasibility check, or set to a default — missed OTIF on one side, over-servicing on the other.Delivery Scheduling
Carrier operationsBooked to a default carrier or to premium freight the service level never required, and small orders shipped un-consolidated. The biggest cost-to-serve leak.Carrier & Booking
In-transitDelays and shortfalls surface too late to re-plan, so recoverable misses become real ones.Exception & In-Transit
Delivery & billingProof-of-delivery gaps, short-ships and discrepancies become disputes and short-pays — revenue leaking after the sale, and DSO drift.Delivery & Billing Handoff

Read down the middle column and the pattern is obvious. Each leak is small on its own. Fix one with one agent and the number barely moves, because the money is leaving from every seam and the whole tail simultaneously.

The leak is distributed. So the coverage has to be distributed too.
Six seams across order-to-delivery, each with the leak it causes and the pod that closes it, all bound together by an orchestrator that carries the operating model

What that buys, and why

Six pods covering the six seams, plus the orchestration that runs them as one operation against your policies. On the deployment I have in mind, that full scope stood up in six weeks.

Not because the agents were special. The agents were pre-built — that is the entire point of them being infrastructure. It went quickly because the coverage was complete and the orchestration was ours to configure rather than ours to invent.

What to do instead of the build-or-buy meeting

  • Stop funding the depreciating layer. Whatever you have budgeted to build harness and orchestration plumbing, redirect it. The market is going to hand you that layer cheaper than you can build it, on a schedule you don't control and don't need to.
  • Write the operating model down before you evaluate anything. Your segmentation, your service policy by segment, your thresholds, what escalates and to whom. If you can't state it, no platform can execute it and no vendor can be held to it.
  • Test any platform on one question: can I take my logic out? If your rules, tolerances and operational knowledge only exist inside their interface and their format, you are not customising a product — you are depositing your IP with them.
  • Scope coverage, not pilots. Two agents prove the technology. They will not move your P&L, because the leak is distributed. Map the seams first, then cover them.

You don't build the agents. You place them where the money leaks and run them as one operation, against rules you wrote and can change.

That's the symphony. The instruments are getting cheaper every quarter — which is exactly why they were never the thing worth owning.