When Your Buyer Is an Agent: Authority Limits for AI That Spends

By NATARAJA Team

Our Executive Authority Brief on AI incidents carries a case worth sitting with. An autonomous treasury system executes FX trades at machine speed. Its authority was designed for continuous execution under all market conditions. Volatility exceeds the expected range, and the system keeps trading, correctly, within the authority it was given, while the losses compound. Afterwards, everyone looks for the operator error. There is none. The failure manifested at execution, but it was created earlier, at the moment someone designed an authority with no boundary for the condition that eventually arrived.

That is the shape of every agent-spending incident that has not happened yet. The agent will not malfunction. It will do exactly what it was permitted to do, and the permission will turn out to be the problem.

Agents are becoming buyers

This is no longer hypothetical. Agents already renew SaaS subscriptions, scale cloud capacity, adjust advertising budgets, reorder inventory, and transact on marketplaces. The next step is visible: agents negotiating with vendors' agents, committing spend at a speed and volume no approval workflow was designed to see.

The right question is not whether an agent can buy. It plainly can, the moment it holds a payment credential. The right questions come from our Executive Guarantees Brief on the Sovereign Decision, and they are worth quoting exactly, because they apply to every autonomous system an organisation runs:

  1. What commitments can this system create?
  2. What institutional exposure can those commitments generate?
  3. How is that authority governed?

A purchasing agent is the most literal case those questions will ever meet. Every purchase is a commitment. Some commitments bind the institution for years. And in most organisations today, the third question has no answer at all.

Access is not authority

The standard objection is that the agent's permissions are already managed: it has a scoped API key, a virtual card with a limit, a service account with defined roles. All of that is access control, and access control answers a different question.

Permission to call the payment API says nothing about what the agent may commit: which vendors, what contract terms, what cumulative exposure, what kinds of obligation. A card limit caps a single charge; it does not stop an annual contract with an auto-renewal clause, a purchase that creates a data-processing relationship, or forty small charges that sum to what one large charge could never have cleared.

We covered the wider version of this in AI agent inventory and non-human identity: the gap between an agent's intended entitlements and its effective ones is where incidents live. Spend authority is the sharpest instance, because the agent that can spend is the agent whose effective entitlements matter most, and most enterprises cannot list which agents hold it.

The mandate: delegation done explicitly

Our Executive Authority Brief on Delegated Authority draws the distinction that decides everything downstream. Implicit delegation gives a system broad objectives with undefined execution limits, and produces ambiguous boundaries, reactive oversight, and post-incident clarification. Explicit delegation defines the authority in advance: a written mandate, quantified limits, defined override rights. Its consequence is the one procurement actually wants: a predictable execution scope with a clear accountability locus.

For a purchasing agent, the mandate needs five elements, and each should be written as a number or a list rather than a sentiment:

  1. A per-transaction ceiling. The most an agent may commit in one decision.
  2. A period budget. Cumulative spend across a defined window, checked against actual spend to date, so many small purchases cannot quietly outrun one large limit.
  3. A vendor scope. An approved list or register, with anything off-list requiring a human by construction.
  4. A category scope. What kinds of purchase this agent exists to make, stated positively, so a renewals agent cannot originate a platform selection.
  5. Explicit exclusions. The terms the agent may never accept alone: auto-renewals beyond a set period, annual commitments, anything that expands access to customer data.

One design rule matters more than the five elements: silence is escalation, not permission. When a request needs a limit the mandate does not state, the check fails rather than waives. The brief's phrase for the alternative is exact: undefined delegation expands silently. Every mandate gap an agent is allowed to interpret in its own favour becomes authority nobody granted.

The gaming screen

Our Executive Guarantees Brief on Boundary Enforcement makes an uncomfortable observation: every boundary simultaneously creates an incentive to route around it. This is familiar with humans, and procurement teams already watch for it. Purchase splitting, amounts sitting just under the approval threshold, repeated small orders to one vendor.

An agent optimising toward an objective will discover the same routes without any intent to deceive, which is precisely why the screen cannot rely on intent. The checks are structural: does this amount sit suspiciously close beneath the ceiling; does this purchase look like one slice of a series; would the reconstructed whole have cleared the limits that the slices individually do? Where a split is suspected, the honest output is the reconstruction itself, with the question it forces: this looks like part of a larger purchase, and the larger purchase is what needs approval.

Govern by reversibility, and the gate that only arms when it matters

The objection to any of this is speed, and it deserves a straight answer because we published it ourselves: reality moves at machine speed, organisational processes move at human speed, and the timing gap is fatal. If every agent purchase waits for a human, the agent is pointless.

The resolution is to govern by reversibility. Deterministic checks, ceilings, budget arithmetic, allowlists, term screens, run at machine speed and leave their own attested record; they cost nothing and delay nothing. The human gate exists but stays disarmed for routine, in-mandate, reversible purchases: those clear in seconds with a record that reads as a decision rather than a rubber stamp. The gate arms only when something warrants it: a limit failed, a request looked novel, or the commitment is one the institution cannot walk back.

Reversibility deserves the emphasis, because it outranks amount. A USD 150 monthly renewal that cancels any time is the easy quadrant regardless of how many zeros you add. A small purchase that connects a vendor's platform to customer data for model training is irreversible in the sense that matters: cancelling the contract later does not return the institution to its prior state, because the data has already been processed. The amount is not the risk. The binding is.

What this looks like when it runs

We built this as a governed protocol and ran the two cases that define its behaviour. Both were synthetic tests on our own platform, which is the honest label for them, but the mechanics were the production mechanics.

The routine case: a replenishment agent requests a USD 150 monthly Grammarly renewal, in-mandate on every check. Five limit checks pass with the mandate lines quoted, the novelty screen returns routine with the split-series check explicitly clear, reversibility is a month at no exit cost. The run completes with no human involved and an attested record of exactly what was checked. Machine speed, evidence included.

The escalation: the same agent requests a USD 48,000 annual AI-platform licence, paid upfront, auto-renewing, including a connection to the customer data warehouse. The limits check fails on all five limits, with the arithmetic shown (96 times the ceiling). The novelty screen flags four independent grounds, including the one that matters most: a renewals agent originating an enterprise platform selection is not a purchase, it is a procurement decision arriving through the wrong door. The commitment assessment classifies it irreversible on data substance rather than contract terms. The run halts, the budget owner is summoned by email, and nothing exists as an authorization until a person decides.

The second case is the demonstration that counts, because the refusal is the product. An approval system proves itself on the request it does not wave through.

Frequently asked questions

What governance frameworks should enterprise retailers put in place before allowing AI agents to transact on their platforms?

Four layers, in order of construction. First, identity: every transacting agent holds its own credential, never shared, so each transaction attributes to one agent and the principal it acts for. Second, an explicit mandate per agent: per-transaction ceiling, period budget, counterparty scope, category scope, and stated exclusions, with silence escalating rather than permitting. Third, runtime enforcement in the transaction path: deterministic checks a correctly functioning agent cannot exceed, a gaming screen for split and threshold-adjacent patterns, and a human gate that arms on limit failures, novelty, or irreversible commitments. Fourth, an attested record per transaction, retained on a horizon that matches disputes rather than debugging. The common mistake is stopping after the first layer and calling payment-credential limits governance: a card limit governs the charge, not the commitment.

How do you set spending limits for AI agents?

Set them as an explicit mandate rather than scattered platform settings: one written document per agent stating the per-transaction ceiling, the period budget checked against actual spend to date, the approved vendors, the permitted purchase categories, and the terms the agent may never accept alone. Two disciplines make the limits real. Enforce them in the execution path, so the agent cannot exceed them however it is prompted. And treat mandate silence as a failed check rather than a permitted action, because every gap an agent may interpret in its own favour becomes authority nobody granted.

Can AI agents make purchases autonomously?

Yes, and for routine, reversible, in-mandate purchases they should: a renewal that cancels within a month does not deserve a human minute, and blocking it wholesale just moves the buying outside governance. The autonomy question is badly posed as all-or-nothing. The workable answer is bounded autonomy: the agent buys alone inside an explicit mandate at machine speed, every purchase leaves an attested record either way, and the human decides only where a limit fails, the request is novel, or the commitment cannot be walked back.

What happens when an AI agent exceeds its authority?

In a governed setup, it cannot complete the action: the request halts at the boundary, the accountable budget owner is summoned with a plain statement of what they are deciding, and the attempt is recorded whether approved or refused. That record matters as much as the block, because a pattern of near-limit attempts is how you discover a mandate that no longer fits the job. In an ungoverned setup, the question answers itself differently: the purchase completes, and the exceeding is discovered by finance, weeks later, as a line item nobody recognises.

Who is accountable when an AI agent buys something?

The principal whose authority the agent acted under, which is exactly why the mandate has to exist in writing before the agent spends. Every purchase should attribute two ways: the agent that executed it, and the named person whose delegated authority covered it. If no mandate covers the purchase, accountability collapses into the post-incident argument our briefs call reactive oversight, where design failure gets relabelled operator error. The treasury case is the caution: the system acted within the authority it was given, so the accountable party is whoever gave it, which is a design-time fact, not an execution-time one.

Where NATARAJA fits

The mechanics above run as a public protocol on our platform: agent-purchase-authority. The mandate enters as evidence, deterministic assembly runs at machine speed, the limit checks quote the mandate lines they relied on, the gaming screen looks for splits, the commitment assessment answers the three questions, and the human gate arms only on escalation. Every run, cleared or halted, leaves an attested record with a hash-chained trace.

If your agents are beginning to spend, request a governed pilot: bring one agent and its actual mandate, or the absence of one, and we will run its next ten purchases governed.

Related reading: AI agent inventory and non-human identity for finding the agents and mapping what they can actually do, and the speed objection, answered for why the gate does not slow the machine down. From the Executive Brief series: Delegated Authority, Sovereign Decision, Boundary Enforcement, and AI Incidents.