Bilal Tahseen
All posts

I Put a Budget Cap and Kill Switch on Every Production Agent

· Bilal Tahseen

I watched an agent pass every tool contract, stay inside its allowlist, and still burn a day's LLM budget before lunch. The call was well-formed. The schema was valid. The write never happened — shadow mode caught that. What nobody capped was the retry storm around a flaky search tool. The model kept rephrasing, the contract kept admitting the call, and the meter sat in a comment in the system prompt: "stop if you spend too much."

The model did not stop.

That's the gap this post is about. Tool contracts decide whether a call is well-formed. Shadow mode decides whether a write is safe. A golden-set eval gate decides whether a build ships. None of those stop a live agent from looping, retrying, or tool-spamming its way through your OpenAI bill. For that you need AI agent cost control in the control plane: hard LLM budget limits, a ledger the model cannot see, and an agent kill switch that does not ask the model for permission.

This is not a drop-in product. The sketches below are the shape of a system I assemble when an agent can spend real money. Your store, your reservation protocol, and your alerting sink will differ. If you paste this into production as-is, you will have a demo with a counter, not a budget.

Why the prompt is not a budget

Teams put spend limits in the system prompt because it feels like control. "Do not exceed 50 tool calls." "Stop after three failed searches." "Keep cost under a few dollars." The model treats those as suggestions. I have seen it lose count, reinterpret a retry as a new query, and keep going because the user sounded urgent.

A prompt cannot:

  • See provider invoices, cached-token discounts, or tool-call overhead the tokenizer never counted
  • Coordinate two concurrent runs for the same tenant so they don't both "stay under" a limit they both exceed
  • Reserve spend before a call, then commit or release after the provider responds
  • Survive a process crash mid-loop without double-counting or under-counting
  • Kill a sibling worker that is already in flight

So I keep the budget outside the model. The agent requests work. Middleware estimates cost. A ledger reserves capacity. The runtime only proceeds if the reservation succeeds. When signals look like a runaway — burn-rate spike, error spike, tool loop — a circuit breaker degrades or kills. The model is a tenant of that system. It is not the accountant.

The control plane: meter, ledger, breaker

I treat spend like an admission-control problem, not a logging problem. Logging tells you what already happened. A ledger tells you whether the next call is allowed to happen.

Budget-aware agent runtime: an incoming request flows through metering middleware, a budget ledger, the agent runtime, tool calls, and a circuit breaker, then splits into a kill-switch stop path or a degrade-to-read-only path

Request. Incoming user or system traffic. Identity must already be resolved: tenant, agent id, run id. If you cannot name the payer, you cannot cap them.

Metering middleware. Classifies the next action (prompt, completion, embedding, tool, retrieval), counts units, and estimates cost in a single currency — I use integer micro-USD, never floats. It does not trust the model's own token estimate. It uses the provider's tokenizer plus a conservative padding factor you tune, labeled as an example: I often pad completions by 15% until the usage callback lands.

Budget ledger. Tracks committed spend, open reservations, and remaining capacity across three hard scopes: this run, this tenant, this UTC day. Reservations are first-class. If the ledger cannot reserve, the call does not start.

Agent runtime. Executes only after reservation. It sees a remaining-budget hint so it can plan, but the hint is not enforcement. Enforcement already happened.

Tool calls. Every tool invocation is a metered event: call count, estimated tokens to parse the result, and any billed external API. Tool spam is a budget event, not just a quality event.

Circuit breaker. Watches live signals: error rate, burn rate, repeated tool signatures. It does not wait for the daily invoice.

From the breaker you get two exits:

  • Kill switch. Stop all execution for that run, and optionally for that tenant, when a hard cap or a runaway signal trips. In-flight calls are cancelled where the SDK allows it. Further reservations fail closed.
  • Degrade, read-only. Soft threshold. Cheaper model, fewer tools, no writes, human handoff. The agent can still answer from context. It cannot spend its way out of the hole.

The principle is boring on purpose: control spend, protect the system, preserve leftover value. A dead agent that already spent the budget is a worse outcome than a degraded one that still answers.

Three hard caps, none of them optional

I never ship a single global "max tokens" integer. One number cannot express a bad run, a noisy tenant, and a bad day at the same time. I enforce three independent ceilings, all outside the model, all fail-closed.

Per-run. Caps one conversation or job. This is how you stop a single loop from eating the tenant's day. Example threshold I start with in staging, not a promised production default: $1.50 USD or 150,000 tokens, whichever hits first. The run dies even if the tenant still has daily headroom.

Per-tenant. Caps one customer (or one internal team) across concurrent runs. Two parallel jobs cannot each "stay under" the per-run cap and overshoot the tenant. Example starting ceiling: $40 USD per UTC day per tenant, plus a concurrent-reservation cap of 3 active runs.

Per-day (account). Caps you, the operator, across every tenant. This is the backstop when a new tenant, a bad deploy, or a retry amplifier hits every agent at once. Example operator ceiling: $250 USD per UTC day for the whole environment, with a page if you cross 70% of that.

USD and tokens are both first-class. Token-only caps lie once you mix models, cached prefixes, tool-call charges, and embeddings. USD-only caps lie when a cheap model loops a million times. I convert tokens to micro-USD with a versioned price table keyed by (provider, model, unit). When a price changes, the table version changes. Old runs keep the version they reserved under so the ledger stays auditable.

Hard means the ledger refuses the reservation. There is no "ask the model if it is done." There is no prompt that can raise the cap. A human can raise a cap; the agent cannot.

Soft thresholds: degrade before you kill

A hard cap that only fires at 100% is a cliff. I put a degrade band under it so the system gets cheaper and safer as spend climbs, instead of running full-fat until it dies.

Example bands I use as a starting policy, not as a claim about anyone's production numbers:

  • Below 50% of the tightest applicable cap: full tool set, primary model, writes still gated by shadow mode.
  • 50–70%: drop the expensive model. Switch to the cheap completer for planning. Keep write tools, but require the existing shadow/promote path.
  • 70–90%: read-only. Strip write tools from the allowlist for this run. Disable web-scale retrieval if that path is priced per call. Offer a human handoff instead of another retry.
  • 90–100%: no new reservations except a single short "I have to stop" completion, itself reserved against a tiny leftover slice. Then kill.
  • Over 100% or kill signal: kill switch. No leftover slice. No apology token stream if it would require a new billed call.

Degrade is a policy object the runtime consults after every ledger snapshot. The model does not choose to degrade. The allowlist shrinks, the model id changes, and write tools disappear from the schema the model is given on the next turn. If you leave the write tools in the prompt and "ask nicely," you are back to hoping.

Human handoff is part of degrade, not a separate product. When the run crosses the read-only band, I enqueue an ops ticket with the run id, tenant, spend so far, last tool signature, and the reason code (BUDGET_DEGRADE, BURN_RATE, TOOL_LOOP). The user-facing message is short: the agent stopped taking actions and a person will continue. Do not spend another completion inventing a soothing essay.

The kill switch: loops, retry storms, tool spam

Budget exhaustion is only one tripwire. Some agents will blow the cap honestly. Others will burn a hole in a few seconds with a tight loop while the USD counter still looks fine because each call is cheap. I trip the kill switch on behavior, not only on dollars.

Signals detected — burn-rate spike, error spike, and tool loop — each branching to continue, degrade to read-only, or trip the kill switch

Three signals I always wire. The numeric cutovers below are example thresholds I put in staging to have something real to tune. They are not universal SLOs.

Burn-rate spike. Spend per minute versus a short baseline for that agent. Example: if the last 2 minutes are 3x the previous 15-minute rate, and absolute spend in that window exceeds $0.40, degrade. If it hits 6x or the per-run cap would trip within the next estimated call, kill. This catches "the retrieval corpus exploded" and "we accidentally pointed staging at the 8x model" faster than a daily budget.

Error spike. Tool or provider errors clustered in one run. Example: 5 failures in 8 attempts, or a 25% error rate over a 20-call window. Continue only if the errors are isolated timeouts with backoff. Degrade if the agent is still retrying the same class of error. Kill if it is amplifying — each error triggers two more calls.

Tool loop detected. Same tool name plus same canonical argument hash, repeated. Example: 4 identical calls in one run, or 6 calls that only differ by a trailing retry nonce the agent invented. This is the classic "search returned nothing, rephrase, search" death spiral. Degrade on the first detection (strip that tool). Kill if it continues on a sibling tool or if the loop shares a signature with a write.

Retry storms sit across those three. A flaky HTTP client with unbounded retry, wrapped by an agent that also retries, is how you turn one 500 into eighty completions. The kill switch has to sit above the HTTP retry policy. If the SDK retries under you, your meter never sees the extras until the invoice does. I disable SDK retries on provider calls and let the control plane own retry: bounded, jittered, and charged against the reservation before the retry goes out.

Kill is fail-closed and sticky for the run. Resetting it is a human action. I do not auto-reset because "the error rate dropped" — of course it dropped, you killed the agent.

Implementation sketch: reservation, not a counter

A counter you increment after the call is an invoice reconcilation tool. It will not stop the call that puts you over. I reserve first, execute, then commit or release.

This is deliberately incomplete. There is no store, no distributed lock, no clock-skew handling, no exactly-once usage callback. Those are the parts that make it real, and they are specific to how you deploy.

type MoneyMicros = number; // integer micro-USD; never a float
type TokenCount = number;

type BudgetScope = 'run' | 'tenant' | 'day';

interface BudgetLimit {
  scope: BudgetScope;
  hardUsd: MoneyMicros;
  hardTokens: TokenCount;
  degradeUsd: MoneyMicros;
  killOnSignal: boolean;
}

interface Reservation {
  id: string;
  runId: string;
  tenantId: string;
  estimatedUsd: MoneyMicros;
  estimatedTokens: TokenCount;
  priceTableVersion: string;
  createdAt: string;
}

type RunMode = 'full' | 'degraded_readonly' | 'killed';

interface LedgerSnapshot {
  reservedUsd: MoneyMicros;
  committedUsd: MoneyMicros;
  committedTokens: TokenCount;
  mode: RunMode;
}

interface BudgetStore {
  // Your transactional store. If this is an in-memory map, you do not have a budget.
  reserve(input: Reservation): Promise<'ok' | 'rejected'>;
  commit(id: string, actual: { usd: MoneyMicros; tokens: TokenCount }): Promise<void>;
  release(id: string): Promise<void>;
  snapshot(runId: string, tenantId: string): Promise<LedgerSnapshot>;
}

interface MeterEvent {
  kind: 'prompt' | 'completion' | 'tool' | 'embedding' | 'retrieval';
  model?: string;
  tool?: string;
  argsHash?: string;
  estimatedUsd: MoneyMicros;
  estimatedTokens: TokenCount;
}

class BudgetLedger {
  constructor(
    private readonly store: BudgetStore,
    private readonly limits: BudgetLimit[],
  ) {}

  async admit(event: MeterEvent, ctx: { runId: string; tenantId: string }): Promise<Reservation> {
    const snap = await this.store.snapshot(ctx.runId, ctx.tenantId);
    if (snap.mode === 'killed') {
      throw new KillSwitchTripped('run already killed');
    }

    for (const limit of this.limits) {
      const used = this.usedFor(limit.scope, snap, ctx);
      if (used.usd + event.estimatedUsd > limit.hardUsd) {
        throw new BudgetRejected(limit.scope, 'usd');
      }
      if (used.tokens + event.estimatedTokens > limit.hardTokens) {
        throw new BudgetRejected(limit.scope, 'tokens');
      }
    }

    const reservation: Reservation = {
      id: crypto.randomUUID(),
      runId: ctx.runId,
      tenantId: ctx.tenantId,
      estimatedUsd: event.estimatedUsd,
      estimatedTokens: event.estimatedTokens,
      priceTableVersion: 'you-must-version-this',
      createdAt: new Date().toISOString(),
    };

    const result = await this.store.reserve(reservation);
    if (result !== 'ok') {
      throw new BudgetRejected('run', 'reservation_conflict');
    }
    return reservation;
  }

  async settle(reservation: Reservation, actual: { usd: MoneyMicros; tokens: TokenCount }): Promise<void> {
    await this.store.commit(reservation.id, actual);
  }

  async abort(reservation: Reservation): Promise<void> {
    await this.store.release(reservation.id);
  }

  private usedFor(
    scope: BudgetScope,
    snap: LedgerSnapshot,
    _ctx: { runId: string; tenantId: string },
  ): { usd: MoneyMicros; tokens: TokenCount } {
    switch (scope) {
      case 'run':
        return { usd: snap.reservedUsd + snap.committedUsd, tokens: snap.committedTokens };
      case 'tenant':
      case 'day':
        // Snapshot must already be scoped. If you re-sum here without a store query, you will double-admit.
        return { usd: snap.reservedUsd + snap.committedUsd, tokens: snap.committedTokens };
      default: {
        const _exhaustive: never = scope;
        throw new Error(`unhandled budget scope: ${_exhaustive}`);
      }
    }
  }
}

Metering sits in front of the runtime, not inside the prompt builder.

class MeteringMiddleware {
  constructor(
    private readonly ledger: BudgetLedger,
    private readonly prices: PriceTable,
    private readonly breaker: CircuitBreaker,
  ) {}

  async aroundCall(
    event: MeterEvent,
    ctx: { runId: string; tenantId: string },
    fn: () => Promise<ProviderResult>,
  ): Promise<ProviderResult> {
    const signal = this.breaker.inspect(ctx.runId);
    this.enforceSignal(signal);

    const reservation = await this.ledger.admit(event, ctx);
    try {
      const result = await fn();
      await this.ledger.settle(reservation, this.prices.actual(result));
      this.breaker.recordSuccess(ctx.runId, event, result);
      return result;
    } catch (error) {
      await this.ledger.abort(reservation);
      this.breaker.recordFailure(ctx.runId, event, error);
      throw error;
    }
  }

  private enforceSignal(signal: BreakerDecision): void {
    switch (signal.action) {
      case 'continue':
        return;
      case 'degrade':
        throw new DegradeRequired(signal.reason);
      case 'kill':
        throw new KillSwitchTripped(signal.reason);
      default: {
        const _exhaustive: never = signal.action;
        throw new Error(`unhandled breaker action: ${_exhaustive}`);
      }
    }
  }
}

The breaker is a pure function over a ring buffer of recent events. If you "just use Datadog alerts" and page a human, the loop is already finished.

type BreakerAction = 'continue' | 'degrade' | 'kill';

interface BreakerDecision {
  action: BreakerAction;
  reason: 'burn_rate' | 'error_spike' | 'tool_loop' | 'ok';
}

class CircuitBreaker {
  inspect(runId: string): BreakerDecision {
    const window = this.window(runId); // you need a per-run ring buffer; this is not it
    if (this.identicalToolLoop(window, 4)) {
      return { action: 'kill', reason: 'tool_loop' };
    }
    if (this.errorRate(window) >= 0.25 && window.length >= 8) {
      return { action: 'degrade', reason: 'error_spike' };
    }
    if (this.burnMultiple(window) >= 3) {
      return { action: 'degrade', reason: 'burn_rate' };
    }
    return { action: 'continue', reason: 'ok' };
  }

  // recordSuccess / recordFailure / window / identicalToolLoop / errorRate / burnMultiple
  // are the actual system. Without them this class is a comment.
}

When KillSwitchTripped fires, I do four things in order: mark the run killed in the ledger so new reservations fail, cancel in-flight HTTP where the client supports abort, drop write tools even if a racing turn already planned them, and emit the alert. Order matters. If you alert first and the mark fails, the next worker keeps spending.

What I log, and what I page

If the ledger is the brake, the log is how you prove the brake worked. I do not log prompts into the spend stream. I log money, identity, and why a decision happened.

Every reservation, commit, and release gets a structured event:

  • timestamp, run_id, tenant_id, agent_id, reservation_id
  • scope that was tightest (run | tenant | day)
  • estimated_usd_micros, actual_usd_micros, estimated_tokens, actual_tokens
  • price_table_version, model, provider
  • event_kind, tool, args_hash (hash, not raw arguments)
  • decision: admit | reject | degrade | kill | handoff
  • reason_code: HARD_CAP_USD, HARD_CAP_TOKENS, BURN_RATE, ERROR_SPIKE, TOOL_LOOP, RETRY_STORM
  • mode_before, mode_after

Every kill or degrade also pages. Example alert policy I use as a starting point, not a vendor recommendation:

  • Page immediately: kill switch trip, per-day operator cap at 70%, reservation failures bursting (10 rejects in 5 minutes for one tenant).
  • Ticket, not page: single-run degrade, a tenant crossing 50% of daily cap, price-table miss (unknown model billed at a conservative fallback).
  • Daily digest: committed vs reserved vs provider invoice, top 10 runs by USD, top 10 args_hash repeats, count of human handoffs.

I reconcile against the provider usage API on a delay — example: every 15 minutes, plus a next-day invoice check. If committed micro-USD and the provider's number diverge by more than 8% (example drift band) over a day, that is an incident against the meter, not against the agent. Drift is how silent under-counting starts.

What I do not log in this stream: raw user text, tool payloads, API keys, or the full prompt. Those belong in a redacted trace system with a different retention policy. The budget log should be safe to keep hot for 90 days so you can answer "why did this tenant die at 14:02."

Failure modes I keep seeing

Meter after the call. You increment a counter in a finally block. One crash, one worker timeout, or one SDK-level retry later, the counter is fiction. Reserve first.

Float dollars. 0.1 + 0.2 will eventually admit a call you thought you blocked. Integer micro-USD.

One cap for everything. A generous daily tenant cap with no per-run cap is how a single loop still ruins that tenant's morning. Three scopes.

Trusting the model's token estimate. The model under-counts tools and over-counts cached prefixes it cannot see. Estimate from the provider tokenizer, pad, then correct on the usage callback.

Degrade that still exposes write tools. You changed the system prompt to "read only" and left update_customer in the schema. The next call will try it. Shrink the allowlist in the contract layer; the prompt is not the policy. That is the same lesson as the tool-contract post.

Kill switch in the agent loop. if (spend > limit) return "I should stop" is not a kill switch. The next planner turn starts again. The kill has to live in the ledger and the process supervisor.

Retry policy under the meter. Axios, the OpenAI SDK, or your queue library retries a 429 ten times. Your ledger reserved once. Disable those retries or charge them.

Shared tenant id of "default". Staging, a misconfigured gateway, or a missing auth header funnels everyone into one bucket. Either that bucket dies instantly or it never dies. Fail closed if identity is missing.

No reservation TTL. A worker dies after reserve and before commit. Capacity leaks until the next deploy. Example TTL I start with: expire unused reservations after 90 seconds and return the micro-USD to the pool.

Alerting only on the invoice. The invoice is a day late. The kill switch is a now problem. Page on reason codes, not on Stripe.

Letting the agent raise its own cap. I have seen a "budget_override" tool added as a convenience. That is a kill switch with a bypass button. Overrides are a human-only admin path, logged, with a ticket id.

Close

If the only thing standing between a production agent and an unbounded bill is a sentence in a system prompt, you do not have AI agent cost control. You have a request for good behavior.

Put LLM budget limits in a ledger the model cannot negotiate with. Reserve before every billed call. Enforce per-run, per-tenant, and per-day caps independently, in USD and in tokens. Degrade on the way up — cheaper model, fewer tools, read-only, human handoff — so you still return something useful. Kill on hard caps and on runaway signals: burn-rate spikes, error spikes, tool loops, retry storms. Log the money and the reason code. Page on kills, not on vibes.

Tool contracts keep the call well-formed. Shadow mode keeps the write from landing blind. A golden-set gate keeps a bad build off the site. The budget cap and the kill switch keep a good build from spending like a bad one. I do not ship a production agent with only the first three.

I work with teams who need this control plane before the agent touches paying traffic. If you are putting budget limits, degrade policies, or an agent kill switch around a live system, hire me.

Building something with AI?

I help teams ship production AI agents, retrieval systems, and document intelligence. Let's talk about yours.