Stack Spend

How to Investigate an AI Spend Spike: An Incident Runbook

GuidesMarch 6, 2026Updated October 4, 2026By 5 min read

The short answer

Investigate an AI spend spike by confirming the account and billing window, separating request volume from model and token mix, and tracing the changed workload to its owner. If costs are still accelerating, pause or bound that workload before optimizing. Reconcile estimates against provider cost records and verify the fix after reporting catches up.

An AI cost incident needs an explanation, a containment decision, and evidence that the fix worked. This runbook gives the responsible engineer a sequence to follow and a record to share with finance.

For: engineering and platform leads responsible for a live AI workload. Use this when: someone asks why the bill jumped and you need a concrete incident decision, not another dashboard screenshot.

First response: confirm, contain, assign

  1. Confirm the provider, organization or project, currency, time zone, and affected period. Compare equal-length windows. Note whether you are reading billed costs, metered usage, or an application estimate.
  2. If a job is visibly generating unintended requests, pause the queue, disable the feature through its existing flag, or stop the worker. Bound retries, concurrency, iterations, and output length. Record the service impact of the intervention.
  3. Name one incident owner. Ask the workload owner to check deployments, routing changes, new customers, and scheduled jobs around the start of the increase.

Cost data can arrive after execution. A lower dashboard total immediately after a change does not prove containment; check that the unwanted requests have stopped in application logs and revisit provider records after they settle.

The FinOps Foundation's anomaly-management guidance connects detection with investigation and resolution. Use that loop to keep an alert from becoming an unowned observation.

Diagnosis table: what changed?

Observation Possible driver Evidence to compare Owner's next action
More requests, similar cost per request Growth, job replay, or retry amplification Requests per completed task; customer and job volume Explain expected growth or stop duplicate work
Similar requests, higher cost per request Model routing, processing mode, or token mix Model distribution; input, cached input, and output tokens Revert the change or test a bounded alternative
Tokens increased after a release Larger prompts, retrieval context, or output Usage per request and the deployed prompt/context configuration Bound context or output and evaluate quality
Application estimate does not match billed cost Missing services, tools, rate tiers, or reporting lag Provider cost categories, applicable rates, and settled periods Reconcile exclusions before claiming savings
One provider dominates the increase Local workload or provider-specific change Project, key, service, and model dimensions where available Route the investigation to the responsible team

These are hypotheses to test. A coincident deployment is a lead, not proof that the deployment caused the bill.

Worked example: separate growth from wasted work

Illustrative example. A synthetic workload completes 10,000 tasks in each comparison window. Its bill increases from $100 to $300. These are invented figures, not provider prices or customer results.

The completed-task count stayed flat, so customer growth alone does not explain the increase. Compare billable attempts per task, token mix, and model routing. If attempts rose from one to three while the request shape stayed the same, investigate retries or job replay. If attempts stayed flat, inspect routing and token usage instead.

Do not tell finance that retries explain the whole bill until the usage records and cost totals support that explanation.

Reconcile against the right provider records

For direct OpenAI usage, compare the organization's usage and cost reporting, then use request logs to trace the workload. The OpenAI high-bill guide covers cached input, paid tools, and native spend controls.

For Anthropic, Bedrock, Vertex AI, or Azure OpenAI, use the relevant provider or cloud billing records for the financial check. A token estimate can help explain activity, but it is not a substitute for an invoice. Preserve the account, time zone, and date window across comparisons.

Missing feature or customer dimensions are an attribution gap. Record what you cannot identify and add the necessary application metadata; do not infer customer-level costs from a provider total. See AI cost attribution by feature, team, and customer.

Copyable AI cost incident record

Incident owner:
Provider / organization / project:
Comparison windows and time zone:
Data source: billed cost / metered usage / application estimate
Expected change:
Observed change:
Largest changed dimension:
Workload owner and supporting evidence:
Containment action and service impact:
Verification: unwanted requests stopped at [time]
Reporting lag or unresolved reconciliation:
Permanent fix / owner / due date:
Next cost check:

Keep credentials, API keys, and customer payloads out of this record. Reference the secured evidence location instead.

Close the incident with evidence

Confirm the workload now completes with bounded execution and acceptable quality. Recheck cost and usage after the provider's reporting delay, using a comparable period. Log whether the fix reduced unnecessary work, changed model mix, or simply stopped the service. Those outcomes mean different things.

Add a budget or anomaly notification, an owner, and a follow-up review. Notifications tell a person to investigate; enforcement requires a provider or application control that changes request execution.

What to check during a StackSpend trial

Connect the provider whose bill you need to explain. Review the available spend history, sync freshness, largest reported drivers, anomaly view, and forecast. Compare the result with the provider's own records and choose who receives the daily Slack or email signal.

StackSpend reads connected cost data; it does not shut down workloads or provide a universal hard spending cap. Detail and history depend on the provider. Its OpenAI connection requires an organization Admin API key, which is not inherently a read-only credential; review the OpenAI setup guide before connecting.

Start a spend-monitoring trial, then evaluate one concrete question: can the responsible person explain the change and decide the next action? For the product workflow, see spend anomaly detection.

What to do next

Copy the incident record into your existing issue tracker, assign the investigation, and set the next provider-cost check. Review the available dimensions before treating monitoring as a substitute for application attribution.

FAQ

What should I check first in an AI spend spike?

Confirm the provider account and comparable billing window. Then compare requests, model mix, input and output tokens, and attempts per task. If a workload is still generating unintended requests, contain it before completing the financial reconciliation.

Does a cost alert stop AI requests?

An alert notifies someone; it does not stop requests. Containment requires an application change or an enforced provider control. Verify the applicable scope and behavior before relying on a limit.

Can a deployment explain a sudden increase?

A deployment can change routing, prompts, retry behavior, or workload volume. Compare its timing with the changed usage dimensions and logs. Timing alone does not establish causation.

What if I cannot identify the feature or customer?

Use the dimensions you can verify and record the missing attribution. Add application metadata or grouping for the next investigation. Do not invent feature or customer allocations from aggregate billing totals.

Can StackSpend automatically stop the workload?

No. StackSpend helps an owner monitor connected costs and investigate changes. Stopping or bounding execution requires an application intervention or an enforced provider control.

References

Give the next spend spike an owner.

Connect a provider, verify the available spend history, and configure a daily signal. Use the trial to check whether the owner can explain changes and act on them.

14-day free trial. No credit card required. Plans from $79/month.
AI Spend Spike: Investigation Runbook — StackSpend