An AI cost incident needs an explanation, a containment decision, and evidence that the fix worked. This runbook gives the responsible engineer a sequence to follow and a record to share with finance.
For: engineering and platform leads responsible for a live AI workload. Use this when: someone asks why the bill jumped and you need a concrete incident decision, not another dashboard screenshot.
First response: confirm, contain, assign
- Confirm the provider, organization or project, currency, time zone, and affected period. Compare equal-length windows. Note whether you are reading billed costs, metered usage, or an application estimate.
- If a job is visibly generating unintended requests, pause the queue, disable the feature through its existing flag, or stop the worker. Bound retries, concurrency, iterations, and output length. Record the service impact of the intervention.
- Name one incident owner. Ask the workload owner to check deployments, routing changes, new customers, and scheduled jobs around the start of the increase.
Cost data can arrive after execution. A lower dashboard total immediately after a change does not prove containment; check that the unwanted requests have stopped in application logs and revisit provider records after they settle.
The FinOps Foundation's anomaly-management guidance connects detection with investigation and resolution. Use that loop to keep an alert from becoming an unowned observation.
Diagnosis table: what changed?
| Observation | Possible driver | Evidence to compare | Owner's next action |
|---|---|---|---|
| More requests, similar cost per request | Growth, job replay, or retry amplification | Requests per completed task; customer and job volume | Explain expected growth or stop duplicate work |
| Similar requests, higher cost per request | Model routing, processing mode, or token mix | Model distribution; input, cached input, and output tokens | Revert the change or test a bounded alternative |
| Tokens increased after a release | Larger prompts, retrieval context, or output | Usage per request and the deployed prompt/context configuration | Bound context or output and evaluate quality |
| Application estimate does not match billed cost | Missing services, tools, rate tiers, or reporting lag | Provider cost categories, applicable rates, and settled periods | Reconcile exclusions before claiming savings |
| One provider dominates the increase | Local workload or provider-specific change | Project, key, service, and model dimensions where available | Route the investigation to the responsible team |
These are hypotheses to test. A coincident deployment is a lead, not proof that the deployment caused the bill.
Worked example: separate growth from wasted work
Illustrative example. A synthetic workload completes 10,000 tasks in each comparison window. Its bill increases from $100 to $300. These are invented figures, not provider prices or customer results.
The completed-task count stayed flat, so customer growth alone does not explain the increase. Compare billable attempts per task, token mix, and model routing. If attempts rose from one to three while the request shape stayed the same, investigate retries or job replay. If attempts stayed flat, inspect routing and token usage instead.
Do not tell finance that retries explain the whole bill until the usage records and cost totals support that explanation.
Reconcile against the right provider records
For direct OpenAI usage, compare the organization's usage and cost reporting, then use request logs to trace the workload. The OpenAI high-bill guide covers cached input, paid tools, and native spend controls.
For Anthropic, Bedrock, Vertex AI, or Azure OpenAI, use the relevant provider or cloud billing records for the financial check. A token estimate can help explain activity, but it is not a substitute for an invoice. Preserve the account, time zone, and date window across comparisons.
Missing feature or customer dimensions are an attribution gap. Record what you cannot identify and add the necessary application metadata; do not infer customer-level costs from a provider total. See AI cost attribution by feature, team, and customer.
Copyable AI cost incident record
Incident owner:
Provider / organization / project:
Comparison windows and time zone:
Data source: billed cost / metered usage / application estimate
Expected change:
Observed change:
Largest changed dimension:
Workload owner and supporting evidence:
Containment action and service impact:
Verification: unwanted requests stopped at [time]
Reporting lag or unresolved reconciliation:
Permanent fix / owner / due date:
Next cost check:
Keep credentials, API keys, and customer payloads out of this record. Reference the secured evidence location instead.
Close the incident with evidence
Confirm the workload now completes with bounded execution and acceptable quality. Recheck cost and usage after the provider's reporting delay, using a comparable period. Log whether the fix reduced unnecessary work, changed model mix, or simply stopped the service. Those outcomes mean different things.
Add a budget or anomaly notification, an owner, and a follow-up review. Notifications tell a person to investigate; enforcement requires a provider or application control that changes request execution.
What to check during a StackSpend trial
Connect the provider whose bill you need to explain. Review the available spend history, sync freshness, largest reported drivers, anomaly view, and forecast. Compare the result with the provider's own records and choose who receives the daily Slack or email signal.
StackSpend reads connected cost data; it does not shut down workloads or provide a universal hard spending cap. Detail and history depend on the provider. Its OpenAI connection requires an organization Admin API key, which is not inherently a read-only credential; review the OpenAI setup guide before connecting.
Start a spend-monitoring trial, then evaluate one concrete question: can the responsible person explain the change and decide the next action? For the product workflow, see spend anomaly detection.
What to do next
Copy the incident record into your existing issue tracker, assign the investigation, and set the next provider-cost check. Review the available dimensions before treating monitoring as a substitute for application attribution.
FAQ
What should I check first in an AI spend spike?
Confirm the provider account and comparable billing window. Then compare requests, model mix, input and output tokens, and attempts per task. If a workload is still generating unintended requests, contain it before completing the financial reconciliation.
Does a cost alert stop AI requests?
An alert notifies someone; it does not stop requests. Containment requires an application change or an enforced provider control. Verify the applicable scope and behavior before relying on a limit.
Can a deployment explain a sudden increase?
A deployment can change routing, prompts, retry behavior, or workload volume. Compare its timing with the changed usage dimensions and logs. Timing alone does not establish causation.
What if I cannot identify the feature or customer?
Use the dimensions you can verify and record the missing attribution. Add application metadata or grouping for the next investigation. Do not invent feature or customer allocations from aggregate billing totals.
Can StackSpend automatically stop the workload?
No. StackSpend helps an owner monitor connected costs and investigate changes. Stopping or bounding execution requires an application intervention or an enforced provider control.

