⏱️ Reading time: 15 min
On October 3, 2026, Simon Willison published an essay with an uncomfortable idea: almost no pay-per-use service stops spending when it spikes on its own. An alert arrives at three in the morning, but the bill keeps growing while the account owner sleeps. Willison isn’t asking for a better notification: he’s asking for a hard spending limit enabled by default.
📑 En este artículo
- TL;DR
- What Is a Hard Spending Limit?
- Why a Hard Billing Cap Matters
- How an Automatic Spending Cutoff Works
- Defense in Depth: Alerts and Hard Limits Together
- How to Enable Hard Spending Limits on Each Service
- Real-World Use Cases
- Comparison: Alerts, Budgets, and Spend Caps
- Common Mistakes When Setting Up a Spending Cap
- Deep Dive: Why a Usage Counter Doesn’t Always Cut Off in Time
- How to Verify the Spending Cap Is Active
- Frequently Asked Questions
- Does a Hard Spending Limit Cut Off Service Instantly?
- Does OpenAI Let You Set a Hard Spending Limit?
- What’s the Difference Between Spend Caps and a Classic Google Cloud Budget?
- Does Anthropic Offer Hard Spending Limits on the Claude API?
- What Happens to a Request Already in Progress When AWS or Google Cloud Hits the Billing Cap?
- Should You Always Use a Hard Spending Limit in Production?
- References
AWS and Google Cloud have already started offering this feature natively. This article explains what a hard spending limit is, how it differs from a billing alert, and how to enable it today on AWS, Google Cloud, OpenAI, and Anthropic.
TL;DR
- A hard spending limit cuts off API access the moment the usage counter hits the set cap.
- AWS launched a per-project spending limit on September 16, 2026, that pauses the service when it hits the monthly cap.
- Google Cloud has offered Spend Caps since July 2026 to set a financial cap per service within a project.
- OpenAI and Anthropic let you set a hard spending limit from each account’s billing dashboard.
- Relying only on an email alert leaves the bill open: a cutoff via HTTP error closes it.
What Is a Hard Spending Limit?
A hard spending limit is a billing rule that cuts off access to an API or cloud service once it reaches a cap set in advance. It doesn’t depend on an alert being read in time: the provider rejects every new request until the account owner lifts it.
Its soft counterpart is the classic billing alert: an email or a webhook that warns a threshold was crossed, but doesn’t interrupt anything on its own. Most clouds started out with only that option, and it’s still the default in several dashboards.
Why a Hard Billing Cap Matters
Coding agents and personal agents lower the friction for spinning up something that spends real money: calls to paid APIs, hosted applications, systems that bill for extra storage and compute. If that agent gets stuck in a loop (retrying a failed call, for example) nobody finds out until the bill arrives.
A typical case: someone connects a coding agent tool to their own cloud account so it can test real infrastructure, the agent misinterprets an instruction and starts creating instances or calling an expensive model in a loop. Without a hard spending limit, that error only shows up when the billing summary arrives at the end of the month.
Willison sums it up with a simple question: would you rather see errors in the middle of the night, or a bill for several thousand dollars the next morning? His proposal is that the hard cutoff be the default option on every pay-per-use service, with an explicit checkbox for anyone who’d rather assume the risk: something like remove the spending cap, my application won’t shut down if I exceed the configured limit and I accept the resulting charges.
💭 Key takeaway: Willison’s argument isn’t technical, it’s about default design: the safe option should be the one that protects the user’s wallet, not the one that keeps the service running at all costs.
How an Automatic Spending Cutoff Works
There are two ways to implement it, and the difference matters. The first keeps a usage counter that’s checked on every request, before it’s processed: that’s how OpenAI, Anthropic, and, according to Willison, Google Cloud’s new Spend Caps work. The cutoff is nearly instant because it happens within the same request.
The second depends on billing data that gets consolidated with a delay, typical of a classic AWS or Google Cloud budget: the system checks accumulated spending at intervals and, if someone configured it that way, triggers an external function (a Lambda, a Cloud Function) that only then revokes access or disables billing. This second model is never instant, and during the window between crossing the threshold and the function running, spending keeps piling up.
Another nuance: not every counter measures the same thing. OpenAI and Anthropic measure consumption in estimated dollars based on input and output tokens, updated as each response finishes generating. AWS and Google Cloud measure actual accumulated spending across all services on the account, not just AI, so a limit there can cut off compute, storage, and network all at once, not just calls to a model.
flowchart TD
A["Service usage grows"] --> B{"Type of control"}
B -->|"Soft alert"| C["Warning email or webhook"]
C --> D["Service keeps responding"]
D --> A
B -->|"Hard limit"| E["Internal counter hits the cap"]
E --> F["Subsequent requests return an error"]
F --> G["Spending stops growing"]
The diagram sums up the difference: the left branch loops back on itself, spending keeps going while the alert just warns, and the right branch cuts the cycle as soon as the cap is hit.
Defense in Depth: Alerts and Hard Limits Together
A hard spending limit doesn’t replace the early alert, it complements it. The design that comes up most often among teams that already got burned once has two layers: a soft alert at 50% of the monthly budget, so someone checks whether the usage is legitimate, and a hard spending limit at 100%, as a last line of defense if nobody reacts to the first signal.
The reason not to skip the alert and rely only on the cutoff is that a poorly calibrated cap can also take down production without warning. The 50% alert gives enough room to raise the cap in time if usage grew for a legitimate reason (more users, a campaign, a migration), without ever letting the hard cutoff trigger by surprise on a service that actually needed that spending.
How to Enable Hard Spending Limits on Each Service
None of the four providers uses the same name or the same flow, so it’s worth reviewing them separately. In all four cases, the limit is set per account, project, or workspace, not per individual request.
AWS: Budget Actions and the New Per-Project Spending Limit
The classic mechanism, available on any account, is AWS Budgets. A simple budget only warns, but you can add an action (Budget Action) that applies a restrictive IAM policy when the threshold is crossed. To create the base budget from the CLI:
aws budgets create-budget \
--account-id 123456789012 \
--budget '{"BudgetName":"limite-api-ia","BudgetLimit":{"Amount":"100","Unit":"USD"},"TimeUnit":"MONTHLY","BudgetType":"COST"}' \
--notifications-with-subscribers '[{"Notification":{"NotificationType":"ACTUAL","ComparisonOperator":"GREATER_THAN","Threshold":100},"Subscribers":[{"SubscriptionType":"EMAIL","Address":"[email protected]"}]}]'
The command doesn’t return output if it runs successfully. To confirm the budget was created:
aws budgets describe-budgets --account-id 123456789012
The response should include an object with "BudgetName": "limite-api-ia" and the configured amount. This alone is still just an alert: to make it actually cut off access you need to attach a Budget Action of type APPLY_IAM_POLICY, documented in the same AWS Budgets guide. In September 2026, AWS also added a per-project spending limit in a new account experience that pauses the entire project when it hits the monthly cap without manually building the policy, although as Willison reports it’s still in limited rollout for a small group of accounts.
Google Cloud: Classic Budgets and Spend Caps
A classic Google Cloud budget only notifies. It’s created like this (replacing the billing ID with your own account’s, visible under Billing > Account Management):
gcloud billing budgets create \
--billing-account=012345-6789AB-CDEF01 \
--display-name="limite-vertex-ai" \
--budget-amount=100USD \
--threshold-rule=percent=1.0
For that budget to actually cut anything off, you need to go a step further: connect it to a Pub/Sub topic and a Cloud Function that disables billing for the project, a pattern Google documents in its programmatic budget notifications guide. It’s the same trick thousands of developers used before a native option existed. Since July 2026, Google Cloud offers Spend Caps, which sets a direct financial cap on a specific service within a project and pauses it when the limit is reached, with no intermediate Cloud Function; for now it’s only configurable from the console, with no equivalent gcloud command.
OpenAI: Hard Spending Cap in the Dashboard
On an OpenAI account, the cap is configured at platform.openai.com, inside the organization’s billing section. There you can set both a warning (soft limit) and a cutoff (hard limit): once the latter is reached, subsequent API calls return a quota exceeded error instead of being processed, the same type of error described in OpenAI’s usage limits guide.
Anthropic: Spend Limits per Workspace
At console.anthropic.com, the plans and billing section lets you set a monthly spending cap per workspace or per API key. Once it’s reached, subsequent requests to the Claude API are rejected until the next billing cycle or until someone with permissions manually raises the limit.
Real-World Use Cases
A developer testing a coding agent with access to their own card is the most obvious case: a poorly handled retry loop can turn an afternoon of testing into a four-figure bill. Setting a $20 or $50 hard spending limit on the test account makes that error harmless.
On a team, the same mechanism protects against an API key leaked in a public repository: while the provider’s warning about the leak is still on its way, a hard spending limit has already cut off the abuse. It also helps in CI pipelines that call a language model on every run: if a change triggers thousands of calls because of a bug in the retry logic, the limit keeps the error from only being noticed on the next bill.
It also protects learners: many clouds hand out free credits for a limited time, and a hard spending limit set right at the credit amount keeps a test from turning into the first real charge on the card.
Comparison: Alerts, Budgets, and Spend Caps
| Service | Type of Control | What Happens at the Cap | Availability |
|---|---|---|---|
| AWS, per-project spending limit | Automatic cutoff at the project level | The project pauses for that month | Limited rollout since September 2026 |
| AWS Budgets + Budget Actions | Alert plus a scheduled action (IAM) | Denies access if the action was configured | Available on any AWS account |
| Google Cloud Spend Caps | Financial cap per service | The service pauses once it hits the cap | Limited rollout since July 2026 |
| Google Cloud Classic Budgets | Alert only (email or Pub/Sub) | Doesn’t cut off anything on its own | Available on every project |
| OpenAI Usage Limits | Hard spending cap in the dashboard | Subsequent calls return a quota error | Available for paid accounts |
| Anthropic Console Spend Limits | Cap per workspace or key | Subsequent calls are rejected | Available for Claude API accounts |
Common Mistakes When Setting Up a Spending Cap
- Confusing an alert with a cutoff: a classic Google Cloud or AWS budget with no connected action only sends an email, it never stops anything.
- Setting the cap too low: a limit calculated from a quiet day’s traffic takes down production on the first real usage spike.
- Sharing the limit across environments: if a single billing account covers both development and production, a bug in staging can bring down the production API.
- Assuming the cutoff is instant: in the classic-budget-plus-Cloud-Function model, minutes pass between crossing the threshold and the function actually running.
- Not testing the limit before trusting it: deliberately creating a low cap ($1 or $5) and generating test traffic until you confirm it actually cuts off access is the only way to know it works.
⚠️ Heads up: a poorly calibrated hard spending limit is, in practice, a self-inflicted service outage. Test it in a staging environment before applying it to production.
Deep Dive: Why a Usage Counter Doesn’t Always Cut Off in Time
The real-time counter model (OpenAI, Anthropic, Spend Caps) has its own limitation: if several requests arrive in parallel right before the counter registers the first one’s spending, more than one might pass the check at the same time, and actual spending ends up slightly above the nominal cap. It’s the same race condition problem that exists in any distributed rate-limiting system: the check and the counter update aren’t a single atomic operation unless the provider explicitly guarantees it.
The budget-plus-external-function model has the opposite problem, and a worse one: the delay isn’t milliseconds but minutes, because it depends on billing data getting consolidated and on the function (Lambda or Cloud Function) actually triggering and having permission to act. The first type’s cutoff is more reliable for avoiding surprises; the second requires more correctly connected pieces to work as expected.
Granularity matters too: a limit set at the organization level in OpenAI or Anthropic cuts off every API key at once, including the production one, if a test project eats up the shared budget. Splitting keys by project or team, each with its own limit, keeps an isolated experiment from taking down an unrelated service.
sequenceDiagram
participant Client
participant API as AI API
participant Counter as Usage Counter
Client->>API: requests completion
API->>Counter: checks accumulated spend
alt spend below cap
Counter-->>API: within limit
API-->>Client: 200 response with result
else spend reached cap
Counter-->>API: cap reached
API-->>Client: quota exceeded error
end
How to Verify the Spending Cap Is Active
On AWS, aws budgets describe-budget-actions-for-budget --account-id 123456789012 --budget-name limite-api-ia shows whether there’s an associated Budget Action and its status. On Google Cloud, gcloud billing budgets list --billing-account=012345-6789AB-CDEF01 lists the budgets configured for that billing account, though it doesn’t distinguish whether they’re connected to an actual cutoff function: that has to be confirmed by separately checking the Pub/Sub topic and the associated Cloud Function.
On OpenAI and Anthropic there’s no direct way to verify the cap from outside the account: they don’t expose a public endpoint to check it. The only reliable way to confirm it is to generate test traffic on a sandbox account until you exceed a deliberately low cap and verify that the API actually returns an error instead of processing the request.
Your next step: create a $1 budget today on your AWS or Google Cloud account using the commands in this article, and confirm the alert arrives before replicating the setup with the real production amount.
Frequently Asked Questions
Does a Hard Spending Limit Cut Off Service Instantly?
It depends on the mechanism. If the provider checks a counter on every request (OpenAI, Anthropic, Google Cloud Spend Caps), the cutoff is nearly immediate. If it relies on a classic budget plus an external function, minutes can pass.
Does OpenAI Let You Set a Hard Spending Limit?
Yes, from the organization’s billing dashboard at platform.openai.com you can set a hard limit that blocks new API calls once it’s reached.
What’s the Difference Between Spend Caps and a Classic Google Cloud Budget?
A classic budget only sends an alert unless it’s manually connected to a Cloud Function that disables billing. Spend Caps, available since July 2026, pauses the service directly once the cap is reached, with no intermediate steps.
Does Anthropic Offer Hard Spending Limits on the Claude API?
Yes, from console.anthropic.com you can set a monthly spending cap per workspace or API key, and requests are rejected once it’s exceeded.
What Happens to a Request Already in Progress When AWS or Google Cloud Hits the Billing Cap?
Across all four providers described here, the cutoff applies to new requests. A call that already started processing normally completes, and that cost also counts toward the next cycle.
Should You Always Use a Hard Spending Limit in Production?
It’s a good idea on any development, testing, or personal project account. In production it needs to be calibrated against expected real traffic, because a poorly calculated cutoff interrupts service just as effectively as an attack.
References
- Simon Willison: the original essay proposing hard spending limits as the default option.
- AWS Budgets: official documentation for creating budgets and Budget Actions.
- Google Cloud: official guide on programmatic budget notifications and the pattern for disabling billing via a function.
- OpenAI: usage limits guide and quota error handling.
- Anthropic Console: billing dashboard where Claude API spend limits are configured.
📱 Enjoying this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.
Featured image: Foto de imgix en Unsplash
Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.
Leave a comment
0 Comments