> Bron: https://neuralex.nl/en/blog/rate-limits-en-quota
> Rate limits and quotas can slow down or interrupt AI services without warning. Retries, queues, model routing, budget guards and monitoring together turn temporary scarcity into controlled degradation instead of an outage.

[Back to Insights](/en/blog)

Infrastructure ·12 September 2026 ·6 min read

# Rate limits and quotas: how do you stop your service from grinding to a halt?

Rate limits and quotas can slow down or interrupt AI services without warning. Retries, queues, model routing, budget guards and monitoring together turn temporary scarcity into controlled degradation instead of an outage.

An AI service remains reliable only if you assume that external API capacity is limited. LLM providers impose restrictions on request volume, token throughput or total usage over a given period. Once a limit is reached, an API request may be rejected, for example with an HTTP 429 response. If your application is not designed for this behaviour, a temporary capacity constraint immediately becomes an outage for users.

The solution is therefore not simply to request a higher quota. A robust architecture combines controlled retries, exponential backoff, queues, batching, model routing, fallback providers, budget guards and monitoring. These measures do not remove every limit, but they help turn temporary capacity shortages into delay or graceful degradation instead of complete failure.

## What are rate limits and quotas?

A rate limit determines how much work you may send to a service within a given period. A provider might limit requests per minute, tokens per minute or a combination of both.

A quota is broader. It may define the maximum usage per day, month, project, account or billing period.

The distinction matters operationally. When you hit a rate limit, capacity may still be available, but you are consuming it too quickly — waiting and retrying later may solve the problem. When a quota is exhausted, retries normally do not help until the quota resets or is increased.

AI APIs may enforce several limits simultaneously. An application can remain comfortably below its request limit while still exceeding a token limit because a small number of prompts or responses are unusually large. Systems should therefore measure not only the number of API calls but also the amount of work represented by each call.

## Provider-side limits: tokens per minute and requests per minute

LLM providers use rate limits to protect infrastructure and distribute capacity across customers. Two common mechanisms are requests per minute and tokens per minute.

Requests per minute is straightforward: too many calls within a short period trigger throttling. Tokens per minute is more subtle — a few large prompts can consume as much capacity as many smaller requests. Long contexts, large RAG payloads or extensive generated responses can therefore reach a limit even when request volume appears modest. An application that monitors API call counts alone will detect this problem too late.

Provider-side limits may also vary by model, account, region or service tier. They should therefore be treated as configurable infrastructure constraints rather than constants hard-coded throughout the application. HTTP 429 responses should also be treated as normal operational signals, not necessarily as a broken system — they tell the client that it must temporarily reduce the rate at which requests are sent.

## Your own internal limits matter just as much

A provider may define the technical maximum, but your application should usually impose limits before that point. Imagine a hundred users starting computationally heavy analyses at almost the same time: even if the provider accepts every request, the result may still be high costs, increased latency and unstable performance. Internal rate limits therefore protect both system capacity and budget.

Limits can be applied per user, organisation, API key, feature or workload type. An interactive chat request might receive priority over a background process that is summarising thousands of documents. Concurrency limits are also useful: instead of starting an unlimited number of model requests in parallel, you define how many tasks may run simultaneously. Remaining work is placed in a queue. This makes load more predictable and prevents short traffic spikes from destabilising the entire application.

## Exponential backoff and retries

A retry is often the first response to a temporary rate limit, but blindly resending requests can make the situation worse. Exponential backoff is therefore commonly used: after each failed attempt, the application waits longer before trying again. A generic example might be waiting one second, then two seconds, then four seconds.

In practice, jitter is often added as well — a small random variation in waiting time. Without jitter, hundreds of clients that receive a 429 at the same moment may all retry at exactly the same time, creating another traffic spike. Retries should also have a limit: a system that continues retrying indefinitely may fill queues and consume unnecessary resources. Not every error should be retried either. A temporary 429 or network failure may justify a retry; a structurally invalid request usually does not.

## Queueing and batching

Queues are one of the main tools for absorbing bursts in workload. Instead of sending every task directly to a model provider, the application places tasks in a queue. Workers then process them at a rate that stays within available capacity. This works particularly well for non-interactive jobs such as document processing, embeddings, classification or bulk summarisation. User-facing workloads require tighter latency limits, so separate priority queues may be appropriate.

Batching can also reduce pressure on rate limits. When several small operations can be combined into a single request, the number of API calls decreases. Batching is only useful, however, when the provider and model support it efficiently and when the resulting delay remains acceptable.

## Model routing and fallback chains

Not every task needs to use the same model. With [model routing](/en/blog/model-routing-uitgelegd), simple workloads can be sent to smaller models while more complex tasks are reserved for more capable ones. This spreads demand and prevents scarce capacity from being consumed by work that does not require it.

A [fallback chain](/en/blog/fallback-ketens) extends this concept: if the primary model is temporarily unavailable or rate-limited, the application may route the request to a secondary model or another provider. This requires careful handling because providers and models differ in behaviour, context limits, tool support, output formats and pricing. A fallback therefore does not automatically provide equivalent quality — in some cases, a clear error or delayed response is preferable to silently switching to a model that is not suitable for the task.

## Budget guards

Technical availability and financial availability are closely related. An application may never hit a provider's rate limit and still become impractical if it allows unrestricted use of expensive requests — see also [what AI really costs](/en/blog/wat-kost-ai-echt). Budget guards introduce financial constraints alongside technical limits.

They can be applied per user, project, day or workload. Maximum output lengths, context sizes and the number of model calls permitted within a workflow can also be restricted. This is particularly important for AI agents: a single user instruction may trigger multiple follow-up actions, producing far more API traffic than the original request suggests. A spending ceiling or maximum number of agent steps prevents an accidental loop from consuming resources indefinitely.

## Monitoring and alerting

Rate-limit problems are often noticed only after users begin reporting failures. By then, the system is already in trouble. Useful metrics include the number of HTTP 429 responses, retry rates, queue depth, token consumption, concurrent requests, latency and the percentage of traffic handled by fallback models or providers.

Trends matter as much as individual incidents. If usage is steadily approaching a provider limit, you want to see that before the limit is reached. Alerts should therefore trigger not only during outages but also when available headroom begins to shrink. Larger systems also benefit from breaking usage down by model, user, tenant or workload — otherwise, you may know that capacity is being consumed without knowing which part of the application is responsible.

## Conclusion

Rate limits and quotas are not unusual edge cases. They are a normal part of operating services that depend on AI APIs. A production system should assume that external capacity will occasionally become constrained. The most effective response is not a single retry rule but a combination of mechanisms: internal limits, exponential backoff, queues, batching, model routing, fallback chains, budget guards and monitoring.

A well-designed system does not attempt to avoid every capacity limit. It makes sure that reaching one is handled predictably. Temporary scarcity then results in delay or reduced functionality rather than the entire service grinding to a halt.

Reliable AI infrastructure

## Want an AI service that keeps running smoothly under peak load?

We design rate limiting, retries, queues and fallback strategies that fit your architecture and budget. Curious what that would mean for your system?

[Ask your question](/en/contact) [Read what AI really costs](/en/blog/wat-kost-ai-echt)

---
Volledige (opgemaakte) versie: https://neuralex.nl/en/blog/rate-limits-en-quota
