How it works
A rate limiter counts requests per key (a user id, an API key, an IP address) in a time window and rejects the excess with HTTP 429 Too Many Requests, often with a Retry-After header. Common algorithms are fixed windows, sliding windows and the token bucket, which allows short bursts while holding the long-run average.
Counters usually live in a fast shared store such as Redis, so every server sees the same numbers; Upstash offers a ready-made library for serverless apps, and Cloudflare and API gateways can enforce limits before traffic reaches the app. Sign-in, OTP and password-reset endpoints, and anything that calls a paid AI model, need limits most.
Related terms
More in Backend and APIs
Jobs and glue