Skip to main content
Unkey enforces per-user and per-customer rate limits, and a monthly request quota, on an API you host yourself. You don’t run Redis or any other counter store. Attach both to an API key, or call ratelimit.limit with a user id, a customer id, or any other stable string. A rate limit is a sliding window, such as 100 requests per minute. A monthly quota is credits on the key: set credits.refill.interval to monthly, and each successful keys.verifyKey spends from that balance. Credits are an exact total. Rate limit counts are close, and they can run a little high when traffic is split across regions. See Credits and refill.

Add per-user rate limits

Create the key with a per-minute limit and a monthly credit refill, then verify it on each request. externalId ties the key to one customer.
duration is the window in milliseconds. autoApply: true checks that limit on every verification. The monthly refill replaces the balance on refillDay at 00:00 UTC. It doesn’t add to what’s left. For an endpoint with no API key, call ratelimit.limit with your own identifier instead. The same pattern is on ratelimit.limit and multiLimit.

What a limited request looks like

keys.verifyKey and ratelimit.limit return HTTP 200 when your user is over their limit. Read the body. keys.verifyKey sets data.valid to false and data.code to RATE_LIMITED (or USAGE_EXCEEDED when credits run out). Each checked limit is in data.ratelimits, with limit, remaining, reset, and exceeded. ratelimit.limit sets data.success to false and puts limit, remaining, and reset on data. reset is Unix milliseconds, when the current window ends. Your server is the one that returns 429 Too Many Requests to the caller. RFC 6585 defines that status and says the response may include Retry-After, a number of seconds to wait. Set it from reset, and use at least 1 second. Unkey doesn’t send Retry-After or X-RateLimit-* on these two endpoints for your user’s limit, so copy limit, remaining, and reset onto your own response if you want those headers. X-RateLimit-Reset on the Compute gateway is Unix seconds, so divide reset by 1000 if you want to match it. If the API runs on Compute, the gateway returns 429 itself. The response has Retry-After (seconds, never under 1), X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset (Unix seconds). See rate_limited. Your server’s own calls to api.unkey.com have a separate workspace limit. If you exceed it, Unkey returns HTTP 429 with RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset (seconds until the window ends), and Retry-After. Those headers describe your workspace’s budget for the Unkey API, not your users’ limits. in API Management has two forms. A standalone API takes any identifier you choose. Limits on a key or identity are checked when the key is verified. The Compute gateway’s rate limit policy is separate. See the glossary.

Choose a rate limit

Use the standalone API when the endpoint has no key (sign-up, password reset) or when you want to limit something other than the key holder. Use key and identity limits when the request already carries a key, so the check happens during the verification you already make. Both count requests the same way. See How rate limiting works.

A first rate limit

You need a root key with the permissions listed on this page. Create one in the dashboard under Settings > Root Keys. See Permission reference for every permission.
The root key needs ratelimit.*.limit, plus ratelimit.*.create_namespace the first time a namespace name is used.
The namespace email.send names what you’re limiting. The identifier user_123 is who’s being counted. Each user gets ten calls per minute, and the eleventh returns success: false until the window moves on. In the TypeScript SDK, the same call is unkey.ratelimit.limit({ namespace, identifier, limit, duration }).

Next steps

How rate limiting works

Sliding windows, regional counters, cross-region convergence, and what remaining and reset mean.

ratelimit.limit and multiLimit

Request and response fields, bounds, namespaces, cost, and atomic multi-limit checks.

Rate limit overrides

Give one identifier or a wildcard pattern a different limit without changing code.

Key and identity rate limits

Named limits on keys and identities, auto-apply, cost, and inline limits at verification.
Last modified on October 11, 2026