Skip to content
Menu
Guide

Rate limits

Two different limits, on two different things. Requests made with a key draw on that key’s daily quota for the endpoint. Requests for a new key are throttled per IP address, because that endpoint has no key yet.

The limits, as the database enforces them

Daily and monthly allowances are rows in the plan_limit table, resolved at request time. This table is that table: it is queried on every render, so the number published here and the number enforced cannot drift apart.

Every plan and endpoint row currently registered.
PlanEndpointPer dayPer month
freeall endpoints1,00020,000
freespeech_to_text20400
freetranslate501,000
institutionall endpoints500,000unlimited
teamall endpoints50,0001,000,000
teamspeech_to_text1,00020,000
teamtranslate5,000100,000

Some rows name a metering key that no page in this reference documents — translate and speech_to_text. They are in the table because the plans define them, not because the public v1 surface has an endpoint for them.

How a limit is resolved

A request is metered under the endpoint key its route was wrapped with — words for search, entry lookup and word-of-the-day; languages; stats. The limit applied is the row for that exact (plan, endpoint) pair if one exists, and otherwise the plan’s * row.

If neither row exists, the request is refused rather than allowed: an installation whose plan table is misconfigured fails closed instead of serving unmetered traffic. The refusal is still reported as quota_exceeded.

Counters are per developer, per endpoint, per UTC day. An over-limit request is still counted, so an abusive client stays visible in the usage data instead of disappearing the moment it crosses the line. The quota resets at 00:00 UTC.

The headers that report it

Every metered response carries the day’s state, so a client can slow itself down without guessing:

HeaderMeaning
X-RateLimit-LimitThe daily limit the plan resolves to for this endpoint.
X-RateLimit-RemainingRequests left today, floored at zero.
X-RateLimit-UsedRequests counted today, after this one.
Retry-AfterSeconds until 00:00 UTC. Set on a metered route only when the response is quota_exceeded, and on the key-registration route as 900 when it is rate_limited.

The headers are attached to successful responses and to failures alike. The OpenAPI document declares X-RateLimit-Limit and X-RateLimit-Remaining on the search operation; the third and the Retry-After come from the shared pipeline that every metered route runs through.

Creating a key: 5 attempts per 15 minutes, per IP

POST /api/v1/developers is unauthenticated and therefore cannot be metered by key. It is throttled by client IP instead: at most 5 attempts in a fixed 15-minute window. The window opens with the first attempt and is not extended by later ones.

Past the limit the route answers HTTP 429 with the code rate_limited and Retry-After: 900 — the full window, in seconds:

{
  "error": {
    "code": "rate_limited",
    "message": "Too many key requests from this address. Try again in 15 minutes."
  }
}

The counter is held in process memory, so it protects a single instance. An installation running more than one container is expected to put a shared limiter or a WAF rate rule in front of the route; the source says as much in the comment above the counter.

Plans

The installation seeds three plans — free, team and institution — and the table above is the current truth for every one of them. The form on the landing page registers a developer on the free plan: the route does not send a plan, and registration defaults to free.

See what the free plan allows →

When a limit is hit

A metered request over its allowance returns HTTP 429 with the code quota_exceeded, a Retry-After in seconds, and a details object naming the plan, the endpoint, the used count and the limit. The full body is on the Errors page.