Rate limits
Two different limits, on two different things. Requests made with a key draw on that key’s daily quota for the endpoint. Requests for a new key are throttled per IP address, because that endpoint has no key yet.
The limits, as the database enforces them
Daily and monthly allowances are rows in the plan_limit table, resolved at request time. This table is that table: it is queried on every render, so the number published here and the number enforced cannot drift apart.
| Plan | Endpoint | Per day | Per month |
|---|---|---|---|
| free | all endpoints | 1,000 | 20,000 |
| free | speech_to_text | 20 | 400 |
| free | translate | 50 | 1,000 |
| institution | all endpoints | 500,000 | unlimited |
| team | all endpoints | 50,000 | 1,000,000 |
| team | speech_to_text | 1,000 | 20,000 |
| team | translate | 5,000 | 100,000 |
Some rows name a metering key that no page in this reference documents — translate and speech_to_text. They are in the table because the plans define them, not because the public v1 surface has an endpoint for them.
How a limit is resolved
A request is metered under the endpoint key its route was wrapped with — words for search, entry lookup and word-of-the-day; languages; stats. The limit applied is the row for that exact (plan, endpoint) pair if one exists, and otherwise the plan’s * row.
If neither row exists, the request is refused rather than allowed: an installation whose plan table is misconfigured fails closed instead of serving unmetered traffic. The refusal is still reported as quota_exceeded.
Counters are per developer, per endpoint, per UTC day. An over-limit request is still counted, so an abusive client stays visible in the usage data instead of disappearing the moment it crosses the line. The quota resets at 00:00 UTC.
The headers that report it
Every metered response carries the day’s state, so a client can slow itself down without guessing:
| Header | Meaning |
|---|---|
X-RateLimit-Limit | The daily limit the plan resolves to for this endpoint. |
X-RateLimit-Remaining | Requests left today, floored at zero. |
X-RateLimit-Used | Requests counted today, after this one. |
Retry-After | Seconds until 00:00 UTC. Set on a metered route only when the response is quota_exceeded, and on the key-registration route as 900 when it is rate_limited. |
The headers are attached to successful responses and to failures alike. The OpenAPI document declares X-RateLimit-Limit and X-RateLimit-Remaining on the search operation; the third and the Retry-After come from the shared pipeline that every metered route runs through.
Creating a key: 5 attempts per 15 minutes, per IP
POST /api/v1/developers is unauthenticated and therefore cannot be metered by key. It is throttled by client IP instead: at most 5 attempts in a fixed 15-minute window. The window opens with the first attempt and is not extended by later ones.
Past the limit the route answers HTTP 429 with the code rate_limited and Retry-After: 900 — the full window, in seconds:
{
"error": {
"code": "rate_limited",
"message": "Too many key requests from this address. Try again in 15 minutes."
}
}The counter is held in process memory, so it protects a single instance. An installation running more than one container is expected to put a shared limiter or a WAF rate rule in front of the route; the source says as much in the comment above the counter.
Plans
The installation seeds three plans — free, team and institution — and the table above is the current truth for every one of them. The form on the landing page registers a developer on the free plan: the route does not send a plan, and registration defaults to free.
When a limit is hit
A metered request over its allowance returns HTTP 429 with the code quota_exceeded, a Retry-After in seconds, and a details object naming the plan, the endpoint, the used count and the limit. The full body is on the Errors page.