> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.talqora.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.talqora.com/_mcp/server.

# Rate limits

> Understand organization-wide request budgets, response headers, and safe retry behavior.

Talqora enforces request rate limits on the server after authentication and before expensive vector-storage or sparse-search work begins. Limits use a rolling 60-second window and are aggregated by organization, not by API key.

## Limits by plan

| Plan       | Control plane | Writes and deletes |    Queries |
| ---------- | ------------: | -----------------: | ---------: |
| Developer  |       120/min |             30/min |     60/min |
| Scale      |     3,000/min |          1,500/min |  3,000/min |
| Enterprise |    20,000/min |         10,000/min | 20,000/min |

The buckets are independent:

* **Control plane** includes authenticated organization, index, API-key, usage, member, and billing operations.
* **Writes and deletes** share one bucket across every index and API key in the organization.
* **Queries** share a separate bucket across dense, sparse, and hybrid retrieval.

Creating five API keys does not provide five times the request capacity. This prevents accidental key fan-out from bypassing an organization's plan. A batch containing 500 vectors consumes one write request, although all 500 rows count toward monthly usage.

## Response headers

Successful rate-limited requests return:

```http
RateLimit-Limit: 600
RateLimit-Remaining: 594
RateLimit-Reset: 41
```

| Header                | Meaning                                                     |
| --------------------- | ----------------------------------------------------------- |
| `RateLimit-Limit`     | Maximum requests permitted in the active 60-second window.  |
| `RateLimit-Remaining` | Requests still available in that bucket.                    |
| `RateLimit-Reset`     | Seconds until the oldest request leaves the rolling window. |
| `Retry-After`         | Included on `429`; minimum seconds to wait before retrying. |

Browser applications can read these headers because the Talqora API exposes them through CORS.

## When a limit is exceeded

The API returns `429 Too Many Requests` before sending work to the underlying search infrastructure.

```json
{
  "status": "error",
  "error": "Queries rate limit exceeded"
}
```

Wait for the number of seconds in `Retry-After`, then retry. Add small random jitter when many workers share one organization so they do not all retry simultaneously.

```typescript
async function talqoraRequest(url: string, init: RequestInit) {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(url, init);
    if (response.status !== 429) return response;

    const retryAfter = Number(response.headers.get("Retry-After") ?? "1");
    const jitter = Math.floor(Math.random() * 250);
    await new Promise(resolve => setTimeout(resolve, retryAfter * 1000 + jitter));
  }
  throw new Error("Talqora rate limit did not recover after five attempts");
}
```

For writes and deletes, reuse the same `Idempotency-Key` and exact request body when retrying. A rate-limited request is rejected before mutation and does not increment monthly usage.

## Rate limits versus monthly limits

A `429` means the organization sent requests too quickly. A `402` means the plan's monthly or stored-resource allowance has been exhausted. Waiting resolves a `429`; it does not resolve a `402`. Upgrade the plan or wait for the next billing-period reset when a monthly allowance is exhausted.