> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.talqora.com/operate-at-scale/rate-limits/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.talqora.com/_mcp/server. # Rate limits > Understand organization-wide request budgets, response headers, and safe retry behavior. Talqora enforces request rate limits on the server after authentication and before expensive vector-storage or sparse-search work begins. Limits use a rolling 60-second window and are aggregated by organization, not by API key. ## Limits by plan | Plan | Control plane | Writes and deletes | Queries | | ---------- | ------------: | -----------------: | ---------: | | Developer | 120/min | 30/min | 60/min | | Scale | 3,000/min | 1,500/min | 3,000/min | | Enterprise | 20,000/min | 10,000/min | 20,000/min | The buckets are independent: * **Control plane** includes authenticated organization, index, API-key, usage, member, and billing operations. * **Writes and deletes** share one bucket across every index and API key in the organization. * **Queries** share a separate bucket across dense, sparse, and hybrid retrieval. Creating five API keys does not provide five times the request capacity. This prevents accidental key fan-out from bypassing an organization's plan. A batch containing 500 vectors consumes one write request, although all 500 rows count toward monthly usage. ## Response headers Successful rate-limited requests return: ```http RateLimit-Limit: 600 RateLimit-Remaining: 594 RateLimit-Reset: 41 ``` | Header | Meaning | | --------------------- | ----------------------------------------------------------- | | `RateLimit-Limit` | Maximum requests permitted in the active 60-second window. | | `RateLimit-Remaining` | Requests still available in that bucket. | | `RateLimit-Reset` | Seconds until the oldest request leaves the rolling window. | | `Retry-After` | Included on `429`; minimum seconds to wait before retrying. | Browser applications can read these headers because the Talqora API exposes them through CORS. ## When a limit is exceeded The API returns `429 Too Many Requests` before sending work to the underlying search infrastructure. ```json { "status": "error", "error": "Queries rate limit exceeded" } ``` Wait for the number of seconds in `Retry-After`, then retry. Add small random jitter when many workers share one organization so they do not all retry simultaneously. ```typescript async function talqoraRequest(url: string, init: RequestInit) { for (let attempt = 0; attempt < 5; attempt += 1) { const response = await fetch(url, init); if (response.status !== 429) return response; const retryAfter = Number(response.headers.get("Retry-After") ?? "1"); const jitter = Math.floor(Math.random() * 250); await new Promise(resolve => setTimeout(resolve, retryAfter * 1000 + jitter)); } throw new Error("Talqora rate limit did not recover after five attempts"); } ``` For writes and deletes, reuse the same `Idempotency-Key` and exact request body when retrying. A rate-limited request is rejected before mutation and does not increment monthly usage. ## Rate limits versus monthly limits A `429` means the organization sent requests too quickly. A `402` means the plan's monthly or stored-resource allowance has been exhausted. Waiting resolves a `429`; it does not resolve a `402`. Upgrade the plan or wait for the next billing-period reset when a monthly allowance is exhausted. > Understand organization-wide request budgets, response headers, and safe retry behavior.