Rate limits
Talqora enforces request rate limits on the server after authentication and before expensive vector-storage or sparse-search work begins. Limits use a rolling 60-second window and are aggregated by organization, not by API key.
Limits by plan
The buckets are independent:
- Control plane includes authenticated organization, index, API-key, usage, member, and billing operations.
- Writes and deletes share one bucket across every index and API key in the organization.
- Queries share a separate bucket across dense, sparse, and hybrid retrieval.
Creating five API keys does not provide five times the request capacity. This prevents accidental key fan-out from bypassing an organization’s plan. A batch containing 500 vectors consumes one write request, although all 500 rows count toward monthly usage.
Response headers
Successful rate-limited requests return:
Browser applications can read these headers because the Talqora API exposes them through CORS.
When a limit is exceeded
The API returns 429 Too Many Requests before sending work to the underlying search infrastructure.
Wait for the number of seconds in Retry-After, then retry. Add small random jitter when many workers share one organization so they do not all retry simultaneously.
For writes and deletes, reuse the same Idempotency-Key and exact request body when retrying. A rate-limited request is rejected before mutation and does not increment monthly usage.
Rate limits versus monthly limits
A 429 means the organization sent requests too quickly. A 402 means the plan’s monthly or stored-resource allowance has been exhausted. Waiting resolves a 429; it does not resolve a 402. Upgrade the plan or wait for the next billing-period reset when a monthly allowance is exhausted.