> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.talqora.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.talqora.com/_mcp/server.

# Indexes

> Design an isolated regional retrieval store for one workload, data boundary, or release.

An index is the durable retrieval boundary in Talqora Vector. It is not a folder, a database table, or a temporary query collection. It is an isolated regional store with a fixed vector contract: **AWS region, dimensions, distance metric, and metadata indexing policy**.

Everything written through the Vector API, Serverless Processing, a connector sync, or a public website crawl ends up in an index. Dense retrieval, BM25 sparse retrieval, hybrid retrieval, Assistant RAG, usage accounting, API-key scopes, and deletion all operate against that same boundary.

Use an index to answer a simple operational question: **which records may be searched together, under the same vector format and data-residency policy?**

## What belongs in an index

An index can hold any record that your application can represent as an embedding plus optional metadata and lexical text. Common examples include:

| Workload           | One record can represent                                        | Useful metadata                                                     |
| ------------------ | --------------------------------------------------------------- | ------------------------------------------------------------------- |
| Product discovery  | a product, SKU, variant, or catalog paragraph                   | `tenant_id`, `category`, `brand`, `in_stock`, `price`, `locale`     |
| RAG knowledge base | a chunk of a policy, contract, handbook, ticket, or webpage     | `source_file`, `page`, `document_type`, `department`, `visibility`  |
| Support copilot    | a resolved case, troubleshooting article, or conversation chunk | `product`, `language`, `status`, `created_at`, `account_id`         |
| Agent memory       | a durable event, decision, observation, or user preference      | `user_id`, `agent_id`, `namespace`, `created_at`                    |
| Recommendations    | an item, content unit, creator, or user profile representation  | `market`, `availability`, `content_type`, `safety_status`           |
| Compliance search  | a paragraph, clause, control, evidence item, or audit finding   | `jurisdiction`, `policy_version`, `retention_class`, `access_level` |

Each record has a stable `id`, a vector, metadata, and optionally `sparse_text`. The vector captures semantic similarity. `sparse_text` makes exact names, codes, dates, legal terms, and uncommon phrases available to lexical search. Metadata is for filtering and authorization-aware retrieval.

An index is **not** the right way to store raw video, arbitrary blobs, user passwords, secrets, or data that should never be returned by a search result. Keep source bytes in your own object storage or use Serverless Processing for supported documents; only send the retrieval representation and metadata needed by the application.

## One index or many?

Create a separate index when any of these are true:

* The data needs a different **AWS region** or residency boundary.
* The embedding model produces a different **dimension count**.
* The workload needs a different distance behavior.
* Records must never appear in the same search result, even by mistake.
* A separate API key, billing boundary, lifecycle, retention policy, or release process is required.

Keep records in the same index when they use the same embedding contract and you can safely separate them with metadata filters. For example, a multi-tenant product search service can use one `products-us` index with a mandatory `tenant_id` filter on every query. A knowledge base can keep policies, ticket summaries, and approved FAQs together when the application is permitted to search all three.

### Practical patterns

**One index per tenant** is simplest when each customer needs an independent API key, region, deletion lifecycle, or hard search boundary. It works well for enterprise RAG and white-label products.

**One index per workload** is usually better for high-cardinality SaaS data. For example, create `catalog-us`, `support-us`, and `agent-memory-us`, then filter by `tenant_id` within each. This reduces operational objects while keeping records with different retrieval intent apart.

**One index per release** is appropriate when changing embedding models or ranking data. Create a branch or a new index, backfill it, evaluate recall and latency, then point the application at the new index. Do not change an existing index's dimensions in place.

**One index per region** is required when the workload must stay close to users or regulated data. A `knowledge-eu` index and a `knowledge-us` index may contain equivalent schemas but have separate physical placement and usage.

## Immutable index contract

Talqora validates the following at creation and never changes them afterward:

| Setting                        | Meaning                                                          | Why it is immutable                                                                                       |
| ------------------------------ | ---------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
| `aws_region`                   | Regional home for dense vectors and retrieval data.              | Moving records changes latency, durability, and residency guarantees. Create a new index to move regions. |
| `dimensions`                   | Number of floating-point values in every vector.                 | A similarity index cannot compare vectors with different shapes.                                          |
| `distance_metric`              | The ranking geometry for dense queries: `cosine` or `euclidean`. | Existing rankings are defined by the metric. Changing it would silently reorder every result.             |
| `non_filterable_metadata_keys` | Metadata keys stored with records but excluded from filtering.   | The physical index policy is chosen before records are written.                                           |

The only mutable index property is its display name. Rename an index with `PATCH /v1/indexes/{index_id}`; this does not move or rewrite data.

## Choose dimensions deliberately

A dimension is one coordinate in an embedding vector. A 1536-dimensional embedding has 1,536 numeric values. Talqora validates the exact count on every write; a 768-dimensional vector cannot be written to a 1536-dimensional index.

Talqora supports **1 through 4,096 dimensions** for direct Vector API writes. The right number is determined by the embedding model, not by a preference for a larger index:

* Use the exact output dimension of the model already used by your application.
* Do not pad, truncate, or mix vectors from models with different dimensions.
* Higher dimensions increase vector bytes, storage, write transfer, query transfer, and local application CPU.
* More dimensions do not automatically improve relevance. Model quality, chunking, metadata, and evaluation data matter more than simply using a larger representation.

When `dimensions` is omitted while creating an index, Talqora uses **1536**. This makes the common File Processing path concise because the managed pipeline generates 1536-dimensional embeddings. Direct-write applications that use another embedding model must explicitly provide that model's exact dimension.

### Why Serverless Processing requires 1536 dimensions

Talqora's Serverless Processing product owns extraction, OCR when necessary, chunking, and embedding generation. Its processing pipeline generates 1536-dimensional embeddings. Therefore an index that receives uploaded files, connector content, or crawled pages must be created with `dimensions: 1536`.

If you only use direct vector writes, choose the dimension your embedding model actually emits. If you want both direct writes and Serverless Processing in one index, generate compatible 1536-dimensional vectors in your direct path as well.

## Choose a distance metric

If `distance_metric` is omitted at index creation, Talqora uses **`cosine`**. This is the normal default for text and multimodal embeddings. Send `euclidean` explicitly only when your embedding model and offline evaluation require it.

`cosine` ranks vectors by directional similarity. It is the normal choice for text and multimodal embeddings that are normalized by the model or client. It answers whether two representations point toward similar concepts, independent of their length.

`euclidean` ranks by geometric distance and preserves magnitude differences. Use it only when the model and your offline evaluation specifically call for Euclidean distance. Do not choose it merely because it sounds more exact; the embedding model's documentation and an evaluation set should drive this decision.

The metric applies to dense retrieval. Sparse retrieval ranks lexical matches, and hybrid retrieval combines dense and sparse candidates. In all cases, enforce a relevance threshold in your application when a low-quality result should be withheld rather than returned.

## Metadata and filtering

Metadata is a JSON object stored beside a vector. Use it for values you need to filter or return with a result, such as a source filename, tenant, product category, language, page number, ACL label, timestamp, or lifecycle state.

```json
{
  "id": "policy-2026:page-12:chunk-3",
  "values": [0.012, 0.483, "..."],
  "metadata": {
    "tenant_id": "acme",
    "source_file": "vendor-retention-policy.pdf",
    "page": 12,
    "jurisdiction": "EU",
    "visibility": "legal"
  },
  "sparse_text": "Vendor retention and subprocessors policy"
}
```

By default, metadata keys are filterable. At creation, you can mark up to ten keys as `non_filterable_metadata_keys` when they are useful to return but should not participate in filters. Key names must be unique and 1–63 characters long.

Filtering is powerful but is not a replacement for authorization. Your service must always apply the tenant, user, or ACL filter required by the requester. Scope API keys to the minimum set of indexes and treat a missing authorization filter as an application bug.

## Region placement

`aws_region` chooses the physical home of an index. Ask `GET /v1/regions` before creation because available regions are controlled by your Talqora account. Current provisioned dense regions are `us-east-1`, `sa-east-1`, `eu-west-1`, and `ap-southeast-1`; availability can vary by account and retrieval capability.

Choose the region closest to the compute that will make most writes and queries, subject to your data-residency requirements. The endpoint hostname is only a routing preference; it does not relocate the index. See [Regional API endpoints](/get-started/regional-api-endpoints) for endpoint selection and [Hybrid retrieval](/build-retrieval/hybrid-retrieval) for sparse availability.

If `aws_region` is omitted when creating an index, Talqora assigns **`us-east-1`**. Send `aws_region` explicitly whenever your workload has a residency or latency requirement for another supported region; a region is immutable after the index is created.

## Create an index

Create an index with a dashboard session token or an unscoped Talqora API key with `write` permission. Index names are lowercase identifiers, 3–63 characters long, and may contain letters, numbers, hyphens, and periods. A key restricted to specific `index_ids` cannot create a new index; it can read, rename, branch, or delete only indexes in its scope.

```bash
curl --fail-with-body https://api.talqora.com/v1/indexes \
  -X POST \
  -H "Authorization: Bearer $TALQORA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "compliance-us",
    "distance_metric": "cosine",
    "non_filterable_metadata_keys": ["raw_source_checksum"]
  }'
```

The response includes the permanent index ID. Use that ID, not the display name, in data-plane URLs and API-key scopes.

```json
{
  "id": "idx_1234...",
  "name": "compliance-us",
  "dimensions": 1536,
  "distance_metric": "cosine",
  "aws_region": "us-east-1",
  "non_filterable_metadata_keys": ["raw_source_checksum"]
}
```

If the regional dense or sparse dependency needed by the selected configuration is unavailable, creation fails before a usable index is returned. Talqora compensates partial creation rather than leaving a half-configured index in the workspace.

## Lifecycle, usage, and branches

`GET /v1/indexes/{index_id}/stats` reports the current cardinality and usage for the index: documents, logical storage, rows written, data written, queries, queried transfer, and activity timestamps. `rows_written` is cumulative accepted write activity; `documents` is the current number of live records, so the two numbers intentionally diverge after upserts and deletes.

Use [index branching](/build-retrieval/index-branching) to create a materialized, isolated copy for an embedding-model migration, ranking experiment, or release candidate. A branch retains the source contract and consumes its own documents, storage, and usage because it is a complete independent retrieval store.

Deleting an index removes its dense vectors, sparse retrieval data, processing sources, and control-plane records. Deletion is irreversible. Revoke or narrow API keys before deleting an index if an application may still send requests to it.

## Design checklist

Before creating an index, answer these questions:

1. What records are allowed to appear in the same result set?
2. Which AWS region satisfies latency and residency requirements?
3. Which embedding model will write vectors, and what exact dimension does it emit?
4. Will the index use Talqora Serverless Processing? If yes, choose 1536 dimensions.
5. Which metadata fields must be filtered on every query for tenant and access isolation?
6. Does this represent a production workload, an experiment, or a release candidate?
7. Which applications need read versus write access, and can each API key be scoped to this index alone?

Once those answers are stable, create the index, issue a least-privilege API key, and proceed to [Write vectors](/build-retrieval/write-vectors) or [File processing](/build-retrieval/file-processing).