Indexes
An index is the durable retrieval boundary in Talqora Vector. It is not a folder, a database table, or a temporary query collection. It is an isolated regional store with a fixed vector contract: AWS region, dimensions, distance metric, and metadata indexing policy.
Everything written through the Vector API, Serverless Processing, a connector sync, or a public website crawl ends up in an index. Dense retrieval, BM25 sparse retrieval, hybrid retrieval, Assistant RAG, usage accounting, API-key scopes, and deletion all operate against that same boundary.
Use an index to answer a simple operational question: which records may be searched together, under the same vector format and data-residency policy?
What belongs in an index
An index can hold any record that your application can represent as an embedding plus optional metadata and lexical text. Common examples include:
Each record has a stable id, a vector, metadata, and optionally sparse_text. The vector captures semantic similarity. sparse_text makes exact names, codes, dates, legal terms, and uncommon phrases available to lexical search. Metadata is for filtering and authorization-aware retrieval.
An index is not the right way to store raw video, arbitrary blobs, user passwords, secrets, or data that should never be returned by a search result. Keep source bytes in your own object storage or use Serverless Processing for supported documents; only send the retrieval representation and metadata needed by the application.
One index or many?
Create a separate index when any of these are true:
- The data needs a different AWS region or residency boundary.
- The embedding model produces a different dimension count.
- The workload needs a different distance behavior.
- Records must never appear in the same search result, even by mistake.
- A separate API key, billing boundary, lifecycle, retention policy, or release process is required.
Keep records in the same index when they use the same embedding contract and you can safely separate them with metadata filters. For example, a multi-tenant product search service can use one products-us index with a mandatory tenant_id filter on every query. A knowledge base can keep policies, ticket summaries, and approved FAQs together when the application is permitted to search all three.
Practical patterns
One index per tenant is simplest when each customer needs an independent API key, region, deletion lifecycle, or hard search boundary. It works well for enterprise RAG and white-label products.
One index per workload is usually better for high-cardinality SaaS data. For example, create catalog-us, support-us, and agent-memory-us, then filter by tenant_id within each. This reduces operational objects while keeping records with different retrieval intent apart.
One index per release is appropriate when changing embedding models or ranking data. Create a branch or a new index, backfill it, evaluate recall and latency, then point the application at the new index. Do not change an existing index’s dimensions in place.
One index per region is required when the workload must stay close to users or regulated data. A knowledge-eu index and a knowledge-us index may contain equivalent schemas but have separate physical placement and usage.
Immutable index contract
Talqora validates the following at creation and never changes them afterward:
The only mutable index property is its display name. Rename an index with PATCH /v1/indexes/{index_id}; this does not move or rewrite data.
Choose dimensions deliberately
A dimension is one coordinate in an embedding vector. A 1536-dimensional embedding has 1,536 numeric values. Talqora validates the exact count on every write; a 768-dimensional vector cannot be written to a 1536-dimensional index.
Talqora supports 1 through 4,096 dimensions for direct Vector API writes. The right number is determined by the embedding model, not by a preference for a larger index:
- Use the exact output dimension of the model already used by your application.
- Do not pad, truncate, or mix vectors from models with different dimensions.
- Higher dimensions increase vector bytes, storage, write transfer, query transfer, and local application CPU.
- More dimensions do not automatically improve relevance. Model quality, chunking, metadata, and evaluation data matter more than simply using a larger representation.
When dimensions is omitted while creating an index, Talqora uses 1536. This makes the common File Processing path concise because the managed pipeline generates 1536-dimensional embeddings. Direct-write applications that use another embedding model must explicitly provide that model’s exact dimension.
Why Serverless Processing requires 1536 dimensions
Talqora’s Serverless Processing product owns extraction, OCR when necessary, chunking, and embedding generation. Its processing pipeline generates 1536-dimensional embeddings. Therefore an index that receives uploaded files, connector content, or crawled pages must be created with dimensions: 1536.
If you only use direct vector writes, choose the dimension your embedding model actually emits. If you want both direct writes and Serverless Processing in one index, generate compatible 1536-dimensional vectors in your direct path as well.
Choose a distance metric
If distance_metric is omitted at index creation, Talqora uses cosine. This is the normal default for text and multimodal embeddings. Send euclidean explicitly only when your embedding model and offline evaluation require it.
cosine ranks vectors by directional similarity. It is the normal choice for text and multimodal embeddings that are normalized by the model or client. It answers whether two representations point toward similar concepts, independent of their length.
euclidean ranks by geometric distance and preserves magnitude differences. Use it only when the model and your offline evaluation specifically call for Euclidean distance. Do not choose it merely because it sounds more exact; the embedding model’s documentation and an evaluation set should drive this decision.
The metric applies to dense retrieval. Sparse retrieval ranks lexical matches, and hybrid retrieval combines dense and sparse candidates. In all cases, enforce a relevance threshold in your application when a low-quality result should be withheld rather than returned.
Metadata and filtering
Metadata is a JSON object stored beside a vector. Use it for values you need to filter or return with a result, such as a source filename, tenant, product category, language, page number, ACL label, timestamp, or lifecycle state.
By default, metadata keys are filterable. At creation, you can mark up to ten keys as non_filterable_metadata_keys when they are useful to return but should not participate in filters. Key names must be unique and 1–63 characters long.
Filtering is powerful but is not a replacement for authorization. Your service must always apply the tenant, user, or ACL filter required by the requester. Scope API keys to the minimum set of indexes and treat a missing authorization filter as an application bug.
Region placement
aws_region chooses the physical home of an index. Ask GET /v1/regions before creation because available regions are controlled by your Talqora account. Current provisioned dense regions are us-east-1, sa-east-1, eu-west-1, and ap-southeast-1; availability can vary by account and retrieval capability.
Choose the region closest to the compute that will make most writes and queries, subject to your data-residency requirements. The endpoint hostname is only a routing preference; it does not relocate the index. See Regional API endpoints for endpoint selection and Hybrid retrieval for sparse availability.
If aws_region is omitted when creating an index, Talqora assigns us-east-1. Send aws_region explicitly whenever your workload has a residency or latency requirement for another supported region; a region is immutable after the index is created.
Create an index
Create an index with a dashboard session token or an unscoped Talqora API key with write permission. Index names are lowercase identifiers, 3–63 characters long, and may contain letters, numbers, hyphens, and periods. A key restricted to specific index_ids cannot create a new index; it can read, rename, branch, or delete only indexes in its scope.
The response includes the permanent index ID. Use that ID, not the display name, in data-plane URLs and API-key scopes.
If the regional dense or sparse dependency needed by the selected configuration is unavailable, creation fails before a usable index is returned. Talqora compensates partial creation rather than leaving a half-configured index in the workspace.
Lifecycle, usage, and branches
GET /v1/indexes/{index_id}/stats reports the current cardinality and usage for the index: documents, logical storage, rows written, data written, queries, queried transfer, and activity timestamps. rows_written is cumulative accepted write activity; documents is the current number of live records, so the two numbers intentionally diverge after upserts and deletes.
Use index branching to create a materialized, isolated copy for an embedding-model migration, ranking experiment, or release candidate. A branch retains the source contract and consumes its own documents, storage, and usage because it is a complete independent retrieval store.
Deleting an index removes its dense vectors, sparse retrieval data, processing sources, and control-plane records. Deletion is irreversible. Revoke or narrow API keys before deleting an index if an application may still send requests to it.
Design checklist
Before creating an index, answer these questions:
- What records are allowed to appear in the same result set?
- Which AWS region satisfies latency and residency requirements?
- Which embedding model will write vectors, and what exact dimension does it emit?
- Will the index use Talqora Serverless Processing? If yes, choose 1536 dimensions.
- Which metadata fields must be filtered on every query for tenant and access isolation?
- Does this represent a production workload, an experiment, or a release candidate?
- Which applications need read versus write access, and can each API key be scoped to this index alone?
Once those answers are stable, create the index, issue a least-privilege API key, and proceed to Write vectors or File processing.