Skip to navigation

Indexes

An index is the durable retrieval boundary in Talqora Vector. It is not a folder, a database table, or a temporary query collection. It is an isolated regional store with a fixed vector contract: AWS region, dimensions, distance metric, and metadata indexing policy.

Everything written through the Vector API, Serverless Processing, a connector sync, or a public website crawl ends up in an index. Dense retrieval, BM25 sparse retrieval, hybrid retrieval, Assistant RAG, usage accounting, API-key scopes, and deletion all operate against that same boundary.

Use an index to answer a simple operational question: which records may be searched together, under the same vector format and data-residency policy?

What belongs in an index

An index can hold any record that your application can represent as an embedding plus optional metadata and lexical text. Common examples include:

WorkloadOne record can representUseful metadata
Product discoverya product, SKU, variant, or catalog paragraphtenant_id, category, brand, in_stock, price, locale
RAG knowledge basea chunk of a policy, contract, handbook, ticket, or webpagesource_file, page, document_type, department, visibility
Support copilota resolved case, troubleshooting article, or conversation chunkproduct, language, status, created_at, account_id
Agent memorya durable event, decision, observation, or user preferenceuser_id, agent_id, namespace, created_at
Recommendationsan item, content unit, creator, or user profile representationmarket, availability, content_type, safety_status
Compliance searcha paragraph, clause, control, evidence item, or audit findingjurisdiction, policy_version, retention_class, access_level

Each record has a stable id, a vector, metadata, and optionally sparse_text. The vector captures semantic similarity. sparse_text makes exact names, codes, dates, legal terms, and uncommon phrases available to lexical search. Metadata is for filtering and authorization-aware retrieval.

An index is not the right way to store raw video, arbitrary blobs, user passwords, secrets, or data that should never be returned by a search result. Keep source bytes in your own object storage or use Serverless Processing for supported documents; only send the retrieval representation and metadata needed by the application.

One index or many?

Create a separate index when any of these are true:

  • The data needs a different AWS region or residency boundary.
  • The embedding model produces a different dimension count.
  • The workload needs a different distance behavior.
  • Records must never appear in the same search result, even by mistake.
  • A separate API key, billing boundary, lifecycle, retention policy, or release process is required.

Keep records in the same index when they use the same embedding contract and you can safely separate them with metadata filters. For example, a multi-tenant product search service can use one products-us index with a mandatory tenant_id filter on every query. A knowledge base can keep policies, ticket summaries, and approved FAQs together when the application is permitted to search all three.

Practical patterns

One index per tenant is simplest when each customer needs an independent API key, region, deletion lifecycle, or hard search boundary. It works well for enterprise RAG and white-label products.

One index per workload is usually better for high-cardinality SaaS data. For example, create catalog-us, support-us, and agent-memory-us, then filter by tenant_id within each. This reduces operational objects while keeping records with different retrieval intent apart.

One index per release is appropriate when changing embedding models or ranking data. Create a branch or a new index, backfill it, evaluate recall and latency, then point the application at the new index. Do not change an existing index’s dimensions in place.

One index per region is required when the workload must stay close to users or regulated data. A knowledge-eu index and a knowledge-us index may contain equivalent schemas but have separate physical placement and usage.

Immutable index contract

Talqora validates the following at creation and never changes them afterward:

SettingMeaningWhy it is immutable
aws_regionRegional home for dense vectors and retrieval data.Moving records changes latency, durability, and residency guarantees. Create a new index to move regions.
dimensionsNumber of floating-point values in every vector.A similarity index cannot compare vectors with different shapes.
distance_metricThe ranking geometry for dense queries: cosine or euclidean.Existing rankings are defined by the metric. Changing it would silently reorder every result.
non_filterable_metadata_keysMetadata keys stored with records but excluded from filtering.The physical index policy is chosen before records are written.

The only mutable index property is its display name. Rename an index with PATCH /v1/indexes/{index_id}; this does not move or rewrite data.

Choose dimensions deliberately

A dimension is one coordinate in an embedding vector. A 1536-dimensional embedding has 1,536 numeric values. Talqora validates the exact count on every write; a 768-dimensional vector cannot be written to a 1536-dimensional index.

Talqora supports 1 through 4,096 dimensions for direct Vector API writes. The right number is determined by the embedding model, not by a preference for a larger index:

  • Use the exact output dimension of the model already used by your application.
  • Do not pad, truncate, or mix vectors from models with different dimensions.
  • Higher dimensions increase vector bytes, storage, write transfer, query transfer, and local application CPU.
  • More dimensions do not automatically improve relevance. Model quality, chunking, metadata, and evaluation data matter more than simply using a larger representation.

When dimensions is omitted while creating an index, Talqora uses 1536. This makes the common File Processing path concise because the managed pipeline generates 1536-dimensional embeddings. Direct-write applications that use another embedding model must explicitly provide that model’s exact dimension.

Why Serverless Processing requires 1536 dimensions

Talqora’s Serverless Processing product owns extraction, OCR when necessary, chunking, and embedding generation. Its processing pipeline generates 1536-dimensional embeddings. Therefore an index that receives uploaded files, connector content, or crawled pages must be created with dimensions: 1536.

If you only use direct vector writes, choose the dimension your embedding model actually emits. If you want both direct writes and Serverless Processing in one index, generate compatible 1536-dimensional vectors in your direct path as well.

Choose a distance metric

If distance_metric is omitted at index creation, Talqora uses cosine. This is the normal default for text and multimodal embeddings. Send euclidean explicitly only when your embedding model and offline evaluation require it.

cosine ranks vectors by directional similarity. It is the normal choice for text and multimodal embeddings that are normalized by the model or client. It answers whether two representations point toward similar concepts, independent of their length.

euclidean ranks by geometric distance and preserves magnitude differences. Use it only when the model and your offline evaluation specifically call for Euclidean distance. Do not choose it merely because it sounds more exact; the embedding model’s documentation and an evaluation set should drive this decision.

The metric applies to dense retrieval. Sparse retrieval ranks lexical matches, and hybrid retrieval combines dense and sparse candidates. In all cases, enforce a relevance threshold in your application when a low-quality result should be withheld rather than returned.

Metadata and filtering

Metadata is a JSON object stored beside a vector. Use it for values you need to filter or return with a result, such as a source filename, tenant, product category, language, page number, ACL label, timestamp, or lifecycle state.

{
"id": "policy-2026:page-12:chunk-3",
"values": [0.012, 0.483, "..."],
"metadata": {
"tenant_id": "acme",
"source_file": "vendor-retention-policy.pdf",
"page": 12,
"jurisdiction": "EU",
"visibility": "legal"
},
"sparse_text": "Vendor retention and subprocessors policy"
}

By default, metadata keys are filterable. At creation, you can mark up to ten keys as non_filterable_metadata_keys when they are useful to return but should not participate in filters. Key names must be unique and 1–63 characters long.

Filtering is powerful but is not a replacement for authorization. Your service must always apply the tenant, user, or ACL filter required by the requester. Scope API keys to the minimum set of indexes and treat a missing authorization filter as an application bug.

Region placement

aws_region chooses the physical home of an index. Ask GET /v1/regions before creation because available regions are controlled by your Talqora account. Current provisioned dense regions are us-east-1, sa-east-1, eu-west-1, and ap-southeast-1; availability can vary by account and retrieval capability.

Choose the region closest to the compute that will make most writes and queries, subject to your data-residency requirements. The endpoint hostname is only a routing preference; it does not relocate the index. See Regional API endpoints for endpoint selection and Hybrid retrieval for sparse availability.

If aws_region is omitted when creating an index, Talqora assigns us-east-1. Send aws_region explicitly whenever your workload has a residency or latency requirement for another supported region; a region is immutable after the index is created.

Create an index

Create an index with a dashboard session token or an unscoped Talqora API key with write permission. Index names are lowercase identifiers, 3–63 characters long, and may contain letters, numbers, hyphens, and periods. A key restricted to specific index_ids cannot create a new index; it can read, rename, branch, or delete only indexes in its scope.

curl --fail-with-body https://api.talqora.com/v1/indexes \
-X POST \
-H "Authorization: Bearer $TALQORA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "compliance-us",
"distance_metric": "cosine",
"non_filterable_metadata_keys": ["raw_source_checksum"]
}'

The response includes the permanent index ID. Use that ID, not the display name, in data-plane URLs and API-key scopes.

{
"id": "idx_1234...",
"name": "compliance-us",
"dimensions": 1536,
"distance_metric": "cosine",
"aws_region": "us-east-1",
"non_filterable_metadata_keys": ["raw_source_checksum"]
}

If the regional dense or sparse dependency needed by the selected configuration is unavailable, creation fails before a usable index is returned. Talqora compensates partial creation rather than leaving a half-configured index in the workspace.

Lifecycle, usage, and branches

GET /v1/indexes/{index_id}/stats reports the current cardinality and usage for the index: documents, logical storage, rows written, data written, queries, queried transfer, and activity timestamps. rows_written is cumulative accepted write activity; documents is the current number of live records, so the two numbers intentionally diverge after upserts and deletes.

Use index branching to create a materialized, isolated copy for an embedding-model migration, ranking experiment, or release candidate. A branch retains the source contract and consumes its own documents, storage, and usage because it is a complete independent retrieval store.

Deleting an index removes its dense vectors, sparse retrieval data, processing sources, and control-plane records. Deletion is irreversible. Revoke or narrow API keys before deleting an index if an application may still send requests to it.

Design checklist

Before creating an index, answer these questions:

  1. What records are allowed to appear in the same result set?
  2. Which AWS region satisfies latency and residency requirements?
  3. Which embedding model will write vectors, and what exact dimension does it emit?
  4. Will the index use Talqora Serverless Processing? If yes, choose 1536 dimensions.
  5. Which metadata fields must be filtered on every query for tenant and access isolation?
  6. Does this represent a production workload, an experiment, or a release candidate?
  7. Which applications need read versus write access, and can each API key be scoped to this index alone?

Once those answers are stable, create the index, issue a least-privilege API key, and proceed to Write vectors or File processing.