Search and filters
POST /v1/indexes/{index_id}/query is the retrieval endpoint for application-owned embeddings, lexical queries, processed files, connector sources, and Assistant RAG foundations. It returns the best valid live records for one index after applying the requested retrieval mode, metadata filter, result count, and relevance threshold.
Search is deliberately index-scoped. An API key cannot query an index outside its configured scope, and an index ID is never inferred from a record ID or metadata value.
Choose a retrieval mode
Use hybrid as the default for natural-language product search, policy retrieval, and assistant grounding. Send one query_text and Talqora creates the index-compatible embedding and uses the same text for BM25. API keys still enforce index scope and read permission. Use dense only when lexical text is unavailable or intentionally irrelevant. Use sparse when a query is primarily a literal identifier or keyword lookup.
Search with plain text
For the simplest public API call, send query_text. Talqora generates the embedding server-side, so no embedding-provider credential is needed in your application. If embedding is unavailable, the API returns a controlled 502 or 503; malformed or blank input returns 422.
query_text works with dense, hybrid, sparse, exact, and regex. For dense and hybrid, it is embedded only when vector is omitted. For hybrid, sparse_query defaults to query_text when omitted. Supplying vector remains fully supported and avoids server-side embedding.
Run a filtered hybrid query
The vector in this short example must be replaced with one that has the exact index dimensions. Alternatively, replace both vector and sparse_query with one query_text field and Talqora handles embedding. In a browser product, keep the Talqora API key on your trusted backend; do not expose it to an untrusted browser.
top_k and thresholds
top_k controls the number of final results returned to the caller. Request only as many as the next stage needs: an interactive result list may use 10–20, while an LLM context builder may retrieve a larger candidate set and rerank or trim it before generation.
sparse_query accepts up to 10,000 characters. Longer text is rejected with 422 and field: "sparse_query"; trim it client-side or summarize a bulk source before querying.
min_score is a relevance guard. It lets an application withhold weak matches rather than always returning a nearest neighbor. Calibrate it with representative queries; score distributions differ across content, embedding models, retrieval modes, and index metrics. Do not copy a threshold from another index without evaluation.
If no result clears the threshold, treat that as useful product behavior: show “no evidence found,” ask a clarifying question, or broaden the search only if your authorization policy allows it.
Metadata filters
Filters constrain the candidate set before results are returned. Use them for tenant isolation, document lifecycle, region, product catalog availability, language, jurisdiction, ACL labels, and any product-specific boundary.
Talqora validates filters before either retrieval path runs. The portable filter subset is equality ($eq), set membership ($in), and boolean composition ($and, $or). Unsupported operators such as $ne, $gt, $gte, $lt, $lte, malformed predicates, and incompatible system provenance types return 422; they are never silently ignored.
Combine predicates with $and when every condition is required. Use $or only when either branch is allowed. Keep filters small and schema-driven; a filter is a request constraint, not a substitute for validating an end user’s access in your own application.
For a compact all-of filter, a flat object is also accepted and is treated as $and. For example, {"tenant": 324, "source": 324} only returns records matching both values. Talqora canonicalizes that shorthand before sending a dense query to S3 Vectors, which requires a single root filter expression.
Authorization pattern
For a multi-tenant index, attach tenant_id to every record and include it in every query filter derived from a trusted session. For per-user knowledge, include both tenant_id and a visibility or principal field. Do not accept an arbitrary tenant filter directly from an untrusted browser request.
API-key scoping and metadata filtering complement each other: API keys limit which indexes a service can reach; filters limit which records within an allowed index may be returned.
Lexical semantics: accents, phrases, and numbers
Sparse retrieval is case-insensitive and accent-tolerant for Spanish terms. For example, autenticacion can match autenticación without replacing the stored original text.
Use double quotes for exact lexical phrase matching: "ingeniero de guardia" is a phrase query. Without quotes, sparse search is an OR-style term query intended for broad keyword retrieval, so a query such as 400 EUR plan bienestar can still return content matching plan and bienestar even if 400 is absent. Use a quoted phrase for a literal phrase, metadata filters for structured constraints, and min_score for a relevance threshold.
Results, provenance, and latency
Results include record IDs, metadata, ranking information, and available text snippets. Processed and connector-derived records preserve provenance such as source_file, page, chunk, job_id, and source_from_connector. Use those fields to render citations, link to an original source, or explain why the system selected a result.
Ranking values should be used for ordering and thresholding within the same retrieval mode and index. Do not present them as a universal probability or compare raw values across different embedding models, metrics, or query types. Every response also reports latency_ms; hybrid responses additionally report dense, sparse, and fusion phase timings. Measure p50 and p95 against your own region, vector dimensions, filter selectivity, and cache state rather than assuming a universal latency target.
Production query loop
- Authenticate the request and resolve the caller’s allowed tenant, user, and resource scope.
- Build a dense vector, sparse query, or both from the user input.
- Apply mandatory metadata filters server-side.
- Query with a bounded
top_kand evaluatedmin_score. - Render source metadata or pass only the selected snippets to an agent.
- Record latency, empty-result rate, click/acceptance signals, and query transfer to improve the system.
See Hybrid retrieval for ranking behavior and Assistant RAG for grounded chat on top of this retrieval layer.
Literal identifiers
Use search_type: "exact" for a copied email, error code, UUID, URL, hash, or version string. Use search_type: "regex" for a bounded pattern over those literal tokens. Both are separate from BM25 phrase search, which remains the right option for human-language text.