Skip to navigation

Hybrid retrieval

Hybrid retrieval combines two independent signals from the same Talqora index:

  1. Dense retrieval finds content that means something similar even when the language changes.
  2. Sparse retrieval finds exact terms, SKU values, ticket IDs, error codes, legal clauses, names, and rare phrases.

Talqora runs each path independently, validates live sparse versions, and merges the ranked candidate lists with reciprocal rank fusion (RRF). The client does not need to normalize a dense distance against a BM25 score, which are different measurements and should not be compared directly.

Why hybrid matters

A customer asking for “footwear for wet mountain routes” may never say “waterproof trail shoe,” but semantic retrieval can find the right catalog entry. A customer searching for ERR_AUTH_429, MSA-2026-07, or a person’s name expects exact vocabulary to dominate. Hybrid search covers both without requiring the application to guess which type of query it received.

Use it for product catalogs, help centers, policy and contract search, support copilots, internal knowledge, and grounded assistants. Dense-only or sparse-only search remains useful when your workload has a clearly single signal.

Request shape

{
"search_type": "hybrid",
"vector": [0.12, 0.18, 0.44],
"sparse_query": "waterproof trail shoe",
"top_k": 10,
"min_score": 0.15,
"filter": {
"$and": [
{"tenant_id": {"$eq": "acme"}},
{"status": {"$eq": "active"}},
{"region": {"$in": ["americas", "eu"]}}
]
}
}

Send the exact-dimensional vector produced by the embedding model for the index. sparse_query should be the original user query or a carefully normalized equivalent; do not replace it with the embedding text if that loses product codes, names, or quoted terms.

How ranking works

Dense and sparse engines each return a ranked candidate list. RRF rewards records that rank well in either list and rewards records that appear strongly in both. It is rank-based, so it avoids pretending that a dense distance and lexical relevance score have compatible units.

This has practical benefits:

  • An exact policy code can surface even when its semantic embedding is average.
  • A paraphrase can surface even when no literal terms overlap.
  • A record supported by both signals rises without manual score calibration.
  • Filtered-out, deleted, or stale sparse versions are removed before final results are returned.

The final result’s score is useful for order and a mode-specific threshold, but it is not a probability of truth. Evaluate thresholds per index using real accepted and rejected results.

Write one dense vector for the semantic content and one concise sparse_text representation for exact matching. For a product, include title, brand, SKU, category, important attributes, and user-visible terminology. For a document chunk, include its extracted text and meaningful headings. Avoid serializing opaque JSON, duplicated boilerplate, or credentials into sparse_text; it harms lexical quality and may expose data.

Keep useful metadata outside sparse_text. Metadata belongs in filters; content that should be matched belongs in the lexical text; text users should see as evidence belongs in a returned snippet or source representation.

Filtering and security

The same metadata filter is applied to both candidate paths. Always include the tenant, principal, visibility, or product boundary required by the caller before a hybrid query reaches Talqora. Hybrid ranking improves retrieval quality; it does not relax authorization.

For example, a global support index can contain many customers only if every record has a durable tenant_id and the service injects the authenticated tenant filter on every request. If hard boundaries are preferable, use one index per tenant and scope keys accordingly.

Tune with evaluation, not intuition

Build a small set of representative queries with expected sources. Include paraphrases, exact identifiers, ambiguous requests, stale terms, and no-answer queries. Compare dense, sparse, and hybrid modes at the same top_k and filter policy. Track recall of expected sources, empty-result behavior, latency, queried transfer, and user acceptance.

If hybrid adds irrelevant literal matches, improve sparse_text, enforce a stronger filter, or adjust your application threshold. If exact IDs are missing, verify those IDs are present in sparse_text and that the record is not deleted or superseded. Use double quotes around a literal phrase; without quotes, sparse terms are intentionally broad. If semantic results are weak, evaluate the embedding model and source chunking before increasing result counts.

Latency observability

The response includes latency_ms for the request and per-phase timings: dense_latency_ms, sparse_latency_ms, and fusion_latency_ms. Collect p50 and p95 separately by region and filter shape. A hybrid request necessarily includes two retrieval paths, so a healthy sparse p95 alone is not a hybrid SLO. Use these fields to isolate whether vector retrieval, lexical retrieval, or merge/snippet work dominates before changing a threshold or increasing concurrency.

When not to use it

Use dense-only retrieval when records have no meaningful lexical text or the query is already a trusted embedding. Use sparse-only retrieval for exact lookup interfaces where semantic expansion would be surprising. Use hybrid by default for human language search where both conceptual similarity and literal evidence are valuable.

See Search and filters for request construction and Sparse retrieval architecture for versioning and lexical consistency.