> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.talqora.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.talqora.com/_mcp/server.

# Hybrid retrieval

> Combine semantic and lexical evidence without forcing incomparable scores into one scale.

Hybrid retrieval combines two independent signals from the same Talqora index:

1. **Dense retrieval** finds content that means something similar even when the language changes.
2. **Sparse retrieval** finds exact terms, SKU values, ticket IDs, error codes, legal clauses, names, and rare phrases.

Talqora runs each path independently, validates live sparse versions, and merges the ranked candidate lists with reciprocal rank fusion (RRF). The client does not need to normalize a dense distance against a BM25 score, which are different measurements and should not be compared directly.

## Why hybrid matters

A customer asking for “footwear for wet mountain routes” may never say “waterproof trail shoe,” but semantic retrieval can find the right catalog entry. A customer searching for `ERR_AUTH_429`, `MSA-2026-07`, or a person's name expects exact vocabulary to dominate. Hybrid search covers both without requiring the application to guess which type of query it received.

Use it for product catalogs, help centers, policy and contract search, support copilots, internal knowledge, and grounded assistants. Dense-only or sparse-only search remains useful when your workload has a clearly single signal.

## Request shape

```json
{
  "search_type": "hybrid",
  "vector": [0.12, 0.18, 0.44],
  "sparse_query": "waterproof trail shoe",
  "top_k": 10,
  "min_score": 0.15,
  "filter": {
    "$and": [
      {"tenant_id": {"$eq": "acme"}},
      {"status": {"$eq": "active"}},
      {"region": {"$in": ["americas", "eu"]}}
    ]
  }
}
```

Send the exact-dimensional vector produced by the embedding model for the index. `sparse_query` should be the original user query or a carefully normalized equivalent; do not replace it with the embedding text if that loses product codes, names, or quoted terms.

## How ranking works

Dense and sparse engines each return a ranked candidate list. RRF rewards records that rank well in either list and rewards records that appear strongly in both. It is rank-based, so it avoids pretending that a dense distance and lexical relevance score have compatible units.

This has practical benefits:

* An exact policy code can surface even when its semantic embedding is average.
* A paraphrase can surface even when no literal terms overlap.
* A record supported by both signals rises without manual score calibration.
* Filtered-out, deleted, or stale sparse versions are removed before final results are returned.

The final result's score is useful for order and a mode-specific threshold, but it is not a probability of truth. Evaluate thresholds per index using real accepted and rejected results.

## Prepare records for hybrid search

Write one dense vector for the semantic content and one concise `sparse_text` representation for exact matching. For a product, include title, brand, SKU, category, important attributes, and user-visible terminology. For a document chunk, include its extracted text and meaningful headings. Avoid serializing opaque JSON, duplicated boilerplate, or credentials into `sparse_text`; it harms lexical quality and may expose data.

Keep useful metadata outside `sparse_text`. Metadata belongs in filters; content that should be matched belongs in the lexical text; text users should see as evidence belongs in a returned snippet or source representation.

## Filtering and security

The same metadata filter is applied to both candidate paths. Always include the tenant, principal, visibility, or product boundary required by the caller before a hybrid query reaches Talqora. Hybrid ranking improves retrieval quality; it does not relax authorization.

For example, a global support index can contain many customers only if every record has a durable `tenant_id` and the service injects the authenticated tenant filter on every request. If hard boundaries are preferable, use one index per tenant and scope keys accordingly.

## Tune with evaluation, not intuition

Build a small set of representative queries with expected sources. Include paraphrases, exact identifiers, ambiguous requests, stale terms, and no-answer queries. Compare dense, sparse, and hybrid modes at the same `top_k` and filter policy. Track recall of expected sources, empty-result behavior, latency, queried transfer, and user acceptance.

If hybrid adds irrelevant literal matches, improve `sparse_text`, enforce a stronger filter, or adjust your application threshold. If exact IDs are missing, verify those IDs are present in `sparse_text` and that the record is not deleted or superseded. Use double quotes around a literal phrase; without quotes, sparse terms are intentionally broad. If semantic results are weak, evaluate the embedding model and source chunking before increasing result counts.

## Latency observability

The response includes `latency_ms` for the request and per-phase timings: `dense_latency_ms`, `sparse_latency_ms`, and `fusion_latency_ms`. Collect p50 and p95 separately by region and filter shape. A hybrid request necessarily includes two retrieval paths, so a healthy sparse p95 alone is not a hybrid SLO. Use these fields to isolate whether vector retrieval, lexical retrieval, or merge/snippet work dominates before changing a threshold or increasing concurrency.

## When not to use it

Use dense-only retrieval when records have no meaningful lexical text or the query is already a trusted embedding. Use sparse-only retrieval for exact lookup interfaces where semantic expansion would be surprising. Use hybrid by default for human language search where both conceptual similarity and literal evidence are valuable.

See [Search and filters](/build-retrieval/search-and-filters) for request construction and [Sparse retrieval architecture](/build-retrieval/sparse-retrieval-architecture) for versioning and lexical consistency.