Agentic and regex search
Talqora has two specialized retrieval modes in addition to dense, sparse, and hybrid search. Both return the same evidence objects, filters, snippets, and source metadata as the standard query endpoint.
Literal exact search
Set search_type to exact to match one literal token without BM25 tokenization. This is the mode for a copied email address, error code, UUID, URL, hash, or version identifier. Punctuation is significant.
Talqora keeps two lexical representations for newly written sparse_text: normalized text for BM25 and literal identifier terms for exact/regex matching. Re-upsert or reprocess records written before this capability to add literal terms; the original source data and vector IDs remain unchanged.
Regex literal search
Set search_type to regex and send the pattern in sparse_query. This runs against a bounded representation of the original indexed text, rather than BM25 tokens, and is appropriate for emails, error codes, document identifiers, URLs, UUIDs, hashes, version strings, and carefully bounded textual patterns. It is not a replacement for semantic retrieval.
Regex is a case-insensitive substring matcher: anchors (^, $), word boundaries, Unicode properties, lookarounds, inline flags, backreferences, empty alternatives, match-all patterns, and nested quantifiers are rejected with 422. Every regex must contain at least three consecutive literal letters or digits. Snippets are untrusted plain text; render them as text, never HTML.
Patterns are limited to 256 characters. Lookarounds, inline flags, and backreferences are rejected with 422 so a caller cannot accidentally submit an expensive or ambiguous pattern. Use sparse search for normal full-text terms and quoted sparse queries for exact phrases.
Agentic retrieval
Set search_type to agentic and send an agentic_query. Talqora asks its retrieval planner to turn the question into a small set of independent retrieval queries. Each query runs hybrid retrieval concurrently. The API merges those result sets with reciprocal-rank fusion and returns the plan in agentic_plan for observability.
Agentic retrieval is available on Scale and Enterprise plans. It has hard server-side limits: the input is at most 2,000 characters, the planner can create at most four subqueries, and top_k is at most 10. These limits cap model use, fan-out, transfer, and latency. Developer-plan requests receive 402 with code: "agentic_retrieval_requires_scale" and an upgrade_url that a client can present directly. The planner never writes data, changes filters, or follows retrieved instructions; retrieved source text remains untrusted evidence.
planner_latency_ms measures planning only. latency_ms is the full operation including all parallel retrieval steps. Use normal hybrid search for high-throughput request paths and reserve agentic retrieval for ambiguous, multi-part questions where a retrieval plan materially improves recall.