> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.talqora.com/build-retrieval/assistant-rag/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.talqora.com/_mcp/server. # Assistant RAG Every index on a **Scale or Enterprise** organization can power multiple Assistant RAGs. Assistant RAG is currently a Beta capability. Developer organizations can see the same console tab, but assistant creation and chat are not enabled until the organization upgrades. Give each assistant its own name, instructions, and chat model, then test grounded answers in the playground. The assistant retrieves from the index when a turn asks for index knowledge. Grounded answers include the source files and pages used as evidence. Conversational turns such as greetings, explicit requests not to search, and assistants with `search_context` disabled skip retrieval and return no sources or retrieval traces. Keep instructions specific: define tone, required citations, what the assistant must refuse, and what it should say when the index does not contain enough evidence. ## Configure assistants Open an index, choose **Assistant RAG**, then select **New assistant**. Write the assistant instructions, choose one of the supported models, and save. Talqora intentionally exposes only this stable set in the console: | Console label | Runtime model | | --------------------- | ---------------------------------------- | | General-purpose model | Select an available model in the console | | gpt-5.6 sol | `gpt-5` | | gpt-5.6 luna | `gpt-5` | | gpt-5.6 terra | `gpt-5` | The three GPT-5.6 labels are Talqora's product names for the GPT-5 runtime route. Legacy, realtime, and retired models are not exposed in the selector. The playground accepts plain-language questions. For a grounded turn, Talqora generates the query embedding, performs hybrid retrieval, applies the configured relevance threshold, and then supplies the selected evidence to the configured chat model. A greeting or explicit search opt-out goes directly to generation without embedding or retrieval. The tab also shows the organization-wide number of AI interactions remaining for the current monthly plan period. ## Design a grounded assistant An assistant is attached to exactly one index. Create separate assistants when their instructions, model choice, intended audience, or access policy differ. For example, a legal assistant can require citations and refuse operational advice, while a support assistant can use a concise customer-facing tone over a different filtered knowledge index. Instructions should state the answer format, citation expectation, uncertainty behavior, audience, and forbidden behavior. They should not attempt to replace retrieval authorization. Restrict the assistant's index and make any tenant or visibility boundary part of the retrieval design before an LLM sees the evidence. Evaluate the assistant with questions that have known evidence, ambiguous questions, exact-term questions, changed documents, and questions the index cannot answer. Measure citation correctness and refusal quality, not only how fluent the answer sounds. ## Threads and traces Chat is stateful when you provide a `thread_id`. The first request can omit it: Talqora creates a `thread_...` identifier and returns it in the JSON response or the terminal SSE event. Pass that value in the next request to continue the same conversation. A new `thread_id` starts a clean conversation. Talqora stores the user and assistant turns, retrieval sources, end-to-end latency, and provider input/output token counts. The console exposes this as **Conversation traces** so an operator can reopen a thread, inspect its input and output token totals, and review the grounded sources associated with each answer. ```json { "message": "Which policy applies to vendor retention?", "thread_id": "thread_01h...", "top_k": 8, "min_score": 0.15 } ``` List conversations for an assistant: ```text GET /v1/indexes//assistants//threads ``` Read one complete trace: ```text GET /v1/indexes//assistants//threads/ ``` These endpoints require the same dashboard session or read-scoped API key as chat. A thread is always isolated to its assistant, index, and organization. Use one `thread_id` per end-user conversation. Persist it in your application session or conversation record and send it again on retry. A newly supplied ID starts an empty conversation; it does not inherit another thread's turns. The trace endpoints are suitable for operator review and debugging, but source content and end-user messages may be sensitive, so apply your own retention and access policy before exposing them in an internal tool. ## Assistant toolkit Each assistant has an explicit MCP toolkit. Configure it in the Assistant RAG tab, then save the assistant. The tools are exposed only through that assistant's MCP endpoint: | Tool | Purpose | | ---------------- | ---------------------------------------------------------------------------- | | `search_context` | Hybrid retrieval over the assistant index, with optional metadata filtering. | | `source_preview` | Read indexed text previews for record IDs returned by `search_context`. | The assistant configuration accepts a `tools` array: ```json { "name": "Compliance assistant", "model": "talqora-default", "system_prompt": "Cite the policy and page for every answer.", "tools": ["search_context", "source_preview"] } ``` Disabled tools are not advertised through MCP and return an unknown-tool error if called. The hosted chat also respects the `search_context` setting: when it is disabled, chat turns do not retrieve from the index or return sources. The toolkit is intentionally narrow. `search_context` is the only tool needed to retrieve grounded evidence, and `source_preview` lets an agent inspect the text associated with an already-retrieved record. Do not use MCP as a broad database-admin interface. Create a dedicated read-only key scoped to the one index and revoke it independently from application keys. ## MCP endpoint Each index exposes a remote MCP endpoint in the Assistant RAG tab: ```text https://api.talqora.com/v1/indexes//assistants//mcp ``` MCP, or Model Context Protocol, is a standard that lets an AI client discover and call tools from a remote service. Talqora exposes a `search_context` tool through Streamable HTTP. Give the MCP client an API key with `read` access scoped only to the target index. The tool accepts: ```json { "query": "Which policy covers vendor data retention?", "top_k": 8, "filter": {"jurisdiction": "SG"} } ``` It returns grounded text snippets, source metadata, and relevance scores. Do not use a dashboard session or a browser-exposed token for MCP. Create a dedicated read-only API key and rotate it independently. ## Chat API Each assistant also has a chat endpoint: ```text POST https://api.talqora.com/v1/indexes//assistants//chat ``` The console uses this endpoint for the playground. It accepts a user message and optional conversation history. Turns that need index evidence retrieve grounded context and return source snippets; conversational or search-disabled turns return an answer without sources. Authenticate with either a dashboard session or a read-only API key scoped to the target index. Keep API keys server-side. ### Server-sent events For interactive applications, use the streaming endpoint: ```text POST https://api.talqora.com/v1/indexes//assistants//chat/stream Accept: text/event-stream Authorization: Bearer Content-Type: application/json ``` The response is an SSE stream. Events arrive in order and are separated by a blank line: ```text data: {"type":"sources","sources":[...]} data: {"type":"delta","text":"The answer"} data: {"type":"delta","text":" continues."} data: {"type":"done","model":"gpt-5","latency_ms":842,"thread_id":"thread_...","input_tokens":781,"output_tokens":132} ``` For grounded turns, `sources` is emitted before answer text so the client can render provenance immediately. Turns that skip retrieval do not emit `sources`, `search_context`, or retrieval-trace events. Multiple `delta` events contain the answer in display order. `done` is the terminal event and includes the runtime model, measured latency, `thread_id`, and token usage. If generation fails, the stream emits an `error` event and closes. Clients should treat a missing `done` event as an interrupted response and retry with the same `thread_id`. ## Production pattern Use the playground to evaluate the assistant, then choose one of two integration paths: * Call your own RAG application through the Vector API when you need full control over prompts, models, and orchestration. * Connect an MCP-capable agent to the Assistant MCP endpoint when the agent should retrieve context as a tool. In both cases, treat metadata filters as authorization boundaries. Retrieval quality does not replace access control. ## Production checklist 1. Process or write a representative corpus and validate source provenance. 2. Create an assistant with clear citation and no-evidence instructions. 3. Use a read-only API key scoped to the target index for server integrations and MCP clients. 4. Persist `thread_id` per conversation and record latency plus token use. 5. Test known-answer, exact-identifier, ambiguous, and no-evidence questions. 6. Review conversation traces for unsupported claims before expanding access. 7. Monitor plan interaction limits, retrieval latency, source quality, and empty-result behavior as traffic grows.