Assistant RAG
Every index on a Scale or Enterprise organization can power multiple Assistant RAGs. Assistant RAG is currently a Beta capability. Developer organizations can see the same console tab, but assistant creation and chat are not enabled until the organization upgrades. Give each assistant its own name, instructions, and chat model, then test grounded answers in the playground.
The assistant retrieves from the index when a turn asks for index knowledge. Grounded answers include the source files and pages used as evidence. Conversational turns such as greetings, explicit requests not to search, and assistants with search_context disabled skip retrieval and return no sources or retrieval traces. Keep instructions specific: define tone, required citations, what the assistant must refuse, and what it should say when the index does not contain enough evidence.
Configure assistants
Open an index, choose Assistant RAG, then select New assistant. Write the assistant instructions, choose one of the supported models, and save. Talqora intentionally exposes only this stable set in the console:
The three GPT-5.6 labels are Talqora’s product names for the GPT-5 runtime route. Legacy, realtime, and retired models are not exposed in the selector.
The playground accepts plain-language questions. For a grounded turn, Talqora generates the query embedding, performs hybrid retrieval, applies the configured relevance threshold, and then supplies the selected evidence to the configured chat model. A greeting or explicit search opt-out goes directly to generation without embedding or retrieval. The tab also shows the organization-wide number of AI interactions remaining for the current monthly plan period.
Design a grounded assistant
An assistant is attached to exactly one index. Create separate assistants when their instructions, model choice, intended audience, or access policy differ. For example, a legal assistant can require citations and refuse operational advice, while a support assistant can use a concise customer-facing tone over a different filtered knowledge index.
Instructions should state the answer format, citation expectation, uncertainty behavior, audience, and forbidden behavior. They should not attempt to replace retrieval authorization. Restrict the assistant’s index and make any tenant or visibility boundary part of the retrieval design before an LLM sees the evidence.
Evaluate the assistant with questions that have known evidence, ambiguous questions, exact-term questions, changed documents, and questions the index cannot answer. Measure citation correctness and refusal quality, not only how fluent the answer sounds.
Threads and traces
Chat is stateful when you provide a thread_id. The first request can omit it: Talqora creates a thread_... identifier and returns it in the JSON response or the terminal SSE event. Pass that value in the next request to continue the same conversation. A new thread_id starts a clean conversation.
Talqora stores the user and assistant turns, retrieval sources, end-to-end latency, and provider input/output token counts. The console exposes this as Conversation traces so an operator can reopen a thread, inspect its input and output token totals, and review the grounded sources associated with each answer.
List conversations for an assistant:
Read one complete trace:
These endpoints require the same dashboard session or read-scoped API key as chat. A thread is always isolated to its assistant, index, and organization.
Use one thread_id per end-user conversation. Persist it in your application
session or conversation record and send it again on retry. A newly supplied ID
starts an empty conversation; it does not inherit another thread’s turns. The
trace endpoints are suitable for operator review and debugging, but source
content and end-user messages may be sensitive, so apply your own retention
and access policy before exposing them in an internal tool.
Assistant toolkit
Each assistant has an explicit MCP toolkit. Configure it in the Assistant RAG tab, then save the assistant. The tools are exposed only through that assistant’s MCP endpoint:
The assistant configuration accepts a tools array:
Disabled tools are not advertised through MCP and return an unknown-tool error if called. The hosted chat also respects the search_context setting: when it is disabled, chat turns do not retrieve from the index or return sources.
The toolkit is intentionally narrow. search_context is the only tool needed
to retrieve grounded evidence, and source_preview lets an agent inspect the
text associated with an already-retrieved record. Do not use MCP as a broad
database-admin interface. Create a dedicated read-only key scoped to the one
index and revoke it independently from application keys.
MCP endpoint
Each index exposes a remote MCP endpoint in the Assistant RAG tab:
MCP, or Model Context Protocol, is a standard that lets an AI client discover and call tools from a remote service. Talqora exposes a search_context tool through Streamable HTTP. Give the MCP client an API key with read access scoped only to the target index.
The tool accepts:
It returns grounded text snippets, source metadata, and relevance scores. Do not use a dashboard session or a browser-exposed token for MCP. Create a dedicated read-only API key and rotate it independently.
Chat API
Each assistant also has a chat endpoint:
The console uses this endpoint for the playground. It accepts a user message and optional conversation history. Turns that need index evidence retrieve grounded context and return source snippets; conversational or search-disabled turns return an answer without sources. Authenticate with either a dashboard session or a read-only API key scoped to the target index. Keep API keys server-side.
Server-sent events
For interactive applications, use the streaming endpoint:
The response is an SSE stream. Events arrive in order and are separated by a blank line:
For grounded turns, sources is emitted before answer text so the client can render provenance immediately. Turns that skip retrieval do not emit sources, search_context, or retrieval-trace events. Multiple delta events contain the answer in display order. done is the terminal event and includes the runtime model, measured latency, thread_id, and token usage. If generation fails, the stream emits an error event and closes. Clients should treat a missing done event as an interrupted response and retry with the same thread_id.
Production pattern
Use the playground to evaluate the assistant, then choose one of two integration paths:
- Call your own RAG application through the Vector API when you need full control over prompts, models, and orchestration.
- Connect an MCP-capable agent to the Assistant MCP endpoint when the agent should retrieve context as a tool.
In both cases, treat metadata filters as authorization boundaries. Retrieval quality does not replace access control.
Production checklist
- Process or write a representative corpus and validate source provenance.
- Create an assistant with clear citation and no-evidence instructions.
- Use a read-only API key scoped to the target index for server integrations and MCP clients.
- Persist
thread_idper conversation and record latency plus token use. - Test known-answer, exact-identifier, ambiguous, and no-evidence questions.
- Review conversation traces for unsupported claims before expanding access.
- Monitor plan interaction limits, retrieval latency, source quality, and empty-result behavior as traffic grows.