Skip to navigation

File processing jobs

Talqora turns unstructured files into searchable records without requiring you to build an extraction or embedding pipeline. Files are processed asynchronously, so your application never holds an HTTP request open while a document is being read.

The public API accepts Talqora API keys. A read key can list jobs; a write key can create, complete, replace, and delete them. Dashboard sessions can use the same endpoints for the console. Worker endpoints under /v1/internal/ are not public.

Lifecycle

  1. Create a job with POST /v1/indexes/{index_id}/files.
  2. Upload the original bytes directly to the one-time upload_url, sending every returned upload_headers entry unchanged.
  3. Call POST /v1/indexes/{index_id}/files/{job_id}/complete after the upload succeeds. Optionally provide custom metadata; it is copied to every generated chunk.
  4. Poll GET /v1/indexes/{index_id}/files until the job reaches a terminal state.

Jobs move through uploading, queued, processing, completed, completed_with_warnings, and failed. A replacement can also mark the prior job superseded; deleting a source marks it deleted.

completed_with_warnings means the pipeline finished but produced no indexable evidence, for example an empty PDF or a source whose OCR produced no usable text. It always reports vectors_written: 0, error_code: "no_extractable_content", and a warning. Treat it as a source-quality outcome, not as a successful retrieval import.

Create and upload

For the shortest SDK flow, omit checksum_sha256; Talqora still validates the uploaded file type and size at completion. Provide a checksum only when your application needs an additional source-integrity assertion.

job = client.files.create_upload(INDEX_ID, filename="call.wav", content_type="audio/wav")
with open("call.wav", "rb") as source:
requests.put(job.upload_url, data=source, headers=job.upload_headers).raise_for_status()
client.files.complete_upload(INDEX_ID, job.id, request=FileUploadComplete(metadata={"tenant": 324}))

Import directly from cloud storage

When the source already lives in cloud storage, submit its direct HTTPS object URL instead of downloading it through your application. Talqora immediately returns a durable file job and the serverless worker downloads, validates, stages, and processes the object asynchronously.

curl --fail-with-body -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/imports" \
-H "Authorization: Bearer $TALQORA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"source_url":"https://bucket.s3.us-east-1.amazonaws.com/calls/recording.wav?X-Amz-Signature=...","metadata":{"tenant":324}}'

Supported providers are Amazon S3, Google Cloud Storage, Azure Blob Storage, Cloudflare R2 (r2.dev or the S3 API hostname), Supabase Storage, DigitalOcean Spaces, Backblaze B2, and Wasabi. Supply a public or time-limited signed direct object URL; Talqora does not receive or store your provider credentials.

The URL must remain readable until the job reaches queued. Talqora accepts supported filename extensions only, blocks redirects and private network destinations, caps the download at 20 MiB, and validates the staged bytes before processing. Poll the existing GET /v1/indexes/{index_id}/files endpoint using the returned job ID. Failures are recorded on that job with a stable error code such as remote_import_failed; retry with a fresh signed URL if it expired.

JOB=$(curl -sS -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/files" \
-H "Authorization: Bearer $TALQORA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"filename":"compliance-handbook.pdf","content_type":"application/pdf"}')
UPLOAD_URL=$(echo "$JOB" | jq -r .upload_url)
UPLOAD_CONTENT_TYPE=$(echo "$JOB" | jq -r '.upload_headers["Content-Type"]')
UPLOAD_IF_NONE_MATCH=$(echo "$JOB" | jq -r '.upload_headers["If-None-Match"]')
JOB_ID=$(echo "$JOB" | jq -r .id)
curl --fail-with-body -X PUT "$UPLOAD_URL" \
-H "Content-Type: $UPLOAD_CONTENT_TYPE" \
-H "If-None-Match: $UPLOAD_IF_NONE_MATCH" \
--upload-file compliance-handbook.pdf
curl --fail-with-body -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/files/$JOB_ID/complete" \
-H "Authorization: Bearer $TALQORA_API_KEY"

upload_headers is part of the upload contract, not optional metadata. Each value is covered by the S3 signature; omitting or changing a returned header causes 403 Forbidden. SDK users should pass the mapping directly:

job = client.files.create_upload(
INDEX_ID,
filename="compliance-handbook.pdf",
content_type="application/pdf",
)
with open("compliance-handbook.pdf", "rb") as source:
response = requests.put(job.upload_url, data=source, headers=job.upload_headers)
response.raise_for_status()
client.files.complete_upload(INDEX_ID, job.id)

Attach metadata at completion

The upload URL is scoped to exactly one file job. Complete that job with an optional metadata object to attach structured, filterable values to every vector produced from the file:

curl --fail-with-body -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/files/$JOB_ID/complete" \
-H "Authorization: Bearer $TALQORA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"metadata":{"tenant_id":"acme","document_type":"policy","language":"es"}}'

The metadata is inherited by each chunk and merged with system provenance such as source_file, job_id, page, and chunk. Chunk-level provenance wins if there is a key collision. A completion request is idempotent: once the job is queued, a second completion cannot change its metadata or upload another file. Use a separate job and upload URL for every source file. The URL includes an S3 conditional write (If-None-Match: *), so a second PUT to the same URL is rejected with 412 Precondition Failed once the first object exists.

The signed upload URL is short-lived and only accepts the object created for that job. Do not send your Talqora API key to the upload URL.

File validation and limits

Talqora validates the uploaded bytes, not only the filename or declared MIME type, when complete is called. A .pdf must start with the PDF signature, Office documents must be ZIP containers, and images must have their matching signature. A text blob labelled application/pdf is marked failed immediately with 422 and code: "invalid_pdf"; it is never sent to the processing queue.

The current hard limit is 20 MiB per file and 2,000 PDF pages. This keeps each serverless job within predictable download, extraction, OCR, transcription, and embedding bounds. Split larger source archives before upload. An oversized upload returns 413 file_too_large; a PDF over the page limit becomes failed with pdf_page_limit_exceeded. The upload URL expires after 15 minutes.

You may provide checksum_sha256 when creating a job. Talqora computes SHA-256 over the uploaded object at completion, stores it in the job ledger, and rejects a mismatch with 422 checksum_mismatch.

What processing does

Talqora runs the file through a staged, asynchronous pipeline. Each stage receives a durable task identifier, so a transient failure can be retried without uploading the source again or repeating successful work.

1. Type detection and extraction

The processor first identifies the content type from the declared MIME type and the file signature. It preserves the original object in the ingestion bucket while the job is active. Text-native formats are decoded directly:

  • PDF text and page boundaries are extracted page by page.
  • DOCX and PPTX paragraphs, headings, tables, and slide boundaries are normalized into text blocks.
  • XLSX and CSV rows are converted into searchable records while retaining sheet, row, and column metadata.
  • JSON, Markdown, HTML, XML, and plain text are parsed without treating markup or keys as unrelated documents.

The extractor keeps provenance such as filename, page, sheet, row, section, and source job ID. That provenance is attached to every resulting vector and is what lets Search and Assistant RAG cite the original location.

2. OCR for visual content

When a page or image has no useful text layer, the processor sends the visual content through Talqora’s OCR model. OCR is applied selectively, not to every page: native text is preferred because it is faster, cheaper, and more accurate for digitally generated documents. Scanned PDFs, screenshots, PNGs, JPEGs, and image-only pages take the visual path.

The OCR stage returns text plus page-level structure. The processor normalizes reading order, removes repeated headers and footers when they are clearly boilerplate, preserves tables as stable text, and records the page number and original filename. If the OCR provider returns a transient error, the task uses bounded retries with backoff. A persistent failure marks the task and job as failed with an actionable error rather than silently indexing incomplete content.

3. Audio transcription

Audio files are transcribed server-side with gpt-4o-transcribe, then converted to an indexable Markdown rendition with the original audio filename as provenance. Talqora never sends the API key to the browser and does not expose the provider response directly. Supported formats are FLAC, M4A, MP3, MP4, MPEG, MPGA, OGG, WAV, and WebM. The same 20 MiB upload cap applies to audio.

For pricing, one OCR credit is one image or rendered PDF page sent to vision, and one voice credit is one started minute of audio transcription. Credits are capacity limits, not a promise that every input has extractable evidence: a silent or unintelligible recording can complete with warnings and zero vectors.

4. Normalization and chunking

Extracted text is normalized before embedding: whitespace is cleaned, encoding is made consistent, and sections are kept together where possible. The chunker prefers paragraph, heading, table, page, and row boundaries instead of cutting through a sentence. Long sections are split into bounded chunks with a controlled overlap so a fact at a boundary remains discoverable in both neighboring chunks.

Every chunk receives stable metadata, including source_file, job_id, page or sheet information when available, and its chunk ordinal. This metadata is stored with the vector and can be used by search filters.

5. Embedding generation

Each chunk is converted into a dense embedding using the configured embedding service. File processing requires a 1536-dimensional index, so the generated vector is validated before it can be written. Batches are used to reduce network overhead, while the job ledger records which task and chunk produced each write.

The embedding is a representation of the chunk’s meaning, not a copy of the source text. Similar concepts can therefore match even when the query uses different wording. The original text is retained in the retrieval record and returned as a snippet when available, while the embedding is stored in the index’s regional dense vector store.

6. Sparse indexing and dense writes

The same normalized chunk is also sent to Talqora’s sparse retrieval pipeline with its sparse_text and provenance metadata. BM25 postings are published as immutable retrieval segments. Dense vectors and filterable metadata are written to Talqora’s regional vector storage layer. A hybrid query combines candidates from both paths.

The task is acknowledged only after the dense write, sparse write, and usage ledger update succeed. Writes are idempotent at the chunk level, so retrying a task does not create duplicate retrieval records.

Large documents are planned as independent tasks so failed work can be retried without restarting the entire source. A job only becomes completed after every task has finished and its vectors have been recorded. The returned job reports total_tasks, completed_tasks, vectors_written, pages_total, pages_processed, attempt_count, queued_at, started_at, error_code, and warnings.

Supported source types include PDF, DOCX, XLSX, PPTX, CSV, JSON, Markdown, HTML, XML, plain text, PNG, JPEG, WebP, FLAC, M4A, MP3, MP4, MPEG, MPGA, OGG, WAV, and WebM. File processing currently requires an index with 1536 dimensions.

Poll status

curl --fail-with-body "https://api.talqora.com/v1/indexes/$INDEX_ID/files" \
-H "Authorization: Bearer $TALQORA_API_KEY"

Poll with exponential backoff. queued means Talqora accepted the object; processing means extraction, OCR when needed, chunking, embedding, and indexing are underway. waiting_reason describes the current queue or worker stage. A watchdog changes a job that does not start within five minutes to failed with queue_timeout; a job exceeding one hour becomes failed with processing_timeout. Task retries increment attempt_count, and a final extraction failure exposes an actionable error_code such as invalid_pdf, extraction_failed, or pdf_page_limit_exceeded.

Inspect page ranges and task-level progress with:

curl --fail-with-body "https://api.talqora.com/v1/indexes/$INDEX_ID/files/$JOB_ID/tasks" \
-H "Authorization: Bearer $TALQORA_API_KEY"

Each task reports its page range, state, attempts, chunks written, vectors written, timestamps, and terminal error. Retrieved records include a bounded source-text snippet for review.

Replace a source

Create a replacement job if a source changes. The prior file remains searchable until the replacement is accepted, then its vectors and source object are removed and it becomes superseded.

curl --fail-with-body -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/files/$JOB_ID/replace" \
-H "Authorization: Bearer $TALQORA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"filename":"compliance-handbook-v2.pdf","content_type":"application/pdf"}'

Upload to the returned URL and complete the replacement exactly as in the create flow.

Delete a source

curl --fail-with-body -X DELETE "https://api.talqora.com/v1/indexes/$INDEX_ID/files/$JOB_ID" \
-H "Authorization: Bearer $TALQORA_API_KEY"

Deletion removes the source object, dense vectors, sparse retrieval records, and corresponding usage from the index. It is irreversible.

The job row remains in GET /files with status: "deleted" as an audit record. deleted_vectors is the number of live vectors removed at deletion time; it can be 0 for a failed, empty, or already-superseded source. A deleted job can never be resumed.

Reliability

Creating a job does not start processing. The explicit completion call prevents partially uploaded objects from entering retrieval. If the completion request times out after the object upload succeeded, retry that same completion URL: it is safe while the job is still uploading and becomes a status read once it has been queued.

Operating large imports

Treat each file job as an observable unit of work, not as a synchronous file upload. Large PDFs and presentations are planned into independent page or content tasks so a transient OCR, parser, embedding, or storage failure can be retried without restarting successful work. The job reports total_tasks, completed_tasks, and vectors_written; use those fields to render progress instead of estimating progress from upload byte size alone.

For a bulk import, create jobs with bounded concurrency, upload each source to its returned URL, and call completion only after the upload succeeds. Poll with exponential backoff. Do not create a second job after a completion timeout: retry the existing job’s completion endpoint first. Repeated job creation is a new source lifecycle and can make source provenance ambiguous.

Plan limits apply server-side to processing jobs, multimodal credits, documents, storage, and assistant use. A client must treat 402 and 429 responses as durable product decisions, not as transient failures. Use the index and organization usage views to estimate capacity before enqueueing a large corpus.

Retrieval quality after processing

Processing produces evidence, not a guarantee of answer quality. After an import, test representative dense, sparse, and hybrid queries. Confirm that results include meaningful source metadata such as filename, page, sheet, row, and chunk ordinal. For a scanned document, review OCR-dependent pages before using the corpus in a high-stakes assistant.

Use a relevance threshold in the search or Assistant RAG layer. A completed job means the source is indexed successfully; it does not mean every query has sufficient evidence. Configure Assistant instructions to cite source files and pages, and to say when the index does not contain an answer.

Exact evidence and repeated documents

Sparse retrieval is BM25-style lexical ranking, not a database equality predicate. A bare number such as 129 is weak in a long, repetitive PDF and may retrieve another page that shares surrounding terms. Use a quoted phrase such as "Página 129", combine it with the page metadata filter when the page is known, and set a workload-specific min_score. Quotes are phrase-search syntax; unquoted terms are broad lexical matching. Search responses include snippets so clients can verify the returned evidence before displaying it or passing it to an assistant.

Untrusted source content

Every file-generated chunk is tagged untrusted_source: true. The processor also sets prompt_injection_detected: true when it sees common instruction-override patterns. These flags are retrieval metadata, not a guarantee that all hostile text is detected. Treat retrieved chunks as data, never as system or developer instructions. Production assistants should keep their system prompt outside retrieved context, cite sources, and ignore instructions embedded inside source text.

Source lifecycle design

Use one file job per source whose lifecycle should be independently updateable. For example, upload each policy, contract, product catalog, or spreadsheet as its own source rather than concatenating unrelated documents into a monolith. This lets you replace a new policy version or delete an expired contract without rebuilding the entire index.

Keep a mapping in your application from the source system’s object ID to the Talqora job ID. Store the job ID, source version, index ID, and last successful processing timestamp. That mapping makes it possible to reconcile a changed source, retry a failed import, and demonstrate the provenance of retrieval evidence during an audit.

Crawl a public website

Scale and Enterprise organizations can enqueue a website crawl from the same Processing surface. Crawls are separate jobs from uploaded files: a crawl discovers public pages, extracts Markdown, and submits each page to the same chunking, embedding, dense, and sparse pipeline. Crawl execution is asynchronous and handled by a serverless worker through Cloudflare Browser Rendering. Talqora respects robots.txt and Content Signals; it does not bypass CAPTCHA or bot protection.

Create a crawl job with a starting URL, page limit, and link depth:

curl --fail-with-body -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/crawl-jobs" \
-H "Authorization: Bearer $TALQORA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://docs.example.com","page_limit":25,"max_depth":2}'

The response includes a Talqora crawl job ID and queued status. Poll it with:

curl --fail-with-body "https://api.talqora.com/v1/indexes/$INDEX_ID/crawl-jobs" \
-H "Authorization: Bearer $TALQORA_API_KEY"

When the job is completed, the discovered pages appear as normal processed sources and can be searched with dense, sparse, hybrid, or Assistant RAG retrieval. Use a read key to monitor jobs and a write key to create them. Developer plans receive a 402 response because internet crawling is a paid processing capability.

Failed file jobs retain their original object so they can be retried without another browser upload:

curl --fail-with-body -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/files/$JOB_ID/retry" \
-H "Authorization: Bearer $TALQORA_API_KEY"