> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.talqora.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.talqora.com/_mcp/server.

# Browser Markdown Fetch

> Fetch rendered public or authenticated pages as Markdown through an admin-approved Browserbase context.

Browser Markdown Fetch retrieves one or more rendered pages as Markdown without exposing a website password, browser cookie, or Browserbase credential to Talqora users. It is intended for a small, explicit set of pages that your workspace is authorized to access; it is not a web crawler and it does not index a page automatically.

## Tenant approval and access

This capability is disabled for every workspace by default. A **Talqora administrator must approve it for each tenant** in the admin console before an organization owner can create a browser context or fetch a page. The API returns `402` while it is disabled. Members and API keys cannot create contexts, invoke fetches, or receive a Live View URL.

An owner creates a Context for one or more exact allowed hostnames. Talqora stores only Browserbase's opaque context identifier. The browser profile, cookies, local storage, and authenticated session state stay encrypted in Browserbase; Talqora does **not** accept, store, log, or return a website password, cookie, or token.

Contexts are tenant-isolated and can only be used by the organization that created them. A context can only render the exact hostnames declared at creation. Delete a context to revoke the Browserbase profile and its stored authentication state.

## Authenticate through Live View

Use a dashboard session and create a Context. The response contains a short-lived Browserbase Live View URL. Open it, navigate to the permitted site, and complete the login yourself, including MFA. Finish the session after login so its Context persists the new state.

```bash
curl --fail-with-body -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/browser-fetch/contexts" \
  -H "Authorization: Bearer $SUPABASE_ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "label":"Support portal",
    "allowed_hosts":["support.example.com"]
  }'
```

```json
{
  "id":"bctx_…",
  "label":"Support portal",
  "allowed_hosts":["support.example.com"],
  "status":"connecting",
  "live_view_url":"https://…",
  "session_id":"…"
}
```

While this login session is open the Context is `connecting` and cannot be used for a fetch. When login is complete, release the Live View session. This persists the Browserbase Context, changes it to `ready`, and stops its billed browser time:

```bash
curl --fail-with-body -X POST \
  "https://api.talqora.com/v1/indexes/$INDEX_ID/browser-fetch/contexts/$CONTEXT_ID/sessions/$SESSION_ID/complete" \
  -H "Authorization: Bearer $SUPABASE_ACCESS_TOKEN"
```

The Browserbase session and the durable Context are different things: the session is an ephemeral remote Chromium instance; the Context contains the returning login state. Browserbase may invalidate a site login, so your application should treat a rendered login page as a re-authentication signal and open a new Live View rather than trying to submit credentials programmatically.

## Fetch rendered Markdown

Fetch up to ten URLs in one request. With `context_id`, every URL must exactly match one of the Context's allowed hosts. Without one, Talqora creates an ephemeral Browserbase Context for the public render and deletes it afterward.

```bash
curl --fail-with-body -X POST "https://api.talqora.com/v1/indexes/$INDEX_ID/browser-fetch" \
  -H "Authorization: Bearer $SUPABASE_ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "context_id":"bctx_…",
    "urls":[
      "https://support.example.com/articles/getting-started",
      "https://support.example.com/articles/billing"
    ]
  }'
```

Each result includes `url`, `final_url`, `title`, `status_code`, and `markdown`. A failed page is represented in that page's result with an `error`; other URLs in the same request can still succeed. The response body is capped at 1 MB of Markdown per page.

Talqora validates HTTPS destinations before navigation, rejects private, loopback, reserved, URL-credential, and non-standard-port targets, and never follows an authenticated Context to a hostname outside its allowlist. This protects tenant browser sessions from SSRF and cross-site credential exposure.

## Context lifecycle

```bash
# List contexts owned by the active tenant.
curl --fail-with-body "https://api.talqora.com/v1/indexes/$INDEX_ID/browser-fetch/contexts" \
  -H "Authorization: Bearer $SUPABASE_ACCESS_TOKEN"

# Permanently revoke a Context and remove it from Browserbase.
curl --fail-with-body -X DELETE \
  "https://api.talqora.com/v1/indexes/$INDEX_ID/browser-fetch/contexts/$CONTEXT_ID" \
  -H "Authorization: Bearer $SUPABASE_ACCESS_TOKEN"
```

## Operating guidance

* Use one Context per tenant, website, and login. Do not share a Context between tenants or use two browser sessions against the same Context concurrently.
* Keep the hostname allowlist narrow. A Context is a sensitive browser identity, not a general-purpose proxy.
* Browser Markdown Fetch is a retrieval utility. If fetched content should become searchable, review it first, then upload or import the approved Markdown through Talqora File Processing.
* Browserbase browser regions are currently limited to the regions Browserbase publishes. Talqora uses its configured Browserbase region; this is independent from the immutable AWS region of your Talqora index.
* Respect each site's terms, consent requirements, robots policy where applicable, and access controls. The feature does not bypass CAPTCHA, MFA, rate limits, or bot protection.