# Scrapfly Documentation

## Table of Contents

### Dashboard

- [Intro](https://scrapfly.io/docs)
- [Project](https://scrapfly.io/docs/project)
- [Account](https://scrapfly.io/docs/account)
- [Workspace & Team](https://scrapfly.io/docs/workspace-and-team)
- [Billing](https://scrapfly.io/docs/billing)

### Products

#### MCP Server

- [Getting Started](https://scrapfly.io/docs/mcp/getting-started)
- [Tools & API Spec](https://scrapfly.io/docs/mcp/tools)
- [Authentication](https://scrapfly.io/docs/mcp/authentication)
- [Examples & Use Cases](https://scrapfly.io/docs/mcp/examples)
- [FAQ](https://scrapfly.io/docs/mcp/faq)
##### Integrations

- [Overview](https://scrapfly.io/docs/mcp/integrations)
- [Claude Desktop](https://scrapfly.io/docs/mcp/integrations/claude-desktop)
- [Claude Code](https://scrapfly.io/docs/mcp/integrations/claude-code)
- [ChatGPT](https://scrapfly.io/docs/mcp/integrations/chatgpt)
- [Cursor](https://scrapfly.io/docs/mcp/integrations/cursor)
- [Cline](https://scrapfly.io/docs/mcp/integrations/cline)
- [Windsurf](https://scrapfly.io/docs/mcp/integrations/windsurf)
- [Zed](https://scrapfly.io/docs/mcp/integrations/zed)
- [Roo Code](https://scrapfly.io/docs/mcp/integrations/roo-code)
- [VS Code](https://scrapfly.io/docs/mcp/integrations/vscode)
- [LangChain](https://scrapfly.io/docs/mcp/integrations/langchain)
- [LlamaIndex](https://scrapfly.io/docs/mcp/integrations/llamaindex)
- [CrewAI](https://scrapfly.io/docs/mcp/integrations/crewai)
- [OpenAI](https://scrapfly.io/docs/mcp/integrations/openai)
- [n8n](https://scrapfly.io/docs/mcp/integrations/n8n)
- [Make](https://scrapfly.io/docs/mcp/integrations/make)
- [Zapier](https://scrapfly.io/docs/mcp/integrations/zapier)
- [Vapi AI](https://scrapfly.io/docs/mcp/integrations/vapi)
- [Agent Builder](https://scrapfly.io/docs/mcp/integrations/agent-builder)
- [Custom Client](https://scrapfly.io/docs/mcp/integrations/custom-client)


#### Web Scraping API

- [Getting Started](https://scrapfly.io/docs/scrape-api/getting-started)
- [API Specification]()
- [Monitoring](https://scrapfly.io/docs/monitoring)
- [Customize Request](https://scrapfly.io/docs/scrape-api/custom)
- [Debug](https://scrapfly.io/docs/scrape-api/debug)
- [Unblocker (formerly ASP)](https://scrapfly.io/docs/scrape-api/unblocker)
- [Proxy](https://scrapfly.io/docs/scrape-api/proxy)
- [Proxy Mode](https://scrapfly.io/docs/scrape-api/proxy-mode)
- [Proxy Mode - Screaming Frog](https://scrapfly.io/docs/scrape-api/proxy-mode/screaming-frog)
- [Proxy Mode - Apify](https://scrapfly.io/docs/scrape-api/proxy-mode/apify)
- [(Auto) Data Extraction](https://scrapfly.io/docs/scrape-api/extraction)
- [Javascript Rendering](https://scrapfly.io/docs/scrape-api/javascript-rendering)
- [Javascript Scenario](https://scrapfly.io/docs/scrape-api/javascript-scenario)
- [SSL](https://scrapfly.io/docs/scrape-api/ssl)
- [DNS](https://scrapfly.io/docs/scrape-api/dns)
- [Cache](https://scrapfly.io/docs/scrape-api/cache)
- [Batch (Multi-URL Scraping)](https://scrapfly.io/docs/scrape-api/batch)
- [Session](https://scrapfly.io/docs/scrape-api/session)
- [Webhook](https://scrapfly.io/docs/scrape-api/webhook)
- [Schedule](https://scrapfly.io/docs/scrape-api/schedule)
- [Screenshot](https://scrapfly.io/docs/scrape-api/screenshot)
- [Errors](https://scrapfly.io/docs/scrape-api/errors)
- [Timeout](https://scrapfly.io/docs/scrape-api/understand-timeout)
- [Throttling](https://scrapfly.io/docs/throttling)
- [Troubleshoot](https://scrapfly.io/docs/scrape-api/troubleshoot)
- [Billing](https://scrapfly.io/docs/scrape-api/billing)
- [FAQ](https://scrapfly.io/docs/scrape-api/faq)

#### Crawler API

- [Getting Started](https://scrapfly.io/docs/crawler-api/getting-started)
- [API Specification]()
- [Retrieving Results](https://scrapfly.io/docs/crawler-api/results)
- [WARC Format](https://scrapfly.io/docs/crawler-api/warc-format)
- [Data Extraction](https://scrapfly.io/docs/crawler-api/extraction-rules)
- [Search](https://scrapfly.io/docs/crawler-api/search)
- [Prompt & Extract](https://scrapfly.io/docs/crawler-api/prompt)
- [Auto Refresh](https://scrapfly.io/docs/crawler-api/refresh)
- [Webhook](https://scrapfly.io/docs/crawler-api/webhook)
- [Schedule](https://scrapfly.io/docs/crawler-api/schedule)
- [Billing](https://scrapfly.io/docs/crawler-api/billing)
- [Errors](https://scrapfly.io/docs/crawler-api/errors)
- [Troubleshoot](https://scrapfly.io/docs/crawler-api/troubleshoot)
- [FAQ](https://scrapfly.io/docs/crawler-api/faq)

#### Screenshot API

- [Getting Started](https://scrapfly.io/docs/screenshot-api/getting-started)
- [API Specification]()
- [Accessibility Testing](https://scrapfly.io/docs/screenshot-api/accessibility)
- [Webhook](https://scrapfly.io/docs/screenshot-api/webhook)
- [Schedule](https://scrapfly.io/docs/screenshot-api/schedule)
- [Billing](https://scrapfly.io/docs/screenshot-api/billing)
- [Errors](https://scrapfly.io/docs/screenshot-api/errors)

#### Extraction API

- [Getting Started](https://scrapfly.io/docs/extraction-api/getting-started)
- [API Specification]()
- [Rules Template](https://scrapfly.io/docs/extraction-api/rules-and-template)
- [LLM Extraction](https://scrapfly.io/docs/extraction-api/llm-prompt)
- [AI Auto Extraction](https://scrapfly.io/docs/extraction-api/automatic-ai)
- [Webhook](https://scrapfly.io/docs/extraction-api/webhook)
- [Billing](https://scrapfly.io/docs/extraction-api/billing)
- [Errors](https://scrapfly.io/docs/extraction-api/errors)
- [FAQ](https://scrapfly.io/docs/extraction-api/faq)

#### Data API


#### Proxy Saver

- [Getting Started](https://scrapfly.io/docs/proxy-saver/getting-started)
- [Fingerprints](https://scrapfly.io/docs/proxy-saver/fingerprints)
- [Optimizations](https://scrapfly.io/docs/proxy-saver/optimizations)
- [SSL Certificates](https://scrapfly.io/docs/proxy-saver/certificates)
- [Protocols](https://scrapfly.io/docs/proxy-saver/protocols)
- [Pacfile](https://scrapfly.io/docs/proxy-saver/pacfile)
- [Secure Credentials](https://scrapfly.io/docs/proxy-saver/security)
- [Billing](https://scrapfly.io/docs/proxy-saver/billing)

#### Cloud Browser API

- [Getting Started](https://scrapfly.io/docs/cloud-browser-api/getting-started)
- [Proxy & Geo-Targeting](https://scrapfly.io/docs/cloud-browser-api/proxy)
- [Unblock API](https://scrapfly.io/docs/cloud-browser-api/unblock)
- [Captcha Solver](https://scrapfly.io/docs/cloud-browser-api/captcha-solver)
- [File Downloads](https://scrapfly.io/docs/cloud-browser-api/file-downloads)
- [Session Resume](https://scrapfly.io/docs/cloud-browser-api/session-resume)
- [Human-in-the-Loop](https://scrapfly.io/docs/cloud-browser-api/human-in-the-loop)
- [Debug Mode](https://scrapfly.io/docs/cloud-browser-api/debug-mode)
- [Browser Extensions](https://scrapfly.io/docs/cloud-browser-api/extensions)
- [Native Browser MCP](https://scrapfly.io/docs/cloud-browser-api/mcp)
- [DevTools Protocol](https://scrapfly.io/docs/cloud-browser-api/cdp-reference)
##### Integrations

- [Puppeteer](https://scrapfly.io/docs/cloud-browser-api/puppeteer)
- [Playwright](https://scrapfly.io/docs/cloud-browser-api/playwright)
- [Selenium](https://scrapfly.io/docs/cloud-browser-api/selenium)
- [Vercel Agent Browser](https://scrapfly.io/docs/cloud-browser-api/agent-browser)
- [Browser Use](https://scrapfly.io/docs/cloud-browser-api/browser-use)
- [Stagehand](https://scrapfly.io/docs/cloud-browser-api/stagehand)

- [Billing](https://scrapfly.io/docs/cloud-browser-api/billing)
- [Errors](https://scrapfly.io/docs/cloud-browser-api/errors)


### Tools

- [Antibot Detector](https://scrapfly.io/docs/tools/antibot-detector)

### SDK

- [Golang](https://scrapfly.io/docs/sdk/golang)
- [Python](https://scrapfly.io/docs/sdk/python)
- [Rust](https://scrapfly.io/docs/sdk/rust)
- [TypeScript](https://scrapfly.io/docs/sdk/typescript)
- [Scrapy](https://scrapfly.io/docs/sdk/scrapy)

### Integrations

- [Getting Started](https://scrapfly.io/docs/integration/getting-started)
- [LangChain](https://scrapfly.io/docs/integration/langchain)
- [LlamaIndex](https://scrapfly.io/docs/integration/llamaindex)
- [CrewAI](https://scrapfly.io/docs/integration/crewai)
- [Zapier](https://scrapfly.io/docs/integration/zapier)
- [Make](https://scrapfly.io/docs/integration/make)
- [n8n](https://scrapfly.io/docs/integration/n8n)

### Academy

- [Overview](https://scrapfly.io/academy)
- [Web Scraping Overview](https://scrapfly.io/academy/scraping-overview)
- [Tools](https://scrapfly.io/academy/tools-overview)
- [Reverse Engineering](https://scrapfly.io/academy/reverse-engineering)
- [Static Scraping](https://scrapfly.io/academy/static-scraping)
- [HTML Parsing](https://scrapfly.io/academy/html-parsing)
- [Dynamic Scraping](https://scrapfly.io/academy/dynamic-scraping)
- [Hidden API Scraping](https://scrapfly.io/academy/hidden-api-scraping)
- [Headless Browsers](https://scrapfly.io/academy/headless-browsers)
- [Hidden Web Data](https://scrapfly.io/academy/hidden-web-data)
- [JSON Parsing](https://scrapfly.io/academy/json-parsing)
- [Data Processing](https://scrapfly.io/academy/data-processing)
- [Scaling](https://scrapfly.io/academy/scaling)
- [Walkthrough Summary](https://scrapfly.io/academy/walkthrough-summary)
- [Scraper Blocking](https://scrapfly.io/academy/scraper-blocking)
- [Proxies](https://scrapfly.io/academy/proxies)

---

# Crawl Search

 A crawl started with `search: true` builds a search index over its own pages while it runs. When the crawl finishes, that index is queryable: you send a natural-language query or a set of exact identifiers, and the API returns ranked passages with the URL they came from and a link back to the full page.

 Search works over **one crawl or a collection of crawls in a single request**. Nothing is created before searching, there is no collection to define, no index to merge and no import step. The crawl *is* the corpus.

> **What enabling search does with your data** `search: true` sends the text of every indexed page to Scrapfly's embedding provider (Google Vertex AI) to compute the vectors. The index also stores the matched passages themselves, as text, next to the crawl in your project and environment, and it is kept for the crawl's retention window and deleted with the crawl. See [Data handling and retention](#data-handling) before enabling it on sensitive targets.

## Enable it at crawl time

 `search` is a boolean on the crawl configuration and defaults to `false`. It is read once, when the crawl is created.

 ```
curl -X POST 'https://api.scrapfly.io/crawl?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -d '{
        "url": "https://web-scraping.dev/products",
        "page_limit": 500,
        "content_formats": ["markdown"],
        "search": true
    }'

```

> **There is no backfill** The index is built from data that only exists while the page is being crawled. A crawl that ran without `search: true`, including every crawl that finished before this feature shipped, has no index and cannot be given one later. There is no retrofit endpoint and none is planned. To search an old crawl, run it again with `search: true`.

 What gets indexed:

- Only pages the crawl actually visited. Failed and skipped URLs are not in the index.
- One text per page, picked in this order: `markdown` if the crawl produced it, else `text`, else `clean_html`, else the raw `html` converted to text. The chosen format is reported per result as `source_format`. Requesting `content_formats: ["markdown"]` gives the best index, because markdown is already stripped of navigation and boilerplate.
- Text is split into overlapping chunks of roughly 800 tokens with 100 tokens of overlap, so a passage that straddles a chunk boundary is still findable. A result is one chunk, not one page.
- Nothing else is indexed. Binary formats (`page_screenshot`, `page_pdf`) and the JSON formats (`extracted_data`, `page_metadata`) never reach the index, so search does not find a page by a field you extracted from it. Each row carries only the chunk text and the handful of attributes the [filters](#filters) work on: URL, host, title, HTTP status, content type and source format.

## Index lifecycle

 The index is built continuously during the crawl and published at the end, after the crawl's own success classification. Its state lives on the crawl, in the `search` block of `GET /crawl/{crawler_uuid}/status`:

 ```
{
  "status": "DONE",
  "is_success": true,
  "search": {
    "status": "READY",
    "documents": 412,
    "vectors": 18432,
    "dropped": 0,
    "index": "IVF_PQ",
    "built_at": "2026-08-25T18:00:00Z",
    "error": null
  }
}

```

 | Status | Meaning | Searchable |
|---|---|---|
| `DISABLED` | No index, and this is not an error. Either the crawl ran without `search: true`, or your plan does not include the feature, or nothing indexable was crawled (for example a crawl that only fetched PDFs). Terminal. | No |
| `BUILDING` | The crawl is still running, or is paused, and rows are still being added. | No, see [a crawl that is still running](#mid-crawl) |
| `READY` | The crawl finished and every visited page made it into the index. Immutable until the crawl refreshes: a crawl with [refresh](https://scrapfly.io/docs/crawler-api/refresh) armed re-embeds changed pages on its next run and republishes the index under a new generation. | Yes |
| `PARTIAL` | The index was published but is known to be incomplete: some documents were dropped while the crawl was running faster than they could be embedded, or the final drain hit its deadline. `dropped` tells you how many. Queries work and the results are correct, they are just drawn from fewer pages than the crawl visited. | Yes |
| `FAILED` | The build could not be published. `error` carries the reason. The crawl itself is unaffected: its pages, artifacts and `/contents` are all intact. | No |

> **The index never breaks the crawl** Indexing runs beside the crawl and is isolated from it. A failed or degraded index never fails the crawl, never slows it down and never changes what `/contents`, WARC or HAR return. That is why `PARTIAL` exists as an outcome instead of the crawl stalling to wait for the embedding provider.

#### Being told when it is ready

 Subscribe to the `crawler_search_ready` and `crawler_search_failed` [webhook events](https://scrapfly.io/docs/crawler-api/webhook#events) instead of polling. They carry the same `search` block shown above. Note that `crawler_search_ready` also fires for a `PARTIAL` index, so read `payload.search.status` rather than assuming `READY`.

#### A paused crawl has no index

 An index is only published for a crawl that reached `DONE`. A paused crawl is resumed automatically and keeps growing, so publishing an index for it would advertise a complete corpus that is still being written. While a crawl is paused its `search.status` stays `BUILDING`, and the collection endpoint skips it with reason `search_not_ready`.

## Endpoints

 | Endpoint | Use |
|---|---|
| `GET /crawl/{crawler_uuid}/search` | One crawl. Query string form, convenient from a browser or a shell. |
| `POST /crawl/search` | One or many crawls. The collection is `crawl_ids` in the request body. |

 The two are the same implementation: the single-crawl route is the collection route with a one-element list, so ranking, filters, paging and the response envelope are identical.

#### Single crawl

 ```
curl -G 'https://api.scrapfly.io/crawl/{crawler_uuid}/search' \
    --data-urlencode 'key={{ YOUR_API_KEY }}' \
    --data-urlencode 'q=TLS fingerprint' \
    --data-urlencode 'limit=20' \
    --data-urlencode 'mode=hybrid'

```

#### A collection of crawls

 ```
curl -X POST 'https://api.scrapfly.io/crawl/search?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -d '{
        "query": "TLS fingerprint",
        "crawl_ids": [
            "0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55",
            "0199a1b0-2c44-7c8e-b3f2-8ad1c9e07731"
        ],
        "limit": 20,
        "mode": "hybrid",
        "filters": {
            "url_prefix": "https://web-scraping.dev/docs/"
        }
    }'

```

 ```
import httpx

resp = httpx.post(
    "https://api.scrapfly.io/crawl/search",
    params={"key": "{{ YOUR_API_KEY }}"},
    json={
        "query": "TLS fingerprint",
        "crawl_ids": ["0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55"],
        "limit": 20,
        "mode": "hybrid",
    },
    timeout=30,
)
payload = resp.json()

for hit in payload["results"]:
    print(hit["rank"], round(hit["score"], 3), hit["url"])
    print(hit["text"][:200])

# Never assume every requested crawl answered.
for skipped in payload["skipped"]:
    print("skipped", skipped["crawler_uuid"], skipped["reason"])

```

## Request

 | Field | Type | Default | Notes |
|---|---|---|---|
| `query` | string | required | The search text. On the single-crawl `GET` route it is the `q` query parameter (`query` is accepted as an alias). |
| `crawl_ids` | string\[\] | required | Crawler UUIDs to search. Up to **256** per request, no duplicates. Body field of `POST /crawl/search` only, the single-crawl route takes the UUID in the path. |
| `limit` | int | `10` | Results per page. **Hard cap of 50**, same as `/contents`. |
| `mode` | enum | `hybrid` | `vector`, `fts` or `hybrid`. See [Modes](#modes). |
| `filters` | object | `{}` | Narrow the corpus before ranking. See [Filters](#filters). |
| `cursor` | string | `null` | Opaque token from a previous response. Sent alone, it returns the next page. See [Paging](#paging). |

#### Limits

- `limit` above 50 returns [`ERR::CRAWLER::CONFIG_ERROR`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::CONFIG_ERROR) with `api_param: "limit"`. In a fan-out, `limit` sizes every per-crawl candidate list as well as the merged one, which is why the ceiling is low.
- More than 256 entries in `crawl_ids` returns [`ERR::CRAWLER::SEARCH_TOO_MANY_CRAWLS`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::SEARCH_TOO_MANY_CRAWLS). Split the collection and merge the pages yourself, or narrow it with `filters.crawler_uuid`.
- An unknown key inside `filters` is rejected, not ignored. A silently dropped filter would widen the data you are looking at without telling you.
- Every UUID is authorized individually against your key's project and environment. A crawl you do not own answers `403`, the same as every other crawl endpoint, and one unauthorized UUID fails the whole request rather than being silently skipped.

## Modes

 Crawled pages are full of exact identifiers that embeddings handle badly: `CVE-2026-12345`, `SKU-891237`, `X-Amz-Credential`, `__cf_bm`. They are also full of prose that keyword search handles badly. Hybrid runs both and fuses the two rankings, which is why it is the default.

 | Mode | What it does | Reach for it when |
|---|---|---|
| `hybrid` | Runs the vector and full-text legs over the same filtered rows and fuses them by rank (reciprocal rank fusion). Default. | You do not know in advance whether the query is prose or an identifier. Start here. |
| `vector` | Semantic similarity only. Finds passages that mean the same thing in other words. | Conceptual questions: "how do they handle refunds", "pages about GDPR consent". |
| `fts` | BM25 full-text only. The tokenizer is tuned for code and identifiers, so `SKU_8912`, long URLs, hashes and header names survive intact, and phrase queries work. | Exact strings: part numbers, error codes, header names, a specific sentence. |

## Filters

 `filters` is a small closed set of keys, not a query language. Filters are pushed down into each crawl and applied **before** that crawl's ranking, so a crawl whose top matches are all excluded still contributes its next-best rows instead of contributing nothing.

 | Key | Type | Matches |
|---|---|---|
| `url_prefix` | string | Exact prefix match on the page URL. |
| `host` | string or string\[\] | Host component of the page URL. Useful when a crawl followed subdomains. Up to 64 entries. |
| `source_format` | enum | `markdown`, `text`, `clean_html` or `html`. |
| `content_type` | string | The content type stored on the row, for example `application/markdown`. |
| `http_status` | int or \[min, max\] | The response status of the page the chunk came from, where the row carries one. |
| `crawler_uuid` | string or string\[\] | Narrows within an already-authorized `crawl_ids`. It cannot widen the request: only the crawls named in `crawl_ids` are ever opened, so a UUID here that is not in `crawl_ids` simply matches nothing. Body form only. |

 ```
{
  "query": "returns policy",
  "crawl_ids": ["0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55"],
  "filters": {
    "url_prefix": "https://web-scraping.dev/docs/",
    "host": ["web-scraping.dev", "blog.web-scraping.dev"],
    "source_format": "markdown",
    "http_status": [200, 299]
  }
}

```

 On the single-crawl `GET` route the same filters are flat query parameters (`url_prefix`, `host`, `source_format`, `content_type`, `http_status`) rather than a nested object, and each takes a single value. `crawler_uuid` is not offered there: the path already names the one crawl.

 There is deliberately no `url_regex`. A regular expression cannot be pushed into the index, so it degrades into a full scan of every crawl in the request. Use `url_prefix` and `host`, or filter the results on your side.

## Response

 ```
{
  "query": "TLS fingerprint",
  "mode": "hybrid",
  "limit": 20,
  "completeness": "exact",
  "crawls": [
    { "crawler_uuid": "0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55", "documents": 412, "vectors": 18432, "index": "IVF_PQ" },
    { "crawler_uuid": "019a77d3-9b02-7f41-8c60-1e4d5a2b7c98", "documents": 96, "vectors": 3110, "index": "FLAT" }
  ],
  "skipped": [
    { "crawler_uuid": "0199a1b0-2c44-7c8e-b3f2-8ad1c9e07731", "reason": "search_not_ready", "status": "BUILDING" }
  ],
  "results": [
    {
      "rank": 1,
      "score": 0.927,
      "scores": { "vector": 0.91, "fts": 12.4, "rrf": 0.0312 },
      "crawler_uuid": "0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55",
      "url": "https://web-scraping.dev/docs/tls",
      "title": "TLS fingerprinting",
      "source_format": "markdown",
      "content_type": "application/markdown",
      "chunk_id": 3,
      "text": "JA3 and JA4 summarise the ClientHello into a short hash ...",
      "warc_offset": 728271,
      "warc_end": 746643,
      "contents_url": "https://api.scrapfly.io/crawl/0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55/contents?url=https%3A%2F%2Fweb-scraping.dev%2Fdocs%2Ftls&formats=markdown"
    }
  ],
  "stats": { "duration_ms": 412, "crawls_searched": 2, "candidates": 150, "gcs_gets": 27 },
  "crawls_requested": 3,
  "crawls_searched": 2,
  "crawls_pruned_exact": 0,
  "crawls_skipped_deadline": [],
  "crawls_failed": [],
  "theta": 0.7412,
  "max_ub_unsearched": null,
  "cursor": "eyJ2IjoxLCJvIjoyMH0"
}

```

 | Field | Meaning |
|---|---|
| `results[].rank` | Position in the merged ranking, starting at 1 and continuing across pages. |
| `results[].score` | The score results are ordered by. Rank on this, not on the individual legs. |
| `results[].scores` | The per-leg scores that produced it: `vector` similarity, `fts` BM25 and the fused `rrf` value. Diagnostic. BM25 in particular is relative to one crawl's corpus statistics, so comparing raw `fts` values across crawls is meaningless. |
| `results[].text` | The matched chunk, not the whole page. It is content written by the crawled site, so escape it before rendering, exactly as you would any scraped content. |
| `results[].chunk_id` | Index of the chunk within its page, so several hits on one page stay distinguishable. |
| `results[].contents_url` | Ready-made link back to the full page. See [From a hit to the page](#expand). |
| `results[].warc_offset`, `warc_end` | Byte range of the source record inside the crawl's WARC, for readers that hold the artifact. |
| `crawls` | The crawls that entered the search with a usable index, with their document and vector counts. Not the same as the crawls that were opened: a pruned crawl is listed here and counted in `crawls_pruned_exact`. `crawls_searched` is the number actually opened. |
| `skipped` | Requested crawls that contributed nothing, each with a reason. Never fatal. |
| `stats` | Wall clock, crawls searched, candidates considered and storage reads for this request. |

#### Why a crawl is skipped

 | Reason | What to do |
|---|---|
| `search_not_enabled` | The crawl ran without `search: true`. Re-crawl with the flag, there is no backfill. |
| `search_not_ready` | The index is still `BUILDING`, which includes a paused crawl. Wait for `crawler_search_ready` ([a running crawl is not searchable](#mid-crawl)). |
| `search_failed` | The build failed. `GET /crawl/{crawler_uuid}/status` carries the error. |
| `search_disabled` | The index is `DISABLED`: nothing indexable was crawled, or the feature is not enabled on the account. |
| `incompatible_index` | That crawl was indexed differently, so its scores cannot be merged with the rest of this request. Search it in its own request. |
| `filtered_out` | A `host` filter excluded that crawl outright: the summary beside its index records which hosts it holds, and none of them matched. Nothing to do, this is the filter working before any storage read. |
| `deadline` | The request budget ran out before the crawl was opened. It is also listed in `crawls_skipped_deadline` and `completeness` is `partial`. Retry, narrow `crawl_ids`, or accept the partial answer. |
| `error` | The crawl was opened and its query failed. It is also listed in `crawls_failed`. Retry the request. |

## The completeness envelope

 Every response states how complete it is. This is the part of the API most worth understanding, because the honest answer looks alarming at first glance:

 ```
{
  "crawls_requested": 214,
  "crawls_searched": 37,
  "crawls_pruned_exact": 171,
  "crawls_skipped_deadline": [],
  "crawls_failed": [],
  "completeness": "exact",
  "theta": 0.7412,
  "max_ub_unsearched": 0.6903
}
```

 You asked for 214 crawls, 37 were opened, and the answer is still `exact`. That is correct, and it is not a shortcut.

 Each crawl publishes a small summary beside its index that bounds the best score any row inside it could possibly achieve for a given query. Reading those summaries costs one batched lookup and no storage reads. The planner works through the crawls cheapest-first, and it keeps `theta`: the score of the worst result currently good enough to be on your page. A crawl whose best possible score is below `theta` cannot put a row on that page, so opening it would cost a round trip and change nothing.

 `max_ub_unsearched` is the highest of those bounds among the crawls that were never opened. When it sits below `theta` (0.6903 against 0.7412 here), the results are provably identical to what an exhaustive search of all 214 crawls would have returned. `crawls_pruned_exact` counts the crawls eliminated that way.

 | `completeness` | Means |
|---|---|
| `exact` | No unopened crawl could have beaten what you got, whatever `crawls_searched` says. This is the normal outcome. |
| `partial` | The deadline arrived before that proof was finished. The results are real and correctly ranked among themselves, but a crawl listed in `crawls_skipped_deadline` might have held something better. `max_ub_unsearched` tells you the best score you could possibly have missed. |

> **Crawls are never silently dropped** Every UUID you send appears somewhere: in `crawls`, in `skipped`, in `crawls_skipped_deadline` or in `crawls_failed`. If you are reconciling coverage, sum those four rather than trusting `results`.

## Paging

 Paging is by cursor, never by offset. The first request materialises a globally ordered candidate pool of up to 500 rows and returns the top `limit`; the `cursor` carries the rest of that ordering. Later pages are pure lookups against the pool, with no second search and therefore no rank drift, no duplicates and no rows that quietly move between pages.

 ```
# Page 2 repeats the request and adds the cursor. query and crawl_ids stay
# required on every page: they are what the request is authorized against.
# The ordering itself lives in the cursor, so the ranking cannot drift.
curl -X POST 'https://api.scrapfly.io/crawl/search?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -d '{
        "query": "TLS fingerprint",
        "crawl_ids": ["0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55"],
        "limit": 20,
        "cursor": "eyJ2IjoxLCJvIjoyMH0"
    }'

```

- `query` and `crawl_ids` stay required on a cursor request, and `crawl_ids` is still authorized crawl by crawl. Resend the request you paged from and add the cursor.
- The cursor wins over everything it overlaps. Changing `query`, `filters` or `mode` alongside it does **not** start a new search, it returns the next page of the ordering the cursor already carries. To search for something else, drop the cursor.
- `cursor` is `null` on the last page.
- Cursors are short-lived. Page through promptly rather than storing one for later.
- The pool is capped at 500 rows. Beyond that, narrow the query or the filters instead of paging deeper.

## From a hit to the page

 A result is a passage. When you need the whole document, follow `contents_url`: it is a ready-made call to the crawl's [`/contents`](https://scrapfly.io/docs/crawler-api/results#query-content) endpoint, already pinned to the matched URL and to the format the chunk was indexed from. Append your API key and fetch it.

 `warc_offset` and `warc_end` are the byte range of the same record inside the crawl's WARC artifact. They are there for pipelines that already hold the artifact and want to read the record locally instead of calling the API. Everyone else should use `contents_url`.

## A crawl that is still running is not searchable yet

 An index only answers queries once it has been published, which happens when the crawl reaches a terminal state. Until then its status is `BUILDING` and both routes treat it the same way: the crawl comes back under `skipped` with reason `search_not_ready` and contributes no rows. On the single-crawl route that means a `200` with an empty `results` list, not an error.

 So do not poll `/search` to watch a long crawl fill up. Wait for `crawler_search_ready`, or poll `search.status` on `GET /crawl/{crawler_uuid}/status`, and start querying once it leaves `BUILDING`. A paused crawl stays `BUILDING` for as long as it is paused.

#### Caching

 A `READY` index over a finished, non-refreshing crawl is immutable, so those responses carry `Cache-Control: private, max-age=3600, immutable`. It is `private`, not `public` like `/contents`: the response carries an `X-Scrapfly-Api-Cost` a shared cache would keep replaying, and authorization for this endpoint can arrive in a header rather than the URL. A crawl with [refresh](https://scrapfly.io/docs/crawler-api/refresh) armed gets `private, max-age=60, must-revalidate` instead, since its next run can change the index. Results that include a `BUILDING` or `PARTIAL` crawl are returned `no-store`, so a half-built index's answer is never cached as if it were final.

## Data handling and retention

- **Page text leaves the crawl to be embedded.** With `search: true`, the text of every indexed page is sent to Scrapfly's embedding provider, Google Vertex AI. That is the whole point of the flag, and it is the reason it is opt-in per crawl rather than an account setting.
- **The index stores the passages themselves.** Alongside the vectors, each row keeps its chunk of text, so a search result can return the matched passage without re-reading the crawl's WARC. The index lives under the same user, project and environment prefix as the crawl it belongs to, and is subject to the same access control.
- **It lives exactly as long as the crawl.** The index inherits the crawl's retention window and is deleted with the crawl, whether that deletion comes from your plan's retention policy or from an explicit delete. There is no separate index retention to manage and no separate deletion to request.
- **The embedding is frozen per crawl.** How a crawl was indexed is recorded when the index is built and never changes afterwards. Crawls indexed differently cannot be ranked together, which surfaces as `incompatible_index`.
- **No backfill.** Crawls that predate this feature, or that ran without the flag, have no index and cannot be given one.

## Errors

 Only the request as a whole can fail. Anything wrong with one crawl inside an otherwise valid request is reported as a [`skipped` entry on a `200`](#skipped), on the single-crawl route as much as on the collection one. Branch on `skipped`, not on a status code, for "this crawl had no index".

 | Code | Cause |
|---|---|
| [`ERR::CRAWLER::CONFIG_ERROR`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::CONFIG_ERROR) | Empty `query`, empty or duplicated `crawl_ids`, `limit` outside 1-50, unknown `mode`, unknown filter key, malformed filter value or an unreadable `cursor`. `api_param` names the offending field. |
| [`ERR::CRAWLER::SEARCH_TOO_MANY_CRAWLS`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::SEARCH_TOO_MANY_CRAWLS) | `crawl_ids` exceeded 256 entries. |
| `403` | One of the UUIDs is not yours, or not in this key's project and environment. Authorization is all-or-nothing: one unauthorized UUID fails the whole request instead of being skipped. |

 The crawler error catalog also carries [`ERR::CRAWLER::SEARCH_NOT_ENABLED`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::SEARCH_NOT_ENABLED), [`ERR::CRAWLER::SEARCH_NOT_READY`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::SEARCH_NOT_READY), [`ERR::CRAWLER::SEARCH_INCOMPATIBLE_INDEX`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::SEARCH_INCOMPATIBLE_INDEX) and [`ERR::CRAWLER::SEARCH_DISABLED_BY_COMPLIANCE`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::SEARCH_DISABLED_BY_COMPLIANCE) for the corresponding per-crawl conditions. A search request does not return them today, it reports those crawls under `skipped` instead. Handle the reason strings and you are covered either way. The full catalog, with the HTTP status and retry guidance for each code, is on the [Errors](https://scrapfly.io/docs/crawler-api/errors) page.

## Next steps

- Ask questions over the same index with [Prompt &amp; Extract](https://scrapfly.io/docs/crawler-api/prompt).
- Subscribe to [`crawler_search_ready`](https://scrapfly.io/docs/crawler-api/webhook#events) instead of polling for readiness.
- Fetch whole pages behind a hit with [`/contents`](https://scrapfly.io/docs/crawler-api/results#query-content).
- Understand what indexing adds to a crawl's bill on the [Billing](https://scrapfly.io/docs/crawler-api/billing#search) page.
