# Scrapfly Documentation

## Table of Contents

### Dashboard

- [Intro](https://scrapfly.io/docs)
- [Project](https://scrapfly.io/docs/project)
- [Account](https://scrapfly.io/docs/account)
- [Workspace & Team](https://scrapfly.io/docs/workspace-and-team)
- [Billing](https://scrapfly.io/docs/billing)

### Products

#### MCP Server

- [Getting Started](https://scrapfly.io/docs/mcp/getting-started)
- [Tools & API Spec](https://scrapfly.io/docs/mcp/tools)
- [Authentication](https://scrapfly.io/docs/mcp/authentication)
- [Examples & Use Cases](https://scrapfly.io/docs/mcp/examples)
- [FAQ](https://scrapfly.io/docs/mcp/faq)
##### Integrations

- [Overview](https://scrapfly.io/docs/mcp/integrations)
- [Claude Desktop](https://scrapfly.io/docs/mcp/integrations/claude-desktop)
- [Claude Code](https://scrapfly.io/docs/mcp/integrations/claude-code)
- [ChatGPT](https://scrapfly.io/docs/mcp/integrations/chatgpt)
- [Cursor](https://scrapfly.io/docs/mcp/integrations/cursor)
- [Cline](https://scrapfly.io/docs/mcp/integrations/cline)
- [Windsurf](https://scrapfly.io/docs/mcp/integrations/windsurf)
- [Zed](https://scrapfly.io/docs/mcp/integrations/zed)
- [Roo Code](https://scrapfly.io/docs/mcp/integrations/roo-code)
- [VS Code](https://scrapfly.io/docs/mcp/integrations/vscode)
- [LangChain](https://scrapfly.io/docs/mcp/integrations/langchain)
- [LlamaIndex](https://scrapfly.io/docs/mcp/integrations/llamaindex)
- [CrewAI](https://scrapfly.io/docs/mcp/integrations/crewai)
- [OpenAI](https://scrapfly.io/docs/mcp/integrations/openai)
- [n8n](https://scrapfly.io/docs/mcp/integrations/n8n)
- [Make](https://scrapfly.io/docs/mcp/integrations/make)
- [Zapier](https://scrapfly.io/docs/mcp/integrations/zapier)
- [Vapi AI](https://scrapfly.io/docs/mcp/integrations/vapi)
- [Agent Builder](https://scrapfly.io/docs/mcp/integrations/agent-builder)
- [Custom Client](https://scrapfly.io/docs/mcp/integrations/custom-client)


#### Web Scraping API

- [Getting Started](https://scrapfly.io/docs/scrape-api/getting-started)
- [API Specification]()
- [Monitoring](https://scrapfly.io/docs/monitoring)
- [Customize Request](https://scrapfly.io/docs/scrape-api/custom)
- [Debug](https://scrapfly.io/docs/scrape-api/debug)
- [Unblocker (formerly ASP)](https://scrapfly.io/docs/scrape-api/unblocker)
- [Proxy](https://scrapfly.io/docs/scrape-api/proxy)
- [Proxy Mode](https://scrapfly.io/docs/scrape-api/proxy-mode)
- [Proxy Mode - Screaming Frog](https://scrapfly.io/docs/scrape-api/proxy-mode/screaming-frog)
- [Proxy Mode - Apify](https://scrapfly.io/docs/scrape-api/proxy-mode/apify)
- [(Auto) Data Extraction](https://scrapfly.io/docs/scrape-api/extraction)
- [Javascript Rendering](https://scrapfly.io/docs/scrape-api/javascript-rendering)
- [Javascript Scenario](https://scrapfly.io/docs/scrape-api/javascript-scenario)
- [SSL](https://scrapfly.io/docs/scrape-api/ssl)
- [DNS](https://scrapfly.io/docs/scrape-api/dns)
- [Cache](https://scrapfly.io/docs/scrape-api/cache)
- [Batch (Multi-URL Scraping)](https://scrapfly.io/docs/scrape-api/batch)
- [Session](https://scrapfly.io/docs/scrape-api/session)
- [Webhook](https://scrapfly.io/docs/scrape-api/webhook)
- [Schedule](https://scrapfly.io/docs/scrape-api/schedule)
- [Screenshot](https://scrapfly.io/docs/scrape-api/screenshot)
- [Errors](https://scrapfly.io/docs/scrape-api/errors)
- [Timeout](https://scrapfly.io/docs/scrape-api/understand-timeout)
- [Throttling](https://scrapfly.io/docs/throttling)
- [Troubleshoot](https://scrapfly.io/docs/scrape-api/troubleshoot)
- [Billing](https://scrapfly.io/docs/scrape-api/billing)
- [FAQ](https://scrapfly.io/docs/scrape-api/faq)

#### Crawler API

- [Getting Started](https://scrapfly.io/docs/crawler-api/getting-started)
- [API Specification]()
- [Retrieving Results](https://scrapfly.io/docs/crawler-api/results)
- [WARC Format](https://scrapfly.io/docs/crawler-api/warc-format)
- [Data Extraction](https://scrapfly.io/docs/crawler-api/extraction-rules)
- [Search](https://scrapfly.io/docs/crawler-api/search)
- [Prompt & Extract](https://scrapfly.io/docs/crawler-api/prompt)
- [Auto Refresh](https://scrapfly.io/docs/crawler-api/refresh)
- [Webhook](https://scrapfly.io/docs/crawler-api/webhook)
- [Schedule](https://scrapfly.io/docs/crawler-api/schedule)
- [Billing](https://scrapfly.io/docs/crawler-api/billing)
- [Errors](https://scrapfly.io/docs/crawler-api/errors)
- [Troubleshoot](https://scrapfly.io/docs/crawler-api/troubleshoot)
- [FAQ](https://scrapfly.io/docs/crawler-api/faq)

#### Screenshot API

- [Getting Started](https://scrapfly.io/docs/screenshot-api/getting-started)
- [API Specification]()
- [Accessibility Testing](https://scrapfly.io/docs/screenshot-api/accessibility)
- [Webhook](https://scrapfly.io/docs/screenshot-api/webhook)
- [Schedule](https://scrapfly.io/docs/screenshot-api/schedule)
- [Billing](https://scrapfly.io/docs/screenshot-api/billing)
- [Errors](https://scrapfly.io/docs/screenshot-api/errors)

#### Extraction API

- [Getting Started](https://scrapfly.io/docs/extraction-api/getting-started)
- [API Specification]()
- [Rules Template](https://scrapfly.io/docs/extraction-api/rules-and-template)
- [LLM Extraction](https://scrapfly.io/docs/extraction-api/llm-prompt)
- [AI Auto Extraction](https://scrapfly.io/docs/extraction-api/automatic-ai)
- [Webhook](https://scrapfly.io/docs/extraction-api/webhook)
- [Billing](https://scrapfly.io/docs/extraction-api/billing)
- [Errors](https://scrapfly.io/docs/extraction-api/errors)
- [FAQ](https://scrapfly.io/docs/extraction-api/faq)

#### Data API


#### Proxy Saver

- [Getting Started](https://scrapfly.io/docs/proxy-saver/getting-started)
- [Fingerprints](https://scrapfly.io/docs/proxy-saver/fingerprints)
- [Optimizations](https://scrapfly.io/docs/proxy-saver/optimizations)
- [SSL Certificates](https://scrapfly.io/docs/proxy-saver/certificates)
- [Protocols](https://scrapfly.io/docs/proxy-saver/protocols)
- [Pacfile](https://scrapfly.io/docs/proxy-saver/pacfile)
- [Secure Credentials](https://scrapfly.io/docs/proxy-saver/security)
- [Billing](https://scrapfly.io/docs/proxy-saver/billing)

#### Cloud Browser API

- [Getting Started](https://scrapfly.io/docs/cloud-browser-api/getting-started)
- [Proxy & Geo-Targeting](https://scrapfly.io/docs/cloud-browser-api/proxy)
- [Unblock API](https://scrapfly.io/docs/cloud-browser-api/unblock)
- [Captcha Solver](https://scrapfly.io/docs/cloud-browser-api/captcha-solver)
- [File Downloads](https://scrapfly.io/docs/cloud-browser-api/file-downloads)
- [Session Resume](https://scrapfly.io/docs/cloud-browser-api/session-resume)
- [Human-in-the-Loop](https://scrapfly.io/docs/cloud-browser-api/human-in-the-loop)
- [Debug Mode](https://scrapfly.io/docs/cloud-browser-api/debug-mode)
- [Browser Extensions](https://scrapfly.io/docs/cloud-browser-api/extensions)
- [Native Browser MCP](https://scrapfly.io/docs/cloud-browser-api/mcp)
- [DevTools Protocol](https://scrapfly.io/docs/cloud-browser-api/cdp-reference)
##### Integrations

- [Puppeteer](https://scrapfly.io/docs/cloud-browser-api/puppeteer)
- [Playwright](https://scrapfly.io/docs/cloud-browser-api/playwright)
- [Selenium](https://scrapfly.io/docs/cloud-browser-api/selenium)
- [Vercel Agent Browser](https://scrapfly.io/docs/cloud-browser-api/agent-browser)
- [Browser Use](https://scrapfly.io/docs/cloud-browser-api/browser-use)
- [Stagehand](https://scrapfly.io/docs/cloud-browser-api/stagehand)

- [Billing](https://scrapfly.io/docs/cloud-browser-api/billing)
- [Errors](https://scrapfly.io/docs/cloud-browser-api/errors)


### Tools

- [Antibot Detector](https://scrapfly.io/docs/tools/antibot-detector)

### SDK

- [Golang](https://scrapfly.io/docs/sdk/golang)
- [Python](https://scrapfly.io/docs/sdk/python)
- [Rust](https://scrapfly.io/docs/sdk/rust)
- [TypeScript](https://scrapfly.io/docs/sdk/typescript)
- [Scrapy](https://scrapfly.io/docs/sdk/scrapy)

### Integrations

- [Getting Started](https://scrapfly.io/docs/integration/getting-started)
- [LangChain](https://scrapfly.io/docs/integration/langchain)
- [LlamaIndex](https://scrapfly.io/docs/integration/llamaindex)
- [CrewAI](https://scrapfly.io/docs/integration/crewai)
- [Zapier](https://scrapfly.io/docs/integration/zapier)
- [Make](https://scrapfly.io/docs/integration/make)
- [n8n](https://scrapfly.io/docs/integration/n8n)

### Academy

- [Overview](https://scrapfly.io/academy)
- [Web Scraping Overview](https://scrapfly.io/academy/scraping-overview)
- [Tools](https://scrapfly.io/academy/tools-overview)
- [Reverse Engineering](https://scrapfly.io/academy/reverse-engineering)
- [Static Scraping](https://scrapfly.io/academy/static-scraping)
- [HTML Parsing](https://scrapfly.io/academy/html-parsing)
- [Dynamic Scraping](https://scrapfly.io/academy/dynamic-scraping)
- [Hidden API Scraping](https://scrapfly.io/academy/hidden-api-scraping)
- [Headless Browsers](https://scrapfly.io/academy/headless-browsers)
- [Hidden Web Data](https://scrapfly.io/academy/hidden-web-data)
- [JSON Parsing](https://scrapfly.io/academy/json-parsing)
- [Data Processing](https://scrapfly.io/academy/data-processing)
- [Scaling](https://scrapfly.io/academy/scaling)
- [Walkthrough Summary](https://scrapfly.io/academy/walkthrough-summary)
- [Scraper Blocking](https://scrapfly.io/academy/scraper-blocking)
- [Proxies](https://scrapfly.io/academy/proxies)

---

# Prompt &amp; Extract

 `/crawl/prompt` and `/crawl/extract` answer questions using the pages a crawl collected. Both retrieve the most relevant passages from one or more crawl indexes, hand only those passages to a model, and return either a cited answer or structured data. Neither one lets the model browse: it sees the retrieved text and nothing else.

 Both endpoints require an index, which means the crawl must have been started with `search: true`. See [Crawl Search](https://scrapfly.io/docs/crawler-api/search) for how the index is built, what it costs you in data terms and why there is no way to add one after the fact.

> **Answers are derived from third-party content** The passages that ground an answer were written by the sites you crawled, not by you and not by Scrapfly. A page author can put text on their page that tries to steer the model, and some do. Read [Untrusted content and prompt injection](#injection) before you feed these responses into anything that acts on them.

## How a prompt is answered

 ```
prompt ─▶ retrieve from the crawl indexes  ─▶  merge and rank
       ─▶ assemble the top passages into a context (capped, see Retrieval budget)
       ─▶ generate ─▶ stream tokens ─▶ validate citations
```

 Retrieval is exactly the [search](https://scrapfly.io/docs/crawler-api/search) path: same modes, same filters, same per-crawl pruning, same completeness rules. The `search` object in the request body is that search request. What changes is the last step: instead of returning ranked passages to you, the API turns them into the model's context.

## POST /crawl/prompt

 | Endpoint | Use |
|---|---|
| `POST /crawl/prompt` | One or many crawls, listed in `crawl_ids`. |
| `POST /crawl/{crawler_uuid}/prompt` | A single crawl. Same implementation, the UUID comes from the path. |

 ```
curl -N -X POST 'https://api.scrapfly.io/crawl/prompt?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -H 'Accept: text/event-stream' \
    -d '{
        "prompt": "Compare the pricing models described across these websites.",
        "crawl_ids": [
            "0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55",
            "0199a1b0-2c44-7c8e-b3f2-8ad1c9e07731"
        ],
        "search": { "limit": 30, "mode": "hybrid" },
        "generation": { "model": "gemini-2.5-flash-lite", "stream": true }
    }'

```

 | Field | Type | Default | Notes |
|---|---|---|---|
| `prompt` | string | required | Your question or instruction. This is the only text that is treated as an instruction. |
| `crawl_ids` | string\[\] | required | Crawls to retrieve from. Same authorization, same per-request cap and same `skipped` semantics as [`/crawl/search`](https://scrapfly.io/docs/crawler-api/search#limits). |
| `search` | object | `{}` | The retrieval request: `limit`, `mode` and `filters`, with the same meanings and the same cap of 50. `limit` defaults to 30 here, not to the 10 `/crawl/search` uses, because an answer needs more passages than a result list does. This is how you scope an answer to a section of a site. |
| `generation.model` | enum | `gemini-2.5-flash-lite` | One of the [supported models](#models). An unknown name is rejected. |
| `generation.stream` | bool | `true` | `true` streams [Server-Sent Events](#sse). `false` returns the same information as one JSON object once generation has finished. |

 Generation runs at temperature 0, so the same prompt over the same index gives the same answer.

## The event stream

 A streaming request answers `Content-Type: text/event-stream` and emits frames as they are produced. There is no `Content-Length` and the response is not compressed, so a client that reads line by line sees tokens as they arrive.

 ```
event: source
data: { "id": 1, "crawler_uuid": "0198c4f2-...", "url": "https://web-scraping.dev/pricing", "title": "Pricing", "score": 0.92 }

event: source
data: { "id": 2, "crawler_uuid": "0199a1b0-...", "url": "https://example.org/plans", "title": "Plans", "score": 0.88 }

:keepalive

event: token
data: "Both sites"

event: token
data: " price per seat"

event: done
data: { "sources_used": [1, 2], "sources_dropped": 0, "truncated": false, "api_credit": 3 }

```

 | Frame | When | Payload |
|---|---|---|
| `event: source` | First, one per retrieved passage, before a single token of the answer. Render them immediately: they are what the answer is grounded in. | `id` (the citation number), `crawler_uuid`, `url`, `title`, `score`. |
| `:keepalive` | Every 15 seconds while retrieval is still running. It is an SSE comment, not an event. | None. Skip any line starting with `:`. |
| `event: token` | Repeatedly, as the model generates. | A JSON string. Concatenate them in order to rebuild the answer. |
| `event: error` | On failure. It can arrive after tokens have already been sent, because generation can fail mid-answer. This frame is the authoritative failure signal. | `code` and `message`. |
| `event: done` | Last frame of a stream that completed, including one that found no sources and generated no answer. It never follows an `error` frame: a failed stream ends with `error` instead and carries no `done`. A stream that ends with neither was cut off. | `sources_used`, `sources_dropped`, `truncated`, `api_credit`. `api_credit` is the authoritative record of what the call was charged, and it is the only price the response reports. |

 ```
import json
import httpx

body = {
    "prompt": "Compare the pricing models described across these websites.",
    "crawl_ids": ["0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55"],
    "search": {"limit": 30, "mode": "hybrid"},
    "generation": {"stream": True},
}

sources = {}
answer = []

with httpx.stream(
    "POST",
    "https://api.scrapfly.io/crawl/prompt",
    params={"key": "{{ YOUR_API_KEY }}"},
    json=body,
    headers={"Accept": "text/event-stream"},
    timeout=None,
) as response:
    event = None
    for line in response.iter_lines():
        if line.startswith(":"):        # keepalive comment, ignore
            continue
        if line.startswith("event: "):
            event = line[7:]
        elif line.startswith("data: "):
            payload = json.loads(line[6:])
            if event == "source":
                sources[payload["id"]] = payload["url"]
            elif event == "token":
                answer.append(payload)
            elif event == "error":
                raise RuntimeError(payload["code"])

print("".join(answer))
print(sources)

```

#### Reading `done`

- `sources_used` is the server-validated set of citation ids the answer actually cites. It is a subset of the `source` frames you received.
- `sources_dropped` counts retrieved passages the model never saw: those that did not fit the [context budget](#budget), plus any hit whose text came back empty. A non-zero value means the answer saw fewer sources than your `search.limit` asked for.
- `truncated: true` means the model hit its output ceiling and the answer stops mid-thought. The text you received is still valid, it is just incomplete. Ask a narrower question or pick a model with a larger output limit.

#### Non-streaming

 Set `generation.stream` to `false` and you get one JSON object with `prompt`, `answer`, `sources` (the same objects the `source` frames carry), a `search` block holding `completeness` and the `skipped` crawls, and the same `sources_used`, `sources_dropped`, `truncated` and `api_credit` fields the `done` frame carries. Nothing is returned until generation completes, so keep your client timeout generous.

 ```
curl -X POST 'https://api.scrapfly.io/crawl/prompt?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -d '{
        "prompt": "Which of these sites offers a free tier?",
        "crawl_ids": ["0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55"],
        "generation": { "stream": false }
    }'

```

## Citations

 The answer cites sources as `[1]`, `[2]`, and those numbers are the `id` field of the `source` frames. Mapping a citation back to a page is a lookup in the map you built while reading the stream:

 ```
[1]  ->  { "id": 1, "crawler_uuid": "0198c4f2-...", "url": "https://web-scraping.dev/pricing" }
```

 Citations are validated on our side before they reach you. If the model invents a reference to an id that was not in its context, that reference is removed rather than forwarded, and the id never appears in `sources_used`. A citation you receive therefore always resolves to a real page from a crawl you own. It does **not** guarantee that the cited page actually supports the sentence it is attached to: for anything consequential, open the URL.

## Retrieval budget

 The context handed to the model is bounded by the retrieval budget, not by the size of your corpus. Searching two crawls and searching two hundred produce contexts of the same size.

 ```
100 crawls  ->  ranked retrieval  ->  ~500 candidates  ->  merge
            ->  the top 20-50 passages  ->  model
```

 | Bound | Value | Why |
|---|---|---|
| Passages retrieved | `search.limit`, 30 by default, capped at 50 | Same ceiling as `/crawl/search`. |
| Assembled context | 120,000 characters | Below the model's own limit on purpose. Very large contexts push generation past the provider's deadline and degrade answer quality more than they help. |
| Per source | 8,000 characters | So one long passage cannot crowd every other source out of the context. |

 The context is filled in rank order, best first. When the next source no longer fits it is dropped whole, never truncated, because half a source produces a citation to text the model never fully saw. Everything dropped this way is counted in `sources_dropped` on the `done` frame.

 Practical consequence: raising `search.limit` past what fits changes nothing except the `sources_dropped` counter. To widen coverage, narrow the question or use `search.filters` so that the passages that do fit are the right ones.

## Models

 | Model | Notes |
|---|---|
| `gemini-2.5-flash-lite` | Default. Fast, large output ceiling, the model the rest of the platform tracks. |
| `gemini-3.5-flash-lite` | Newer lite model. |
| `gemini-2.5-flash` | Stronger reasoning, higher token price. |
| `gemini-2.0-flash` | Cheapest per token, smaller output ceiling. |
| `gemini-2.0-flash-lite` | Cheapest of the family, smaller output ceiling. |
| `gemini-1.5-pro-002` | Legacy. Most expensive, kept for compatibility. |

 Any other value is rejected with [`ERR::CRAWLER::CONFIG_ERROR`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::CONFIG_ERROR) rather than being passed through to the provider.

## POST /crawl/extract

 `/crawl/extract` is the same retrieval with a structured answer. You supply a JSON Schema and get back an object that conforms to it, instead of prose.

 ```
curl -X POST 'https://api.scrapfly.io/crawl/extract?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -d '{
        "prompt": "Extract every plan with its monthly price in USD.",
        "crawl_ids": ["0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55"],
        "search": { "limit": 40, "filters": { "url_prefix": "https://web-scraping.dev/pricing" } },
        "schema": {
            "type": "object",
            "properties": {
                "plans": {
                    "type": "array",
                    "items": {
                        "type": "object",
                        "properties": {
                            "name": { "type": "string" },
                            "monthly_price_usd": { "type": "number" }
                        },
                        "required": ["name", "monthly_price_usd"]
                    }
                }
            },
            "required": ["plans"]
        }
    }'

```

 | Field | Type | Notes |
|---|---|---|
| `prompt` | string | What to extract. Also drives retrieval. |
| `schema` | object | JSON Schema for the result. Keep it tight: types, enums and `required` are what the response is validated against, and a loose schema validates almost anything. |
| `crawl_ids`, `search` | string\[\], object | Identical to `/crawl/prompt`. |

 The response is a single JSON object. There is no streaming variant: a partial structured document is not useful. The schema is handed to the model as its response schema, so generation is constrained to it rather than asked for it in prose: the shape, the types and the enum members you declared are what the model is able to emit. It is a strong constraint, not a post-hoc validator on our side, so keep validating on yours for anything you act on.

## Untrusted content and prompt injection

 This is the part of the feature worth reading twice. A prompt answer is assembled from pages written by whoever owns the sites you crawled. Those authors can see that crawlers read their pages, and some of them write text aimed at whatever reads it next: instructions to ignore your question, to repeat the content of another source, to attribute a claim to a page that never made it, or, for `/crawl/extract`, to push a particular price or contact address into a field your code will trust.

#### What Scrapfly does about it

- **Your prompt is the only instruction.** Retrieved passages are placed in delimited source blocks and the model is told they are data to be quoted, never instructions to follow.
- **The model has no capabilities.** No tools, no function calling, no fetching, no code execution. It reads text and writes text. An injected instruction has nothing to reach for.
- **Citations are validated server side.** References to sources that were not in the context are stripped before the answer reaches you.
- **Extraction is schema-constrained.** Your JSON Schema is the model's response schema, so an injected value cannot introduce a field, a type or an enum member you did not declare. What it constrains is the shape, never the truth of a value that does fit.
- **The model never sees your storage.** Source blocks carry the URL, the title and the crawler UUID, nothing about where the artifact lives.

#### What remains your responsibility

- **Treat the answer as untrusted input.** Do not pass it unchecked to a shell, a database, a downstream agent, or anything that executes. It is derived from content you do not control.
- **Escape it before rendering.** Answers and passage text can contain HTML, markdown, links and image references from the crawled page. Render as plain text, or escape and strip. Auto-linking source content turns an injected string into a live phishing link in your own UI.
- **Check the citations that matter.** A valid citation proves the source was in the context, not that it supports the sentence. Open the URL for anything consequential.
- **Constrain `/crawl/extract` with your schema, then validate the result.** Enums and numeric ranges are the only automatic defence against a value that was planted rather than published, they only work if you declare them, and they bound the shape rather than the content.
- **Scope the retrieval.** `search.filters` keeps a multi-crawl answer from drawing on a site you did not mean to include in that question.

## Errors

 Both endpoints reject a request the same way `/crawl/search` does: an empty `prompt`, a bad model name, a bad `search.limit` or an unknown filter gets [`ERR::CRAWLER::CONFIG_ERROR`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::CONFIG_ERROR), too many crawls gets [`ERR::CRAWLER::SEARCH_TOO_MANY_CRAWLS`](https://scrapfly.io/docs/crawler-api/error/ERR::CRAWLER::SEARCH_TOO_MANY_CRAWLS), and a UUID you do not own gets `403`. A crawl with no queryable index is not an error here either: retrieval reports it under `search.skipped` and the answer is generated from the remaining crawls. When **no** crawl in the request has an index, retrieval returns nothing and the stream closes on a `done` frame carrying `"answer_empty": true` with no token before it.

 A failure that happens after the stream has opened arrives as an `event: error` frame with an HTTP status of `200`, because the status line was already sent. Branch on the frame, not on the status code. The full catalog is on the [Errors](https://scrapfly.io/docs/crawler-api/errors) page.

## Next steps

- Retrieve passages without a model on the [Crawl Search](https://scrapfly.io/docs/crawler-api/search) page.
- Fetch the whole page behind a citation with [`/contents`](https://scrapfly.io/docs/crawler-api/results#query-content).
- See what a prompt costs on the [Billing](https://scrapfly.io/docs/crawler-api/billing#search) page.
- For structured extraction from a single page during the crawl itself, use [extraction rules](https://scrapfly.io/docs/crawler-api/extraction-rules) instead.
