# Scrapfly Documentation

## Table of Contents

### Dashboard

- [Intro](https://scrapfly.io/docs)
- [Project](https://scrapfly.io/docs/project)
- [Account](https://scrapfly.io/docs/account)
- [Workspace & Team](https://scrapfly.io/docs/workspace-and-team)
- [Billing](https://scrapfly.io/docs/billing)

### Products

#### MCP Server

- [Getting Started](https://scrapfly.io/docs/mcp/getting-started)
- [Tools & API Spec](https://scrapfly.io/docs/mcp/tools)
- [Authentication](https://scrapfly.io/docs/mcp/authentication)
- [Examples & Use Cases](https://scrapfly.io/docs/mcp/examples)
- [FAQ](https://scrapfly.io/docs/mcp/faq)
##### Integrations

- [Overview](https://scrapfly.io/docs/mcp/integrations)
- [Claude Desktop](https://scrapfly.io/docs/mcp/integrations/claude-desktop)
- [Claude Code](https://scrapfly.io/docs/mcp/integrations/claude-code)
- [ChatGPT](https://scrapfly.io/docs/mcp/integrations/chatgpt)
- [Cursor](https://scrapfly.io/docs/mcp/integrations/cursor)
- [Cline](https://scrapfly.io/docs/mcp/integrations/cline)
- [Windsurf](https://scrapfly.io/docs/mcp/integrations/windsurf)
- [Zed](https://scrapfly.io/docs/mcp/integrations/zed)
- [Roo Code](https://scrapfly.io/docs/mcp/integrations/roo-code)
- [VS Code](https://scrapfly.io/docs/mcp/integrations/vscode)
- [LangChain](https://scrapfly.io/docs/mcp/integrations/langchain)
- [LlamaIndex](https://scrapfly.io/docs/mcp/integrations/llamaindex)
- [CrewAI](https://scrapfly.io/docs/mcp/integrations/crewai)
- [OpenAI](https://scrapfly.io/docs/mcp/integrations/openai)
- [n8n](https://scrapfly.io/docs/mcp/integrations/n8n)
- [Make](https://scrapfly.io/docs/mcp/integrations/make)
- [Zapier](https://scrapfly.io/docs/mcp/integrations/zapier)
- [Vapi AI](https://scrapfly.io/docs/mcp/integrations/vapi)
- [Agent Builder](https://scrapfly.io/docs/mcp/integrations/agent-builder)
- [Custom Client](https://scrapfly.io/docs/mcp/integrations/custom-client)


#### Web Scraping API

- [Getting Started](https://scrapfly.io/docs/scrape-api/getting-started)
- [API Specification]()
- [Monitoring](https://scrapfly.io/docs/monitoring)
- [Customize Request](https://scrapfly.io/docs/scrape-api/custom)
- [Debug](https://scrapfly.io/docs/scrape-api/debug)
- [Unblocker (formerly ASP)](https://scrapfly.io/docs/scrape-api/unblocker)
- [Proxy](https://scrapfly.io/docs/scrape-api/proxy)
- [Proxy Mode](https://scrapfly.io/docs/scrape-api/proxy-mode)
- [Proxy Mode - Screaming Frog](https://scrapfly.io/docs/scrape-api/proxy-mode/screaming-frog)
- [Proxy Mode - Apify](https://scrapfly.io/docs/scrape-api/proxy-mode/apify)
- [(Auto) Data Extraction](https://scrapfly.io/docs/scrape-api/extraction)
- [Javascript Rendering](https://scrapfly.io/docs/scrape-api/javascript-rendering)
- [Javascript Scenario](https://scrapfly.io/docs/scrape-api/javascript-scenario)
- [SSL](https://scrapfly.io/docs/scrape-api/ssl)
- [DNS](https://scrapfly.io/docs/scrape-api/dns)
- [Cache](https://scrapfly.io/docs/scrape-api/cache)
- [Batch (Multi-URL Scraping)](https://scrapfly.io/docs/scrape-api/batch)
- [Session](https://scrapfly.io/docs/scrape-api/session)
- [Webhook](https://scrapfly.io/docs/scrape-api/webhook)
- [Schedule](https://scrapfly.io/docs/scrape-api/schedule)
- [Screenshot](https://scrapfly.io/docs/scrape-api/screenshot)
- [Errors](https://scrapfly.io/docs/scrape-api/errors)
- [Timeout](https://scrapfly.io/docs/scrape-api/understand-timeout)
- [Throttling](https://scrapfly.io/docs/throttling)
- [Troubleshoot](https://scrapfly.io/docs/scrape-api/troubleshoot)
- [Billing](https://scrapfly.io/docs/scrape-api/billing)
- [FAQ](https://scrapfly.io/docs/scrape-api/faq)

#### Crawler API

- [Getting Started](https://scrapfly.io/docs/crawler-api/getting-started)
- [API Specification]()
- [Retrieving Results](https://scrapfly.io/docs/crawler-api/results)
- [WARC Format](https://scrapfly.io/docs/crawler-api/warc-format)
- [Data Extraction](https://scrapfly.io/docs/crawler-api/extraction-rules)
- [Search](https://scrapfly.io/docs/crawler-api/search)
- [Prompt & Extract](https://scrapfly.io/docs/crawler-api/prompt)
- [Auto Refresh](https://scrapfly.io/docs/crawler-api/refresh)
- [Webhook](https://scrapfly.io/docs/crawler-api/webhook)
- [Schedule](https://scrapfly.io/docs/crawler-api/schedule)
- [Billing](https://scrapfly.io/docs/crawler-api/billing)
- [Errors](https://scrapfly.io/docs/crawler-api/errors)
- [Troubleshoot](https://scrapfly.io/docs/crawler-api/troubleshoot)
- [FAQ](https://scrapfly.io/docs/crawler-api/faq)

#### Screenshot API

- [Getting Started](https://scrapfly.io/docs/screenshot-api/getting-started)
- [API Specification]()
- [Accessibility Testing](https://scrapfly.io/docs/screenshot-api/accessibility)
- [Webhook](https://scrapfly.io/docs/screenshot-api/webhook)
- [Schedule](https://scrapfly.io/docs/screenshot-api/schedule)
- [Billing](https://scrapfly.io/docs/screenshot-api/billing)
- [Errors](https://scrapfly.io/docs/screenshot-api/errors)

#### Extraction API

- [Getting Started](https://scrapfly.io/docs/extraction-api/getting-started)
- [API Specification]()
- [Rules Template](https://scrapfly.io/docs/extraction-api/rules-and-template)
- [LLM Extraction](https://scrapfly.io/docs/extraction-api/llm-prompt)
- [AI Auto Extraction](https://scrapfly.io/docs/extraction-api/automatic-ai)
- [Webhook](https://scrapfly.io/docs/extraction-api/webhook)
- [Billing](https://scrapfly.io/docs/extraction-api/billing)
- [Errors](https://scrapfly.io/docs/extraction-api/errors)
- [FAQ](https://scrapfly.io/docs/extraction-api/faq)

#### Data API


#### Proxy Saver

- [Getting Started](https://scrapfly.io/docs/proxy-saver/getting-started)
- [Fingerprints](https://scrapfly.io/docs/proxy-saver/fingerprints)
- [Optimizations](https://scrapfly.io/docs/proxy-saver/optimizations)
- [SSL Certificates](https://scrapfly.io/docs/proxy-saver/certificates)
- [Protocols](https://scrapfly.io/docs/proxy-saver/protocols)
- [Pacfile](https://scrapfly.io/docs/proxy-saver/pacfile)
- [Secure Credentials](https://scrapfly.io/docs/proxy-saver/security)
- [Billing](https://scrapfly.io/docs/proxy-saver/billing)

#### Cloud Browser API

- [Getting Started](https://scrapfly.io/docs/cloud-browser-api/getting-started)
- [Proxy & Geo-Targeting](https://scrapfly.io/docs/cloud-browser-api/proxy)
- [Unblock API](https://scrapfly.io/docs/cloud-browser-api/unblock)
- [Captcha Solver](https://scrapfly.io/docs/cloud-browser-api/captcha-solver)
- [File Downloads](https://scrapfly.io/docs/cloud-browser-api/file-downloads)
- [Session Resume](https://scrapfly.io/docs/cloud-browser-api/session-resume)
- [Human-in-the-Loop](https://scrapfly.io/docs/cloud-browser-api/human-in-the-loop)
- [Debug Mode](https://scrapfly.io/docs/cloud-browser-api/debug-mode)
- [Browser Extensions](https://scrapfly.io/docs/cloud-browser-api/extensions)
- [Native Browser MCP](https://scrapfly.io/docs/cloud-browser-api/mcp)
- [DevTools Protocol](https://scrapfly.io/docs/cloud-browser-api/cdp-reference)
##### Integrations

- [Puppeteer](https://scrapfly.io/docs/cloud-browser-api/puppeteer)
- [Playwright](https://scrapfly.io/docs/cloud-browser-api/playwright)
- [Selenium](https://scrapfly.io/docs/cloud-browser-api/selenium)
- [Vercel Agent Browser](https://scrapfly.io/docs/cloud-browser-api/agent-browser)
- [Browser Use](https://scrapfly.io/docs/cloud-browser-api/browser-use)
- [Stagehand](https://scrapfly.io/docs/cloud-browser-api/stagehand)

- [Billing](https://scrapfly.io/docs/cloud-browser-api/billing)
- [Errors](https://scrapfly.io/docs/cloud-browser-api/errors)


### Tools

- [Antibot Detector](https://scrapfly.io/docs/tools/antibot-detector)

### SDK

- [Golang](https://scrapfly.io/docs/sdk/golang)
- [Python](https://scrapfly.io/docs/sdk/python)
- [Rust](https://scrapfly.io/docs/sdk/rust)
- [TypeScript](https://scrapfly.io/docs/sdk/typescript)
- [Scrapy](https://scrapfly.io/docs/sdk/scrapy)

### Integrations

- [Getting Started](https://scrapfly.io/docs/integration/getting-started)
- [LangChain](https://scrapfly.io/docs/integration/langchain)
- [LlamaIndex](https://scrapfly.io/docs/integration/llamaindex)
- [CrewAI](https://scrapfly.io/docs/integration/crewai)
- [Zapier](https://scrapfly.io/docs/integration/zapier)
- [Make](https://scrapfly.io/docs/integration/make)
- [n8n](https://scrapfly.io/docs/integration/n8n)

### Academy

- [Overview](https://scrapfly.io/academy)
- [Web Scraping Overview](https://scrapfly.io/academy/scraping-overview)
- [Tools](https://scrapfly.io/academy/tools-overview)
- [Reverse Engineering](https://scrapfly.io/academy/reverse-engineering)
- [Static Scraping](https://scrapfly.io/academy/static-scraping)
- [HTML Parsing](https://scrapfly.io/academy/html-parsing)
- [Dynamic Scraping](https://scrapfly.io/academy/dynamic-scraping)
- [Hidden API Scraping](https://scrapfly.io/academy/hidden-api-scraping)
- [Headless Browsers](https://scrapfly.io/academy/headless-browsers)
- [Hidden Web Data](https://scrapfly.io/academy/hidden-web-data)
- [JSON Parsing](https://scrapfly.io/academy/json-parsing)
- [Data Processing](https://scrapfly.io/academy/data-processing)
- [Scaling](https://scrapfly.io/academy/scaling)
- [Walkthrough Summary](https://scrapfly.io/academy/walkthrough-summary)
- [Scraper Blocking](https://scrapfly.io/academy/scraper-blocking)
- [Proxies](https://scrapfly.io/academy/proxies)

---

# FAQ

 Here are some of the most common issues and questions that come up when using Scrapfly Crawler API. See the tag filter on the right for more.

##  What is the Crawler API?

 The Scrapfly Crawler API enables recursive website crawling at scale. Unlike the Web Scraping API which scrapes individual URLs, the Crawler API can automatically discover and scrape entire websites following links and applying URL filtering rules.

 Results are delivered as industry-standard artifacts in [WARC Format](https://scrapfly.io/docs/crawler-api/warc-format) and HAR formats for easy integration with data pipelines. See [Crawler API Getting Started](https://scrapfly.io/docs/crawler-api/getting-started) for more details.

##  What's the difference between Crawler API and Web Scraping API?

 The **Web Scraping API** scrapes individual URLs that you provide. You control exactly which pages to scrape.

 The **Crawler API** automatically discovers URLs by following links from a starting point (seed URL). It can crawl entire websites or sections based on URL filtering rules, making it ideal for bulk data collection and site archival.

 Both APIs use the same underlying scraping infrastructure and support the same features like [Unblocker](https://scrapfly.io/docs/scrape-api/unblocker), [proxies](https://scrapfly.io/docs/scrape-api/proxy), and [JavaScript rendering](https://scrapfly.io/docs/scrape-api/javascript-rendering).

##  How do I start a crawl?

 To start a crawl, send a POST request to the `/crawl` endpoint with your configuration:

- **Seed URLs:** Starting points for the crawl (e.g., homepage, category pages)
- **URL Rules:** Patterns to include/exclude URLs (regex or glob patterns)
- **Limits:** Maximum pages, crawl depth, time limits
- **Scrape Config:** Web Scraping API parameters (Unblocker, proxies, JavaScript rendering, etc.)

 The API returns a crawler UUID which you use to check status and retrieve results. See [Crawler API Getting Started](https://scrapfly.io/docs/crawler-api/getting-started) for detailed examples.

##  How do I control which URLs get crawled?

 Use URL filtering rules to include or exclude specific URL patterns:

- **Include rules:** Only crawl URLs matching these patterns (e.g., `/products/*`)
- **Exclude rules:** Skip URLs matching these patterns (e.g., `/login`, `/cart`)
- **Domain restrictions:** Stay within specific domains or subdomains

 Rules support both glob patterns (`*.html`) and regex patterns for maximum flexibility. See [Crawler API Getting Started](https://scrapfly.io/docs/crawler-api/getting-started) for configuration details.

##  How do I limit the size of my crawl?

 Configure limits to control crawl scope and prevent runaway costs:

- **`page_limit`:** Maximum number of pages to crawl
- **`max_depth`:** Maximum link depth from seed URLs
- **`max_duration`:** Maximum crawl time in seconds
- **`max_api_credit`:** Maximum API credits to spend

 The crawl stops when any limit is reached. Always set appropriate limits to avoid unexpected costs.

##  How do I check the status of my crawl?

 Use the `GET /crawl/{crawler_uuid}/status` endpoint to check crawl status. The response includes:

- **status:** `PENDING`, `SCHEDULED`, `RUNNING`, `PAUSED`, `RESCHEDULE`, `DONE`, or `CANCELLED`
- **is\_finished:** `true` once the crawl has terminated
- **is\_success:** `true` if the crawl succeeded; a failed crawl reports `DONE` with `is_success: false`
- **state:** the run counters, including `urls_visited`, `urls_extracted`, `urls_failed`, `urls_skipped`, `urls_to_crawl`, `api_credit_used`, `duration` and `stop_reason`
- **urls:** a map of each indexed URL to its status, empty until the crawl has written its URL index
- **project** and **env:** the scope the crawl runs in

 Poll this endpoint periodically until `is_finished` is `true`, then check `is_success` to tell a successful crawl from a failed one. See [Crawler API Getting Started](https://scrapfly.io/docs/crawler-api/getting-started) for workflow details.

##  Can I get notified when my crawl completes?

 Yes! Configure a webhook URL when creating the crawl. Scrapfly will send a POST request to your webhook when:

- The crawl completes successfully
- The crawl fails or is cancelled
- Each page is crawled (optional)

 Webhooks eliminate the need for polling and provide real-time notifications. See [Crawler API Webhook](https://scrapfly.io/docs/crawler-api/webhook) for configuration and examples.

##  How do I retrieve crawl results?

 Results are available as downloadable artifacts in multiple formats:

- **WARC:** Industry-standard web archive format (gzipped)
- **HAR:** HTTP Archive format for request/response inspection

 Use the `GET /crawl/{crawler_uuid}/artifact?type={format}` endpoint to download results. See [Crawler API Results](https://scrapfly.io/docs/crawler-api/results) for detailed format specifications.

##  Which artifact format should I use?

 Choose based on your use case:

- **WARC:** Best for archival, compliance, and reprocessing. Standard format used by Internet Archive and libraries.
- **HAR:** Best for debugging, request inspection, and replaying requests in browser dev tools.

 All formats contain the same data - choose what works best for your workflow. See [WARC Format](https://scrapfly.io/docs/crawler-api/warc-format) and [Crawler API Results](https://scrapfly.io/docs/crawler-api/results) for format details.

##  Can the Crawler API extract structured data?

 Yes! Configure extraction rules to extract structured data from crawled pages simultaneously. The Crawler API supports:

- **AI extraction models:** Automatic extraction using predefined models (products, articles, reviews, etc.)
- **LLM prompts:** Custom extraction using AI prompts
- **Template-based:** Define JSON templates with CSS/XPath selectors

 Extracted data is included in the result artifacts alongside raw HTML content. See [Crawler API Extraction Rules](https://scrapfly.io/docs/crawler-api/extraction-rules) for configuration examples.

##  Can I search the pages a crawl collected?

 Yes, if the crawl was started with `search: true`. That builds a search index over the crawled text while the crawl runs. Once it finishes you can query it by meaning or by exact string, across one crawl or a collection of crawls in one request, with `POST /crawl/search`. `POST /crawl/prompt` and `POST /crawl/extract` go one step further and run a model over the retrieved passages to return a cited answer or structured data.

 See [Crawl Search](https://scrapfly.io/docs/crawler-api/search) and [Prompt &amp; Extract](https://scrapfly.io/docs/crawler-api/prompt).

##  What does enabling search do with my data?

 `search: true` sends the text of every indexed page to Scrapfly's embedding provider (Google Vertex AI) to compute the vectors that make semantic search possible. That is what the flag is for, which is why it is opt-in per crawl rather than an account-wide setting.

 The index also stores the matched passages themselves as text, next to the crawl, under the same project and environment prefix and behind the same access control. It inherits the crawl's retention window and is deleted with the crawl. There is no separate index retention and no separate deletion to request. See [Crawl Search](https://scrapfly.io/docs/crawler-api/search).

##  Why is my search index PARTIAL?

 `PARTIAL` means the index was published but does not cover every page the crawl visited. Two things cause it, and both are the same trade-off: indexing runs beside the crawl and is never allowed to slow it down or fail it.

- **Documents were dropped.** The crawl produced pages faster than they could be embedded, so the indexing queue overflowed and the excess was discarded rather than made the crawler wait. `search.dropped` counts them.
- **The final drain hit its deadline.** Pages still queued when the crawl ended were not indexed in time to be published.

 A `PARTIAL` index is fully usable, and the results it returns are correct. They are just drawn from fewer pages than the crawl visited. If full coverage matters, lower `max_concurrency` so pages arrive at a pace the indexer can keep up with, or re-run the crawl. Note that `crawler_search_ready` fires for `PARTIAL` too, so read `payload.search.status` rather than assuming `READY`.

##  Why does my paused crawl have no search index?

 An index is only published for a crawl that reached `DONE`. A paused crawl has not finished: it is resumed automatically and keeps adding pages, so publishing an index for it would advertise a complete corpus over a corpus that is still being written.

 While the crawl is paused, `search.status` stays `BUILDING` and it is not queryable on either route: a search that names it reports it under `skipped` with reason `search_not_ready` and returns no rows from it. Wait for the crawl to finish, or for the `crawler_search_ready` webhook.

##  Can I add a search index to a crawl I already ran?

 No. The index is built from data that only exists while a page is being crawled, so it cannot be reconstructed afterwards from the stored artifacts. A crawl that ran without `search: true`, including every crawl that finished before the feature existed, has no index and cannot be given one. There is no backfill endpoint and none is planned.

 The only way to search an old crawl is to run it again with `search: true`. Searching an unindexed crawl is not an error: the request answers `200` and lists that crawl under `skipped` with reason `search_not_enabled`, on the single-crawl route as much as on the collection one.

##  How is Crawler API billing calculated?

 Crawler API billing is simple: **total cost = sum of all Web Scraping API calls made during the crawl**.

 Each page crawled is billed as a Web Scraping API request based on your enabled features (Unblocker, JavaScript rendering, proxies, screenshots, etc.). The total cost is shown in the crawl status response.

 See [Crawler API Billing](https://scrapfly.io/docs/crawler-api/billing) for detailed cost calculation and examples.

##  How do I estimate the cost of a crawl before running it?

 To estimate costs:

1. Estimate the number of pages that will be crawled (use `page_limit` as upper bound)
2. Check the cost per page in [Crawler API Billing](https://scrapfly.io/docs/crawler-api/billing) based on your scrape configuration
3. Multiply: `estimated_pages × cost_per_page = total_cost`

 Start with a small crawl (low `page_limit`) to test your configuration and measure actual cost per page. You can also set `max_api_credit` to automatically stop the crawl when reaching a credit limit.

##  How do I prevent unexpected crawl costs?

 Always configure crawl limits to prevent runaway costs:

- Set `page_limit` to limit total pages crawled
- Set `max_api_credit` to cap total API credits spent
- Set `max_duration` to limit crawl time
- Configure [project limits](https://scrapfly.io/docs/project) for additional spending controls

 The crawl automatically stops when any limit is reached, preventing unexpected charges.

##  Why did my crawl fail?

 Check the crawl status response for error details. Common reasons include:

- **Invalid configuration:** Check seed URLs, URL rules, and scrape parameters
- **All pages blocked:** Target website blocking requests - enable [Unblocker](https://scrapfly.io/docs/scrape-api/unblocker)
- **No URLs found:** URL filtering rules too restrictive or no links on seed pages
- **Credit limit exceeded:** Check your account quota and [billing](https://scrapfly.io/docs/billing)

 See [Crawler API Errors](https://scrapfly.io/docs/crawler-api/errors) for all error codes and [Crawler API Troubleshooting](https://scrapfly.io/docs/crawler-api/troubleshoot) for debugging tips.

##  Why did my crawl not find any URLs?

 This usually happens due to:

- **Too restrictive URL rules:** Include rules don't match any links on the seed pages
- **JavaScript-loaded links:** Set `rendering_delay` (in milliseconds) in the crawl config to render JavaScript
- **Wrong seed URLs:** Seed pages don't contain links to follow
- **Cross-domain restrictions:** Links point to different domains and same-domain crawling is enforced

 Test your URL rules on a small sample first and check the HAR artifact to inspect discovered links.

##  Does Crawler API support anti-bot bypass?

 Yes! Enable [Unblocker](https://scrapfly.io/docs/scrape-api/unblocker) in the scrape configuration when creating the crawl. The Unblocker is applied to every page crawled, bypassing protections like Cloudflare, PerimeterX, DataDome, etc.

 All Web Scraping API features work with the Crawler API including proxies, JavaScript rendering, sessions, and screenshots.

##  Does Crawler API rotate proxies automatically?

 Yes! The Crawler API automatically rotates proxies between pages to avoid rate limiting and IP blocks. Configure proxy settings in the scrape configuration:

- **proxy\_pool:** Choose datacenter or residential proxies
- **country:** Specify proxy country or list of countries

 See [Proxy documentation](https://scrapfly.io/docs/scrape-api/proxy) for all available proxy options.

##  Can Crawler API handle JavaScript-rendered websites?

 Yes! Set `rendering_delay` to a non-zero value (milliseconds, max 25000) in the crawl configuration. The crawler will use headless browsers to execute JavaScript on every page, making it ideal for modern SPAs and dynamic websites.

 You can also use [JavaScript scenarios](https://scrapfly.io/docs/scrape-api/javascript-scenario) to automate interactions like clicking buttons, scrolling, filling forms, etc. before extracting content.

##  How fast does the Crawler API crawl?

 Crawl speed depends on your account's concurrency limit. The crawler uses your available concurrency to fetch multiple pages in parallel, significantly speeding up large crawls.

 For example, with 100 concurrent requests, the crawler can fetch 100 pages simultaneously. Configure [throttling rules](https://scrapfly.io/docs/throttling) if you need to slow down crawling for specific domains.

##  Does Crawler API support sessions?

 Yes! Configure a session in the scrape configuration to maintain cookies and context across crawled pages. This is useful for:

- Crawling authenticated areas (after login)
- Maintaining shopping cart state
- Preserving user preferences across pages

 See [Session](https://scrapfly.io/docs/scrape-api/session) for session configuration details.

##  Can I monitor crawl progress in real-time?

 Yes! The crawl status endpoint provides real-time progress metrics:

- **Pages crawled:** Total pages successfully scraped
- **Pages pending:** URLs queued for crawling
- **Pages failed:** URLs that failed to scrape
- **Current depth:** Current crawl depth from seed URLs
- **Credits used:** API credits consumed so far

 Additionally, all crawled pages appear in the [Monitoring dashboard](https://scrapfly.io/docs/monitoring) for detailed inspection.

##  Can I cancel a running crawl?

 Yes! Send a `POST` request to `/crawl/{crawler_uuid}/cancel` to cancel a running crawl. The crawler stops and the job stays in status `CANCELLED`. Artifact downloads (WARC/HAR) are only served for a crawl that reached `DONE`, so they are refused with a 400 for a cancelled job. Content for the pages already crawled stays reachable through the `/contents` endpoints.

 You are only charged for pages successfully crawled before cancellation.

##  Can I resume a failed or cancelled crawl?

 Currently, crawls cannot be resumed. If a crawl fails or is cancelled, you need to create a new crawl. You can still retrieve the pages that were successfully scraped through the `/contents` endpoints. Artifact downloads (WARC/HAR) are not available for a crawl that never reached `DONE`.

##  How long are crawl artifacts stored?

 Crawl artifacts are stored according to your plan's log retention policy, typically 7-90 days depending on your subscription. Download artifacts before they expire if you need long-term storage.

 Check your [project settings](https://scrapfly.io/docs/project) for your specific retention period.

##  Can I schedule recurring crawls?

 Yes. The Crawler API has built-in scheduling. `POST /crawl/schedules` pairs a `crawler_config` with a `recurrence` (a 5-field cron expression evaluated in UTC, or an interval and unit) and a `webhook_name` that receives each run.

- `POST /crawl/schedules` to create, `GET /crawl/schedules` to list
- `GET`, `PATCH` and `DELETE /crawl/schedules/{id}` to manage one schedule
- `POST /crawl/schedules/{id}/pause`, `/resume` and `/execute`

 See [Crawler API Schedule](https://scrapfly.io/docs/crawler-api/schedule) for the full contract and SDK examples.

 An external scheduler (cron, AWS EventBridge, GCP Cloud Scheduler, Airflow, n8n) still works if you already run one: trigger `POST /crawl` on your own schedule.

##  Can I use custom extraction templates?

 Yes! Define custom extraction rules using CSS or XPath selectors to extract specific data fields from crawled pages. You can also use AI models or LLM prompts for intelligent extraction.

 See [Crawler API Extraction Rules](https://scrapfly.io/docs/crawler-api/extraction-rules) for template syntax and examples.

##  Why is my crawl running slowly?

 Common reasons for slow crawls:

- **Low concurrency:** Upgrade your plan for higher concurrency limits
- **JavaScript rendering:** Headless browsers are slower than HTTP requests - only enable if needed
- **Slow target website:** Some websites respond slowly or throttle requests
- **Throttling rules:** Check if you have [throttling](https://scrapfly.io/docs/throttling) configured

 Leave `rendering_delay` at 0 for static pages and increase concurrency for faster crawls.

##  Why did my crawl stop before completing?

 Check which limit was reached in the crawl status response:

 The `state.stop_reason` field reports which limit triggered the stop. Possible values:

- `page_limit`: `page_limit` reached
- `max_depth`: `max_depth` reached
- `max_duration`: `max_duration` exceeded
- `max_api_credit`: `max_api_credit` budget exhausted
- `no_more_urls`: all discoverable URLs within rules have been crawled
- `seed_url_failed`: the seed URL itself could not be fetched
- `user_cancelled`: user called `POST /crawl/{uuid}/cancel`
- `no_api_credit_left`: account ran out of credits
- `crawler_error`: an internal crawler error stopped the run

 Increase the relevant limit if you want to crawl more pages.

##  How do I integrate Crawler API with my data pipeline?

 The Crawler API is designed for easy integration:

- **Webhooks:** Get notified when crawls complete and trigger downstream processing
- **WARC to Parquet:** Convert the WARC artifact to Parquet for BigQuery, Snowflake, Databricks, or Pandas, see [WARC Format](https://scrapfly.io/docs/crawler-api/warc-format)
- **WARC artifacts:** Standard format supported by archival and ETL tools
- **REST API:** Easy integration with any programming language or workflow tool

 Use webhooks + WARC for the most streamlined integration with modern data stacks.

##  Does Crawler API work with Scrapfly SDKs?

 Yes! Both [Python SDK](https://scrapfly.io/docs/sdk/python) and [Typescript SDK](https://scrapfly.io/docs/sdk/typescript) support the Crawler API with convenient wrapper methods for creating crawls, checking status, and downloading artifacts.

 SDKs handle authentication, request formatting, and error handling automatically.

##  When should I use Crawler API vs Web Scraping API?

 Use **Crawler API** when:

- You need to scrape entire websites or large sections
- You want automatic URL discovery by following links
- You need archival-quality artifacts (WARC)
- You're building a search engine, data lake, or compliance archive

 Use **Web Scraping API** when:

- You have a specific list of URLs to scrape
- You need real-time scraping with immediate results
- You're building application features (price monitoring, product data, etc.)
- You want maximum control over scraping order and logic
