# Scrapfly Documentation

## Table of Contents

### Dashboard

- [Intro](https://scrapfly.io/docs)
- [Project](https://scrapfly.io/docs/project)
- [Account](https://scrapfly.io/docs/account)
- [Workspace & Team](https://scrapfly.io/docs/workspace-and-team)
- [Billing](https://scrapfly.io/docs/billing)

### Products

#### MCP Server

- [Getting Started](https://scrapfly.io/docs/mcp/getting-started)
- [Tools & API Spec](https://scrapfly.io/docs/mcp/tools)
- [Authentication](https://scrapfly.io/docs/mcp/authentication)
- [Examples & Use Cases](https://scrapfly.io/docs/mcp/examples)
- [FAQ](https://scrapfly.io/docs/mcp/faq)
##### Integrations

- [Overview](https://scrapfly.io/docs/mcp/integrations)
- [Claude Desktop](https://scrapfly.io/docs/mcp/integrations/claude-desktop)
- [Claude Code](https://scrapfly.io/docs/mcp/integrations/claude-code)
- [ChatGPT](https://scrapfly.io/docs/mcp/integrations/chatgpt)
- [Cursor](https://scrapfly.io/docs/mcp/integrations/cursor)
- [Cline](https://scrapfly.io/docs/mcp/integrations/cline)
- [Windsurf](https://scrapfly.io/docs/mcp/integrations/windsurf)
- [Zed](https://scrapfly.io/docs/mcp/integrations/zed)
- [Roo Code](https://scrapfly.io/docs/mcp/integrations/roo-code)
- [VS Code](https://scrapfly.io/docs/mcp/integrations/vscode)
- [LangChain](https://scrapfly.io/docs/mcp/integrations/langchain)
- [LlamaIndex](https://scrapfly.io/docs/mcp/integrations/llamaindex)
- [CrewAI](https://scrapfly.io/docs/mcp/integrations/crewai)
- [OpenAI](https://scrapfly.io/docs/mcp/integrations/openai)
- [n8n](https://scrapfly.io/docs/mcp/integrations/n8n)
- [Make](https://scrapfly.io/docs/mcp/integrations/make)
- [Zapier](https://scrapfly.io/docs/mcp/integrations/zapier)
- [Vapi AI](https://scrapfly.io/docs/mcp/integrations/vapi)
- [Agent Builder](https://scrapfly.io/docs/mcp/integrations/agent-builder)
- [Custom Client](https://scrapfly.io/docs/mcp/integrations/custom-client)


#### Web Scraping API

- [Getting Started](https://scrapfly.io/docs/scrape-api/getting-started)
- [API Specification]()
- [Monitoring](https://scrapfly.io/docs/monitoring)
- [Customize Request](https://scrapfly.io/docs/scrape-api/custom)
- [Debug](https://scrapfly.io/docs/scrape-api/debug)
- [Unblocker (formerly ASP)](https://scrapfly.io/docs/scrape-api/unblocker)
- [Proxy](https://scrapfly.io/docs/scrape-api/proxy)
- [Proxy Mode](https://scrapfly.io/docs/scrape-api/proxy-mode)
- [Proxy Mode - Screaming Frog](https://scrapfly.io/docs/scrape-api/proxy-mode/screaming-frog)
- [Proxy Mode - Apify](https://scrapfly.io/docs/scrape-api/proxy-mode/apify)
- [(Auto) Data Extraction](https://scrapfly.io/docs/scrape-api/extraction)
- [Javascript Rendering](https://scrapfly.io/docs/scrape-api/javascript-rendering)
- [Javascript Scenario](https://scrapfly.io/docs/scrape-api/javascript-scenario)
- [SSL](https://scrapfly.io/docs/scrape-api/ssl)
- [DNS](https://scrapfly.io/docs/scrape-api/dns)
- [Cache](https://scrapfly.io/docs/scrape-api/cache)
- [Batch (Multi-URL Scraping)](https://scrapfly.io/docs/scrape-api/batch)
- [Session](https://scrapfly.io/docs/scrape-api/session)
- [Webhook](https://scrapfly.io/docs/scrape-api/webhook)
- [Schedule](https://scrapfly.io/docs/scrape-api/schedule)
- [Screenshot](https://scrapfly.io/docs/scrape-api/screenshot)
- [Errors](https://scrapfly.io/docs/scrape-api/errors)
- [Timeout](https://scrapfly.io/docs/scrape-api/understand-timeout)
- [Throttling](https://scrapfly.io/docs/throttling)
- [Troubleshoot](https://scrapfly.io/docs/scrape-api/troubleshoot)
- [Billing](https://scrapfly.io/docs/scrape-api/billing)
- [FAQ](https://scrapfly.io/docs/scrape-api/faq)

#### Crawler API

- [Getting Started](https://scrapfly.io/docs/crawler-api/getting-started)
- [API Specification]()
- [Retrieving Results](https://scrapfly.io/docs/crawler-api/results)
- [WARC Format](https://scrapfly.io/docs/crawler-api/warc-format)
- [Data Extraction](https://scrapfly.io/docs/crawler-api/extraction-rules)
- [Search](https://scrapfly.io/docs/crawler-api/search)
- [Prompt & Extract](https://scrapfly.io/docs/crawler-api/prompt)
- [Auto Refresh](https://scrapfly.io/docs/crawler-api/refresh)
- [Webhook](https://scrapfly.io/docs/crawler-api/webhook)
- [Schedule](https://scrapfly.io/docs/crawler-api/schedule)
- [Billing](https://scrapfly.io/docs/crawler-api/billing)
- [Errors](https://scrapfly.io/docs/crawler-api/errors)
- [Troubleshoot](https://scrapfly.io/docs/crawler-api/troubleshoot)
- [FAQ](https://scrapfly.io/docs/crawler-api/faq)

#### Screenshot API

- [Getting Started](https://scrapfly.io/docs/screenshot-api/getting-started)
- [API Specification]()
- [Accessibility Testing](https://scrapfly.io/docs/screenshot-api/accessibility)
- [Webhook](https://scrapfly.io/docs/screenshot-api/webhook)
- [Schedule](https://scrapfly.io/docs/screenshot-api/schedule)
- [Billing](https://scrapfly.io/docs/screenshot-api/billing)
- [Errors](https://scrapfly.io/docs/screenshot-api/errors)

#### Extraction API

- [Getting Started](https://scrapfly.io/docs/extraction-api/getting-started)
- [API Specification]()
- [Rules Template](https://scrapfly.io/docs/extraction-api/rules-and-template)
- [LLM Extraction](https://scrapfly.io/docs/extraction-api/llm-prompt)
- [AI Auto Extraction](https://scrapfly.io/docs/extraction-api/automatic-ai)
- [Webhook](https://scrapfly.io/docs/extraction-api/webhook)
- [Billing](https://scrapfly.io/docs/extraction-api/billing)
- [Errors](https://scrapfly.io/docs/extraction-api/errors)
- [FAQ](https://scrapfly.io/docs/extraction-api/faq)

#### Data API


#### Proxy Saver

- [Getting Started](https://scrapfly.io/docs/proxy-saver/getting-started)
- [Fingerprints](https://scrapfly.io/docs/proxy-saver/fingerprints)
- [Optimizations](https://scrapfly.io/docs/proxy-saver/optimizations)
- [SSL Certificates](https://scrapfly.io/docs/proxy-saver/certificates)
- [Protocols](https://scrapfly.io/docs/proxy-saver/protocols)
- [Pacfile](https://scrapfly.io/docs/proxy-saver/pacfile)
- [Secure Credentials](https://scrapfly.io/docs/proxy-saver/security)
- [Billing](https://scrapfly.io/docs/proxy-saver/billing)

#### Cloud Browser API

- [Getting Started](https://scrapfly.io/docs/cloud-browser-api/getting-started)
- [Proxy & Geo-Targeting](https://scrapfly.io/docs/cloud-browser-api/proxy)
- [Unblock API](https://scrapfly.io/docs/cloud-browser-api/unblock)
- [Captcha Solver](https://scrapfly.io/docs/cloud-browser-api/captcha-solver)
- [File Downloads](https://scrapfly.io/docs/cloud-browser-api/file-downloads)
- [Session Resume](https://scrapfly.io/docs/cloud-browser-api/session-resume)
- [Human-in-the-Loop](https://scrapfly.io/docs/cloud-browser-api/human-in-the-loop)
- [Debug Mode](https://scrapfly.io/docs/cloud-browser-api/debug-mode)
- [Browser Extensions](https://scrapfly.io/docs/cloud-browser-api/extensions)
- [Native Browser MCP](https://scrapfly.io/docs/cloud-browser-api/mcp)
- [DevTools Protocol](https://scrapfly.io/docs/cloud-browser-api/cdp-reference)
##### Integrations

- [Puppeteer](https://scrapfly.io/docs/cloud-browser-api/puppeteer)
- [Playwright](https://scrapfly.io/docs/cloud-browser-api/playwright)
- [Selenium](https://scrapfly.io/docs/cloud-browser-api/selenium)
- [Vercel Agent Browser](https://scrapfly.io/docs/cloud-browser-api/agent-browser)
- [Browser Use](https://scrapfly.io/docs/cloud-browser-api/browser-use)
- [Stagehand](https://scrapfly.io/docs/cloud-browser-api/stagehand)

- [Billing](https://scrapfly.io/docs/cloud-browser-api/billing)
- [Errors](https://scrapfly.io/docs/cloud-browser-api/errors)


### Tools

- [Antibot Detector](https://scrapfly.io/docs/tools/antibot-detector)

### SDK

- [Golang](https://scrapfly.io/docs/sdk/golang)
- [Python](https://scrapfly.io/docs/sdk/python)
- [Rust](https://scrapfly.io/docs/sdk/rust)
- [TypeScript](https://scrapfly.io/docs/sdk/typescript)
- [Scrapy](https://scrapfly.io/docs/sdk/scrapy)

### Integrations

- [Getting Started](https://scrapfly.io/docs/integration/getting-started)
- [LangChain](https://scrapfly.io/docs/integration/langchain)
- [LlamaIndex](https://scrapfly.io/docs/integration/llamaindex)
- [CrewAI](https://scrapfly.io/docs/integration/crewai)
- [Zapier](https://scrapfly.io/docs/integration/zapier)
- [Make](https://scrapfly.io/docs/integration/make)
- [n8n](https://scrapfly.io/docs/integration/n8n)

### Academy

- [Overview](https://scrapfly.io/academy)
- [Web Scraping Overview](https://scrapfly.io/academy/scraping-overview)
- [Tools](https://scrapfly.io/academy/tools-overview)
- [Reverse Engineering](https://scrapfly.io/academy/reverse-engineering)
- [Static Scraping](https://scrapfly.io/academy/static-scraping)
- [HTML Parsing](https://scrapfly.io/academy/html-parsing)
- [Dynamic Scraping](https://scrapfly.io/academy/dynamic-scraping)
- [Hidden API Scraping](https://scrapfly.io/academy/hidden-api-scraping)
- [Headless Browsers](https://scrapfly.io/academy/headless-browsers)
- [Hidden Web Data](https://scrapfly.io/academy/hidden-web-data)
- [JSON Parsing](https://scrapfly.io/academy/json-parsing)
- [Data Processing](https://scrapfly.io/academy/data-processing)
- [Scaling](https://scrapfly.io/academy/scaling)
- [Walkthrough Summary](https://scrapfly.io/academy/walkthrough-summary)
- [Scraper Blocking](https://scrapfly.io/academy/scraper-blocking)
- [Proxies](https://scrapfly.io/academy/proxies)

---

# Auto Refresh

 A crawl is a snapshot. **Auto refresh keeps that snapshot current without creating a new crawl.** On a period you choose, the crawl re-scrapes the URLs it already holds, keeps the ones that did not change, updates the ones that did, and drops the ones that no longer exist. The `crawler_uuid` never changes, so every URL, dashboard link, search index and integration already pointing at that crawl keeps working.

> **This is not a schedule** A [scheduled crawl](https://scrapfly.io/docs/crawler-api/schedule) creates a brand new crawl on every fire: a new `crawler_uuid`, a new artifact set, a new search index, each ageing out on its own retention clock. Use a schedule when you want a history of independent snapshots. Use refresh when you want *one* corpus that stays current.

## What one refresh run does

 A run walks the crawl's own URL index. It does not rediscover the site: pages the original crawl never reached are not visited, and depth, path filters and page limits are not re-applied. For every known URL:

- **Unchanged**: the page is re-scraped and its content fingerprint matches what the crawl already holds. Nothing is rewritten, nothing is re-embedded, nothing is re-indexed.
- **Updated**: the fingerprint differs. The new response is appended to the crawl's WARC, the URL's byte ranges are updated in place, and the page is re-indexed under a new generation.
- **Added**: a URL the run found that the crawl did not already hold, indexed like a normal first visit.
- **Removed**: the URL is gone upstream. It is dropped from the crawl's URL list and from the search index.
- **Failed**: the run could not fetch the URL. The page keeps the content it already had; a failed fetch never deletes data.

 Old bytes stay readable throughout: the crawl's WARC is append-only, so a run adds records rather than rewriting the archive, and a URL's byte ranges move only once its new record exists.

> **A refreshed crawl is not immutable** Without refresh, `GET /crawl/{crawler_uuid}/contents` returns the same bytes forever. With refresh on, **the same URL can return different content after a run**, and a URL that was in the crawl can disappear from it. If you cache, diff or hash crawl output downstream, key it on `refresh.generation`, not on the crawl UUID alone. Code that assumes a crawl is frozen will silently serve stale comparisons.

## Enable it at crawl time

 `refresh` is a boolean on the crawl configuration and defaults to `false`. `refresh_interval` is the period in seconds.

 ```
curl -X POST 'https://api.scrapfly.io/crawl?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -d '{
        "url": "https://web-scraping.dev/products",
        "page_limit": 500,
        "content_formats": ["markdown"],
        "refresh": true,
        "refresh_interval": 86400
    }'

```

 | Parameter | Type | Default | Notes |
|---|---|---|---|
| `refresh` | bool | `false` | Turns auto refresh on for this crawl. |
| `refresh_interval` | int (seconds) | `86400` | Between `3600` (1 hour) and `7776000` (90 days). Requires `refresh: true`. |

 The first run happens one interval *after* the crawl finishes, not immediately: the original crawl already produced fresh data. Use [a manual run](#run-now) if you need one sooner.

> **Why the floor is one hour** The interval decides the cost. A crawl refreshing every minute would re-scrape the whole site 1,440 times a day. Pick the slowest period that still catches the changes you care about: a documentation site rarely needs more than daily, a price list rarely more than hourly.

## Run one now

 `POST /crawl/{crawler_uuid}/refresh` starts one run immediately without touching the schedule. It answers as soon as the run is accepted; poll the status endpoint or the [history](#history) for the outcome.

 ```
curl -X POST 'https://api.scrapfly.io/crawl/{crawler_uuid}/refresh?key={{ YOUR_API_KEY }}'

```

 It is not idempotent. Calling it twice starts two re-scrapes and bills both, so do not retry it blindly on a timeout. A run that is already in flight is rejected with `ERR::CRAWLER::REFRESH_IN_PROGRESS` rather than queued.

## Change the schedule

 `PATCH /crawl/{crawler_uuid}/refresh` changes only the fields you send. Turning a crawl off keeps its interval for when you turn it back on. Turning refresh on for a crawl that was created without it is allowed: the crawl already holds the URL index a run walks.

 The body takes the same two keys `POST /crawl` takes, `refresh` and `refresh_interval`, so a crawl body and a later PATCH name the same things. The `enabled` / `interval_seconds` spelling belongs to the [state block](#state) this call answers with, not to its request. Unknown keys are rejected rather than ignored, so a typo is a `400` and never a silent no-op.

 ```
# Turn it on, daily
curl -X PATCH 'https://api.scrapfly.io/crawl/{crawler_uuid}/refresh?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -d '{ "refresh": true, "refresh_interval": 86400 }'

# Turn it off. The interval is kept for when it is turned back on.
curl -X PATCH 'https://api.scrapfly.io/crawl/{crawler_uuid}/refresh?key={{ YOUR_API_KEY }}' \
    -H 'Content-Type: application/json' \
    -d '{ "refresh": false }'

```

 | Field | Type | Notes |
|---|---|---|
| `refresh` | bool | Turns auto refresh on or off. Omit to leave it as it is. |
| `refresh_interval` | int | `3600` to `7776000`. Omit to leave the current period alone. Cannot be sent alongside `refresh: false`. |

## Reading the state

 Refresh state lives on the crawl, in the `refresh` block of `GET /crawl/{crawler_uuid}/status`:

 ```
{
  "status": "DONE",
  "is_success": true,
  "refresh": {
    "enabled": true,
    "interval_seconds": 86400,
    "status": "SCHEDULED",
    "generation": 4,
    "last_run_at": "2026-09-01T04:00:00Z",
    "next_run_at": "2026-09-02T04:00:00Z",
    "error": null,
    "history": [
      {
        "at": "2026-09-01T04:00:00Z",
        "generation": 4,
        "added": 3,
        "updated": 7,
        "removed": 1,
        "unchanged": 404,
        "failed": 0,
        "duration_ms": 44900,
        "search_status": "READY",
        "error": null,
        "sample_updated": ["https://web-scraping.dev/product/12"],
        "sample_removed": ["https://web-scraping.dev/product/99"]
      }
    ]
  }
}

```

 | Status | Meaning |
|---|---|
| `DISABLED` | The crawl does not refresh itself. Not an error. |
| `SCHEDULED` | A run is due at `next_run_at`. |
| `RUNNING` | A run is in flight. Content may change under you while this lasts. |
| `FAILED` | The last run failed; see `error`. The crawl keeps the content it had and the schedule stays armed, so the next period is attempted normally. |

 `generation` counts completed runs. It is the value to key downstream caches on: a crawl at generation 4 holds different bytes from the same crawl at generation 3.

## The activity timeline

 `GET /crawl/{crawler_uuid}/refresh/history` returns one row per run, oldest first. The same rows are inlined in the `refresh.history` block of the status response.

 ```
curl -G 'https://api.scrapfly.io/crawl/{crawler_uuid}/refresh/history' \
    --data-urlencode 'key={{ YOUR_API_KEY }}' \
    --data-urlencode 'limit=10'

```

 ```
{
  "crawler_uuid": "0198c4f2-1f3a-7a10-9d21-6f0b8a1c4e55",
  "history": [
    {
      "at": "2026-08-31T04:00:00Z",
      "generation": 3,
      "added": 0, "updated": 0, "removed": 0, "unchanged": 412, "failed": 0,
      "duration_ms": 41200,
      "search_status": "READY",
      "error": null,
      "sample_updated": [],
      "sample_removed": []
    },
    {
      "at": "2026-09-01T04:00:00Z",
      "generation": 4,
      "added": 3, "updated": 7, "removed": 1, "unchanged": 404, "failed": 0,
      "duration_ms": 44900,
      "search_status": "READY",
      "error": null,
      "sample_updated": ["https://web-scraping.dev/product/12"],
      "sample_removed": ["https://web-scraping.dev/product/99"]
    }
  ]
}

```

 | Field | Meaning |
|---|---|
| `at` | When the run completed. |
| `generation` | The generation this run produced. |
| `added` | URLs the run found that the crawl did not hold. |
| `updated` | Known URLs whose content changed and were re-indexed. |
| `removed` | Known URLs that no longer exist and were dropped. |
| `unchanged` | Re-scraped with identical content. No re-indexing. |
| `failed` | URLs the run could not fetch. They keep their previous content. |
| `duration_ms` | Wall time of the run. |
| `search_status` | Index state after the run, absent on a crawl without a search index. |
| `sample_updated` / `sample_removed` | Up to ten URLs each, as a sample. The full lists are not returned: a 5,000-page crawl would put 5,000 strings into every status poll. |

 A row where `added`, `updated` and `removed` are all zero means the site stood still. That run cost its scrapes and nothing else.

 The timeline keeps the **50 most recent runs**. Older rows are trimmed rather than paged: the history exists to show recent activity, not as an audit log. Archive it yourself if you need more.

## Getting told what changed

 A crawl whose [webhook](https://scrapfly.io/docs/crawler-api/webhook) subscribes to `crawler_updated` is called once per run that changed something, with the run's counts and the changed URLs in the body. Subscribe to it and nothing has to poll the timeline.

 The event is silent for a run where nothing moved, which is the common case on a stable site, and for a run that failed outright, which changed nothing either. Both are still recorded in `refresh.history`: the timeline is where every run shows up, the webhook where every change does.

## Refresh and the search index

 A crawl with [search](https://scrapfly.io/docs/crawler-api/search) enabled keeps one index across refreshes. A run re-embeds only the pages it marked added or updated and publishes them under a new index generation; unchanged pages keep the vectors they already had, and removed pages are deleted from the index.

 That is the whole economic argument for fingerprinting. A refresh that re-embedded every page would cost as much as the original crawl on every period. A site where nothing changed costs zero embeddings and zero index writes.

## What it costs

 **A refresh bills exactly like a crawl of the same pages.** Every URL the run fetches is a scrape, charged at the same rate as the original crawl: the scrape options it was created with (`unblocker`, `proxy_pool`, `country`, `headers`) carry over.

- An unchanged page still costs its scrape. There is no way to know a page did not change without fetching it. What it saves is the embedding and the index write.
- A run's cost is therefore roughly *pages in the crawl × the crawl's per-page cost*, every `refresh_interval`. A 500-page crawl refreshing daily is 15,000 page scrapes a month.
- Refresh runs do not consume a scheduler slot. They are not schedules and are not counted against your schedule quota.

 Costs are itemised per run in your project's usage, and the [Billing](https://scrapfly.io/docs/crawler-api/billing) page covers how a crawl's per-page cost is computed.

## SDKs and CLI

 ```
from scrapfly import ScrapflyClient, CrawlerConfig
from scrapfly.crawler import Crawl

client = ScrapflyClient(key="{{ YOUR_API_KEY }}")

crawl = Crawl(client, CrawlerConfig(
    url="https://web-scraping.dev/products",
    page_limit=500,
    content_formats=["markdown"],
    refresh=True,
    refresh_interval=86400,
)).crawl().wait()

# Run one now instead of waiting for the period.
state = crawl.refresh_now()
print(state.status, state.next_run_at)

# The activity timeline, newest last.
for entry in crawl.refresh_history(limit=10):
    print(entry.at, entry.added, entry.updated, entry.removed, entry.unchanged)

# Change the schedule later, or turn it off.
crawl.refresh_settings(enabled=False)

```

 The same three calls exist in every SDK: `crawl_refresh_now`, `crawl_refresh_settings` and `crawl_refresh_history` on the client, with `refresh_now()` / `refresh_settings()` / `refresh_history()` on the crawl object as sugar.

 ```
# Start a crawl that refreshes itself daily
scrapfly crawl start https://web-scraping.dev/products \
    --max-pages 500 --content-format markdown \
    --refresh --refresh-interval 86400

# Run one refresh right now
scrapfly crawl refresh "$UUID"

# Change the schedule instead of running one
scrapfly crawl refresh "$UUID" --enable --interval 604800
scrapfly crawl refresh "$UUID" --disable

# The activity timeline
scrapfly crawl refresh-history "$UUID" --limit 10 --pretty

```

## Errors

 | Code | When |
|---|---|
| `ERR::CRAWLER::REFRESH_NOT_ENABLED` | 422. The crawl does not refresh itself. Enable it with a PATCH first. |
| `ERR::CRAWLER::REFRESH_IN_PROGRESS` | 409. A run is already in flight. Runs are never queued; retry later. |
| `ERR::CRAWLER::REFRESH_INTERVAL_INVALID` | 400. `refresh_interval` is outside `3600` to `7776000`, or was sent alongside `refresh: false`. |

 The full catalog, with the HTTP status and retry guidance for each code, is on the [Errors](https://scrapfly.io/docs/crawler-api/errors) page.

## Next steps

- Search the refreshed corpus with [Search](https://scrapfly.io/docs/crawler-api/search).
- Compare with independent snapshots on the [Schedule](https://scrapfly.io/docs/crawler-api/schedule) page.
- Fetch the current bytes for a URL through [`/contents`](https://scrapfly.io/docs/crawler-api/results#query-content).
- Have a run's changes pushed to you with the `crawler_updated` [webhook](https://scrapfly.io/docs/crawler-api/webhook#events).
- See what a run adds to your bill on the [Billing](https://scrapfly.io/docs/crawler-api/billing) page.
