# Scrapfly Documentation

## Table of Contents

### Dashboard

- [Intro](https://scrapfly.io/docs)
- [Project](https://scrapfly.io/docs/project)
- [Account](https://scrapfly.io/docs/account)
- [Workspace & Team](https://scrapfly.io/docs/workspace-and-team)
- [Billing](https://scrapfly.io/docs/billing)

### Products

#### MCP Server

- [Getting Started](https://scrapfly.io/docs/mcp/getting-started)
- [Tools & API Spec](https://scrapfly.io/docs/mcp/tools)
- [Authentication](https://scrapfly.io/docs/mcp/authentication)
- [Examples & Use Cases](https://scrapfly.io/docs/mcp/examples)
- [FAQ](https://scrapfly.io/docs/mcp/faq)
##### Integrations

- [Overview](https://scrapfly.io/docs/mcp/integrations)
- [Claude Desktop](https://scrapfly.io/docs/mcp/integrations/claude-desktop)
- [Claude Code](https://scrapfly.io/docs/mcp/integrations/claude-code)
- [ChatGPT](https://scrapfly.io/docs/mcp/integrations/chatgpt)
- [Cursor](https://scrapfly.io/docs/mcp/integrations/cursor)
- [Cline](https://scrapfly.io/docs/mcp/integrations/cline)
- [Windsurf](https://scrapfly.io/docs/mcp/integrations/windsurf)
- [Zed](https://scrapfly.io/docs/mcp/integrations/zed)
- [Roo Code](https://scrapfly.io/docs/mcp/integrations/roo-code)
- [VS Code](https://scrapfly.io/docs/mcp/integrations/vscode)
- [LangChain](https://scrapfly.io/docs/mcp/integrations/langchain)
- [LlamaIndex](https://scrapfly.io/docs/mcp/integrations/llamaindex)
- [CrewAI](https://scrapfly.io/docs/mcp/integrations/crewai)
- [OpenAI](https://scrapfly.io/docs/mcp/integrations/openai)
- [n8n](https://scrapfly.io/docs/mcp/integrations/n8n)
- [Make](https://scrapfly.io/docs/mcp/integrations/make)
- [Zapier](https://scrapfly.io/docs/mcp/integrations/zapier)
- [Vapi AI](https://scrapfly.io/docs/mcp/integrations/vapi)
- [Agent Builder](https://scrapfly.io/docs/mcp/integrations/agent-builder)
- [Custom Client](https://scrapfly.io/docs/mcp/integrations/custom-client)


#### Web Scraping API

- [Getting Started](https://scrapfly.io/docs/scrape-api/getting-started)
- [API Specification]()
- [Monitoring](https://scrapfly.io/docs/monitoring)
- [Customize Request](https://scrapfly.io/docs/scrape-api/custom)
- [Debug](https://scrapfly.io/docs/scrape-api/debug)
- [Anti Scraping Protection](https://scrapfly.io/docs/scrape-api/anti-scraping-protection)
- [Proxy](https://scrapfly.io/docs/scrape-api/proxy)
- [Proxy Mode](https://scrapfly.io/docs/scrape-api/proxy-mode)
- [Proxy Mode - Screaming Frog](https://scrapfly.io/docs/scrape-api/proxy-mode/screaming-frog)
- [Proxy Mode - Apify](https://scrapfly.io/docs/scrape-api/proxy-mode/apify)
- [(Auto) Data Extraction](https://scrapfly.io/docs/scrape-api/extraction)
- [Javascript Rendering](https://scrapfly.io/docs/scrape-api/javascript-rendering)
- [Javascript Scenario](https://scrapfly.io/docs/scrape-api/javascript-scenario)
- [SSL](https://scrapfly.io/docs/scrape-api/ssl)
- [DNS](https://scrapfly.io/docs/scrape-api/dns)
- [Cache](https://scrapfly.io/docs/scrape-api/cache)
- [Batch (Multi-URL Scraping)](https://scrapfly.io/docs/scrape-api/batch)
- [Session](https://scrapfly.io/docs/scrape-api/session)
- [Webhook](https://scrapfly.io/docs/scrape-api/webhook)
- [Schedule](https://scrapfly.io/docs/scrape-api/schedule)
- [Screenshot](https://scrapfly.io/docs/scrape-api/screenshot)
- [Errors](https://scrapfly.io/docs/scrape-api/errors)
- [Timeout](https://scrapfly.io/docs/scrape-api/understand-timeout)
- [Throttling](https://scrapfly.io/docs/throttling)
- [Troubleshoot](https://scrapfly.io/docs/scrape-api/troubleshoot)
- [Billing](https://scrapfly.io/docs/scrape-api/billing)
- [FAQ](https://scrapfly.io/docs/scrape-api/faq)

#### Crawler API

- [Getting Started](https://scrapfly.io/docs/crawler-api/getting-started)
- [API Specification]()
- [Retrieving Results](https://scrapfly.io/docs/crawler-api/results)
- [WARC Format](https://scrapfly.io/docs/crawler-api/warc-format)
- [Data Extraction](https://scrapfly.io/docs/crawler-api/extraction-rules)
- [Webhook](https://scrapfly.io/docs/crawler-api/webhook)
- [Schedule](https://scrapfly.io/docs/crawler-api/schedule)
- [Billing](https://scrapfly.io/docs/crawler-api/billing)
- [Errors](https://scrapfly.io/docs/crawler-api/errors)
- [Troubleshoot](https://scrapfly.io/docs/crawler-api/troubleshoot)
- [FAQ](https://scrapfly.io/docs/crawler-api/faq)

#### Screenshot API

- [Getting Started](https://scrapfly.io/docs/screenshot-api/getting-started)
- [API Specification]()
- [Accessibility Testing](https://scrapfly.io/docs/screenshot-api/accessibility)
- [Webhook](https://scrapfly.io/docs/screenshot-api/webhook)
- [Schedule](https://scrapfly.io/docs/screenshot-api/schedule)
- [Billing](https://scrapfly.io/docs/screenshot-api/billing)
- [Errors](https://scrapfly.io/docs/screenshot-api/errors)

#### Extraction API

- [Getting Started](https://scrapfly.io/docs/extraction-api/getting-started)
- [API Specification]()
- [Rules Template](https://scrapfly.io/docs/extraction-api/rules-and-template)
- [LLM Extraction](https://scrapfly.io/docs/extraction-api/llm-prompt)
- [AI Auto Extraction](https://scrapfly.io/docs/extraction-api/automatic-ai)
- [Webhook](https://scrapfly.io/docs/extraction-api/webhook)
- [Billing](https://scrapfly.io/docs/extraction-api/billing)
- [Errors](https://scrapfly.io/docs/extraction-api/errors)
- [FAQ](https://scrapfly.io/docs/extraction-api/faq)

#### Data API


#### Proxy Saver

- [Getting Started](https://scrapfly.io/docs/proxy-saver/getting-started)
- [Fingerprints](https://scrapfly.io/docs/proxy-saver/fingerprints)
- [Optimizations](https://scrapfly.io/docs/proxy-saver/optimizations)
- [SSL Certificates](https://scrapfly.io/docs/proxy-saver/certificates)
- [Protocols](https://scrapfly.io/docs/proxy-saver/protocols)
- [Pacfile](https://scrapfly.io/docs/proxy-saver/pacfile)
- [Secure Credentials](https://scrapfly.io/docs/proxy-saver/security)
- [Billing](https://scrapfly.io/docs/proxy-saver/billing)

#### Cloud Browser API

- [Getting Started](https://scrapfly.io/docs/cloud-browser-api/getting-started)
- [Proxy & Geo-Targeting](https://scrapfly.io/docs/cloud-browser-api/proxy)
- [Unblock API](https://scrapfly.io/docs/cloud-browser-api/unblock)
- [Captcha Solver](https://scrapfly.io/docs/cloud-browser-api/captcha-solver)
- [File Downloads](https://scrapfly.io/docs/cloud-browser-api/file-downloads)
- [Session Resume](https://scrapfly.io/docs/cloud-browser-api/session-resume)
- [Human-in-the-Loop](https://scrapfly.io/docs/cloud-browser-api/human-in-the-loop)
- [Debug Mode](https://scrapfly.io/docs/cloud-browser-api/debug-mode)
- [Browser Extensions](https://scrapfly.io/docs/cloud-browser-api/extensions)
- [Native Browser MCP](https://scrapfly.io/docs/cloud-browser-api/mcp)
- [DevTools Protocol](https://scrapfly.io/docs/cloud-browser-api/cdp-reference)
##### Integrations

- [Puppeteer](https://scrapfly.io/docs/cloud-browser-api/puppeteer)
- [Playwright](https://scrapfly.io/docs/cloud-browser-api/playwright)
- [Selenium](https://scrapfly.io/docs/cloud-browser-api/selenium)
- [Vercel Agent Browser](https://scrapfly.io/docs/cloud-browser-api/agent-browser)
- [Browser Use](https://scrapfly.io/docs/cloud-browser-api/browser-use)
- [Stagehand](https://scrapfly.io/docs/cloud-browser-api/stagehand)

- [Billing](https://scrapfly.io/docs/cloud-browser-api/billing)
- [Errors](https://scrapfly.io/docs/cloud-browser-api/errors)


### Tools

- [Antibot Detector](https://scrapfly.io/docs/tools/antibot-detector)

### SDK

- [Golang](https://scrapfly.io/docs/sdk/golang)
- [Python](https://scrapfly.io/docs/sdk/python)
- [Rust](https://scrapfly.io/docs/sdk/rust)
- [TypeScript](https://scrapfly.io/docs/sdk/typescript)
- [Scrapy](https://scrapfly.io/docs/sdk/scrapy)

### Integrations

- [Getting Started](https://scrapfly.io/docs/integration/getting-started)
- [LangChain](https://scrapfly.io/docs/integration/langchain)
- [LlamaIndex](https://scrapfly.io/docs/integration/llamaindex)
- [CrewAI](https://scrapfly.io/docs/integration/crewai)
- [Zapier](https://scrapfly.io/docs/integration/zapier)
- [Make](https://scrapfly.io/docs/integration/make)
- [n8n](https://scrapfly.io/docs/integration/n8n)

### Academy

- [Overview](https://scrapfly.io/academy)
- [Web Scraping Overview](https://scrapfly.io/academy/scraping-overview)
- [Tools](https://scrapfly.io/academy/tools-overview)
- [Reverse Engineering](https://scrapfly.io/academy/reverse-engineering)
- [Static Scraping](https://scrapfly.io/academy/static-scraping)
- [HTML Parsing](https://scrapfly.io/academy/html-parsing)
- [Dynamic Scraping](https://scrapfly.io/academy/dynamic-scraping)
- [Hidden API Scraping](https://scrapfly.io/academy/hidden-api-scraping)
- [Headless Browsers](https://scrapfly.io/academy/headless-browsers)
- [Hidden Web Data](https://scrapfly.io/academy/hidden-web-data)
- [JSON Parsing](https://scrapfly.io/academy/json-parsing)
- [Data Processing](https://scrapfly.io/academy/data-processing)
- [Scaling](https://scrapfly.io/academy/scaling)
- [Walkthrough Summary](https://scrapfly.io/academy/walkthrough-summary)
- [Scraper Blocking](https://scrapfly.io/academy/scraper-blocking)
- [Proxies](https://scrapfly.io/academy/proxies)

---

# Cloud Browser File Downloads

 Cloud Browser supports automatic file download handling through a custom **CDP extension**. When the browser triggers a file download (PDFs, images, spreadsheets, etc.), you can retrieve the downloaded files using special CDP commands.

## How It Works

 When a navigation or user action triggers a file download in the browser:

1. The browser intercepts the download request
2. The file is saved to a temporary directory on the remote browser
3. Download events are emitted so you can track progress
4. You can retrieve the file content using the `ScrapiumBrowser.getDownloads` command

##### Key Points

- Files are returned as **base64-encoded strings**
- Maximum download size is **500 MB** per file, and your client's WebSocket message limit is reached well before it
- Downloads are automatically cleaned up when the session ends
- **Multiple file downloads** are fully supported in a single session
- Browser download popups and confirmations are **automatically bypassed** for seamless automation

 **Automation-Friendly:** Cloud Browser automatically disables browser download protection mechanisms (such as multiple download confirmation popups) that would normally interrupt automation scripts. Your downloads will proceed without any user interaction required.

## CDP Commands

 Cloud Browser extends the standard CDP protocol with custom `ScrapiumBrowser` commands for file download handling:

 | Command | Description |
|---|---|
| `ScrapiumBrowser.getDownloads` | Retrieve all downloaded files as base64-encoded content. Optionally delete files after retrieval. |
| `ScrapiumBrowser.getDownloadsMetadatas` | Get metadata (filename and size) for all downloaded files without retrieving content. |
| `ScrapiumBrowser.getDownload` | Retrieve **one** file by `filename`, returned once under `data`. Optionally delete it after retrieval. |
| `ScrapiumBrowser.getDownloadMetadatas` | Get the size of a single file by `filename`. |
| `ScrapiumBrowser.hasDownloads` | Report whether the download directory holds anything at all. |

 `getDownload` returns the file **once**, so its reply is ~1.34x the file rather than the ~2.7x `getDownloads` costs. Prefer it when you know which file you want.

### ScrapiumBrowser.getDownloads

 Retrieves all files currently in the download directory, encoded as base64 strings. This command takes no parameters.

#### Response

 ```
{
  "files": {
    "document.pdf": "JVBERi0xLjQKJe...",
    "report.xlsx": "UEsDBBQAAAA..."
  },
  "files_by_guid": {
    "9F2C1A4E-...": "JVBERi0xLjQKJe...",
    "B71D0E88-...": "UEsDBBQAAAA..."
  }
}
```

 The same content is returned under two keys:

- `files` is keyed by on-disk filename. Use it when filenames are unique.
- `files_by_guid` is keyed by the download GUID from `Browser.downloadWillBegin`. Chromium writes same-named downloads to the same path, so two files called `invoice.pdf` collapse to one entry in `files` while `files_by_guid` keeps both.

 **Both keys carry the full content**, so this response is roughly **2.7x the file size** on the wire: base64 adds a third, and the payload appears twice. A 40 MB download answers with about 107 MB of JSON in a single WebSocket message. See [Large Downloads and the IO Domain](#large-downloads) before retrieving anything above a few megabytes.

### ScrapiumBrowser.getDownloadsMetadatas

 Retrieves metadata for all downloaded files without transferring the actual content. Useful for checking what files are available before downloading.

#### Response

 ```
{
  "metadata": {
    "document.pdf": 245760,
    "report.xlsx": 102400
  }
}
```

## Code Examples

Download a file after clicking a download button:

   Puppeteer   Playwright JS   Python   Python Async

  ```
const puppeteer = require('puppeteer-core');

const API_KEY = '';
const BROWSER_WS = `wss://browser.scrapfly.io?api_key=${API_KEY}&proxy_pool=public_datacenter_pool`;

async function downloadFile() {
    const browser = await puppeteer.connect({ browserWSEndpoint: BROWSER_WS });
    const page = await browser.newPage();

    // Navigate and trigger download
    await page.goto('https://web-scraping.dev/file-download');
    await page.click('#download-btn');
    await new Promise(resolve => setTimeout(resolve, 3000));

    // Get the CDP session and retrieve downloads
    const client = await page.createCDPSession();
    const result = await client.send('ScrapiumBrowser.getDownloads');

    for (const [filename, base64Content] of Object.entries(result.files)) {
        const buffer = Buffer.from(base64Content, 'base64');
        require('fs').writeFileSync(filename, buffer);
        console.log(`Saved: ${filename} (${buffer.length} bytes)`);
    }
    await browser.close();
}

downloadFile();
```

 ```
const { chromium } = require('playwright');

const API_KEY = '';
const BROWSER_WS = `wss://browser.scrapfly.io?api_key=${API_KEY}&proxy_pool=public_datacenter_pool`;

async function downloadFile() {
    const browser = await chromium.connectOverCDP(BROWSER_WS);
    const context = browser.contexts()[0];
    const page = context.pages()[0] || await context.newPage();
    const cdpSession = await page.context().newCDPSession(page);

    // Navigate and trigger download
    await page.goto('https://web-scraping.dev/file-download');
    await page.click('#download-btn');
    await page.waitForTimeout(3000);

    // Retrieve and save files
    const downloads = await cdpSession.send('ScrapiumBrowser.getDownloads');
    for (const [filename, base64Content] of Object.entries(downloads.files)) {
        const buffer = Buffer.from(base64Content, 'base64');
        require('fs').writeFileSync(filename, buffer);
        console.log(`Saved ${filename} (${buffer.length} bytes)`);
    }
    await browser.close();
}

downloadFile();
```

 ```
import base64
from playwright.sync_api import sync_playwright

API_KEY = ''
BROWSER_WS = f'wss://browser.scrapfly.io?api_key={API_KEY}&proxy_pool=public_datacenter_pool'

def download_file():
    with sync_playwright() as p:
        browser = p.chromium.connect_over_cdp(BROWSER_WS)
        context = browser.contexts[0]
        page = context.pages[0] if context.pages else context.new_page()
        cdp_session = context.new_cdp_session(page)

        # Navigate and trigger download
        page.goto('https://web-scraping.dev/file-download')
        page.click('#download-btn')
        page.wait_for_timeout(3000)

        # Retrieve and save files
        downloads = cdp_session.send('ScrapiumBrowser.getDownloads')
        for filename, base64_content in downloads['files'].items():
            file_bytes = base64.b64decode(base64_content)
            with open(filename, 'wb') as f:
                f.write(file_bytes)
            print(f'Saved {filename} ({len(file_bytes)} bytes)')
        browser.close()

download_file()
```

 ```
import asyncio
import base64
from playwright.async_api import async_playwright

API_KEY = ''
BROWSER_WS = f'wss://browser.scrapfly.io?api_key={API_KEY}&proxy_pool=public_datacenter_pool'

async def download_file():
    async with async_playwright() as p:
        browser = await p.chromium.connect_over_cdp(BROWSER_WS)
        context = browser.contexts[0]
        page = context.pages[0] if context.pages else await context.new_page()
        cdp_session = await context.new_cdp_session(page)

        # Navigate and trigger download
        await page.goto('https://web-scraping.dev/file-download')
        await page.click('#download-btn')
        await page.wait_for_timeout(3000)

        # Retrieve and save files
        downloads = await cdp_session.send('ScrapiumBrowser.getDownloads')
        for filename, base64_content in downloads['files'].items():
            file_bytes = base64.b64decode(base64_content)
            with open(filename, 'wb') as f:
                f.write(file_bytes)
            print(f'Saved {filename} ({len(file_bytes)} bytes)')
        await browser.close()

asyncio.run(download_file())
```

## Monitoring Download Progress

 Cloud Browser emits standard Chrome CDP events that you can listen to for tracking download progress:

 | Event | Description |
|---|---|
| `Browser.downloadWillBegin` | Fired when a download is about to start. Contains URL and suggested filename. |
| `Browser.downloadProgress` | Fired periodically during download. Contains progress state and bytes received. |

### Listening to Download Events

 ```
const puppeteer = require('puppeteer-core');

const API_KEY = '';
const BROWSER_WS = `wss://browser.scrapfly.io?api_key=${API_KEY}&proxy_pool=public_datacenter_pool`;

async function monitorDownloads() {
    const browser = await puppeteer.connect({
        browserWSEndpoint: BROWSER_WS,
    });

    const page = await browser.newPage();
    const client = await page.createCDPSession();

    // Listen for download start
    client.on('Browser.downloadWillBegin', (event) => {
        console.log('Download starting:', {
            url: event.url,
            filename: event.suggestedFilename,
            guid: event.guid
        });
    });

    // Listen for download progress
    client.on('Browser.downloadProgress', (event) => {
        console.log('Download progress:', {
            guid: event.guid,
            state: event.state,  // 'inProgress', 'completed', 'canceled'
            receivedBytes: event.receivedBytes,
            totalBytes: event.totalBytes
        });

        if (event.state === 'completed') {
            console.log(`Download completed: ${event.receivedBytes} bytes`);
        }
    });

    // Navigate and trigger download
    await page.goto('https://web-scraping.dev/file-download');
    await page.click('#download-btn');

    // Wait for download to complete
    await new Promise(resolve => setTimeout(resolve, 5000));

    // Retrieve the file
    const result = await client.send('ScrapiumBrowser.getDownloads');
    console.log('Downloaded files:', Object.keys(result.files));

    await browser.close();
}

monitorDownloads();
```

## Multiple File Downloads

 Cloud Browser fully supports downloading multiple files in a single session. Unlike regular browsers that may prompt for confirmation when triggering multiple downloads, Cloud Browser automatically accepts all downloads without interruption.

### Example: Downloading Multiple Files

 ```
const puppeteer = require('puppeteer-core');

const API_KEY = '';
const BROWSER_WS = `wss://browser.scrapfly.io?api_key=${API_KEY}&proxy_pool=public_datacenter_pool`;

async function downloadMultipleFiles() {
    const browser = await puppeteer.connect({
        browserWSEndpoint: BROWSER_WS,
    });

    const page = await browser.newPage();
    const client = await page.createCDPSession();

    await page.goto('https://web-scraping.dev/file-download');

    // Trigger multiple downloads - no confirmation popups will appear
    await page.click('#download-btn');
    await page.click('#download-pdf');
    await page.click('#download-csv');

    // Wait for all downloads to complete
    await new Promise(resolve => setTimeout(resolve, 5000));

    // Retrieve all downloaded files at once
    const result = await client.send('ScrapiumBrowser.getDownloads');

    console.log(`Downloaded ${Object.keys(result.files).length} files:`);
    for (const [filename, base64Content] of Object.entries(result.files)) {
        const buffer = Buffer.from(base64Content, 'base64');
        require('fs').writeFileSync(filename, buffer);
        console.log(`  - ${filename} (${buffer.length} bytes)`);
    }

    await browser.close();
}

downloadMultipleFiles();
```

 **Tip:** When downloading multiple files, you can use `getDownloadsMetadatas` to check how many files are ready before retrieving them all with `getDownloads`.

## Large Downloads and the IO Domain

 `ScrapiumBrowser.getDownloads` answers with the whole file inside one CDP message. That is convenient for a spreadsheet and a problem for a video: every WebSocket client enforces a maximum message size, and a message above it is a protocol error, not a slow read. Your client closes the connection with code `1009` (message too big) and the browser session dies with it, usually reported as a closed target rather than as a size error.

### Why the limit is a wall and not a slow path

 CDP is a request/response protocol carried over a single WebSocket connection, and a WebSocket delivers **messages, not bytes**. The wire format does allow a message to be split into fragments, but fragmentation is invisible above the socket: the client library reassembles every fragment into one buffer before handing the message over, because a CDP reply is a single JSON document that cannot be parsed until its last byte has arrived. The receiver therefore holds the entire reply in memory, which is why every implementation caps how large that reply may be.

Three consequences follow:

- **You cannot ask for the reply in pieces.** `getDownloads` has no offset or range parameter, and CDP defines no chunking of its own. One command produces one message, and its size is decided by the file, not by you.
- **Going over the cap kills the connection, it does not fail the call.** A message larger than the receiver accepts is a protocol violation, so the receiver closes the socket with status `1009`. There is no response object to inspect, no error field to branch on, and nothing to retry on that connection.
- **The connection is the session.** Your CDP session lives on that socket, so every attached target dies with it. Puppeteer and Playwright report the aftermath rather than the cause: `Target closed`, `Session closed`, or `Protocol error (...): Connection closed`. Check the WebSocket close code to confirm it was a size failure.

 Raising the cap moves the wall, it does not remove it, and it costs memory in your process for every byte you let through. The only way to move a large payload out of the browser without ever building a large message is the [IO domain](#io-domain), covered below.

### Work out the wire size first

 Multiply the file size by **2.7**: base64 costs a third, and the reply carries the content under both `files` and `files_by_guid`.

 | File size | getDownloads reply | Safe in one message? |
|---|---|---|
| 5 MB | ~13 MB | Yes |
| 20 MB | ~53 MB | Yes, with a raised client limit |
| 40 MB | ~107 MB | Only above a 128 MB client limit |
| 100 MB | ~267 MB | No, stream it instead |

 Call `ScrapiumBrowser.getDownloadsMetadatas` before `getDownloads` to learn the real sizes, then decide which path to take.

### Raise your client's message limit

 This is headroom for small and mid-sized files, not an answer for large ones. Raise it so a legitimate reply is not cut off by a default that was never meant for file transfer, then use the IO domain for anything larger.

 Puppeteer and Playwright already raise this well above their library default (256 MB at the time of writing), so a few tens of megabytes work out of the box. A hand-rolled CDP client usually does not, and the Python defaults are small enough to fail on the first real download.

 | Client | Setting | Default |
|---|---|---|
| Node `ws` | `maxPayload` | 100 MB |
| Python `websockets` | `max_size` | 1 MB |
| Python `websocket-client` | no limit, but the whole message is buffered in memory | unlimited |

 ```
// Node: a raw CDP client that can accept a large getDownloads reply
const WebSocket = require('ws');

const ws = new WebSocket(BROWSER_WS, {
    maxPayload: 512 * 1024 * 1024, // 512 MB, well above the largest expected reply
    perMessageDeflate: false,      // the payload is base64, compression only costs CPU
});
```

 ```
# Python: websockets defaults to 1 MB, which fails on any real download
import websockets

async with websockets.connect(BROWSER_WS, max_size=512 * 1024 * 1024) as ws:
    ...
```

 Raising the limit means your process buffers the entire reply in memory. Budget roughly three times the file size for the client alone, and prefer streaming past a few tens of megabytes.

### Stream large payloads with the IO domain

 Chromium's [IO domain](https://scrapfly.io/docs/cloud-browser-api/cdp-reference/IO) exists for exactly this case, and inside CDP it is the only alternative. Instead of one large reply, a command hands back a **stream handle** and you drain it with repeated `IO.read` calls, so no single message is ever large. Every other CDP route for a bulky payload ends up here anyway: `Page.printToPDF` with `transferMode: 'ReturnAsStream'`, the tracing streams, and `IO.resolveBlob` all return an IO handle. Cloud Browser passes the IO domain through to Chromium unchanged.

 **Read 10 MB per chunk for a large body.** That is the balance point: large enough that a 1 GB file is 100 round trips instead of 8000, small enough that the resulting message stays far below any reasonable cap. Budget about **1.34x** your chunk size for the message, since binary chunks come back base64-encoded, so a 10 MB read produces a reply near 13.4 MB. Keep your client's message limit above that, and drop to a smaller chunk if you are memory-bound or on a lossy link where a failed read costs a full re-read.

 **`Network.getResponseBody` is not an option for a large body.** Chromium bounds how much of a response it keeps for the inspector, so a body of a few tens of megabytes is dropped before you can ask for it and the call fails with `Request content was evicted from inspector cache`. Raising your message limit does not help: the bytes are already gone. Fetch the URL in the page and stream the result instead, as below.

 **And it corrupts binary bodies served as text, without any error.** `getResponseBody` returns `base64Encoded: true` only when Chromium classified the resource as binary. For a text content type it returns the body *already decoded* to a string, so bytes that were not valid text in the declared charset come back re-encoded and no longer match the file. The usual symptom is a body about 1.5x the size it should be, because every byte above `0x7F` became two UTF-8 bytes. Always check `base64Encoded` before treating the result as bytes, and serve or fetch binary payloads as `application/octet-stream`.

 | Command | Parameters | Returns |
|---|---|---|
| `<a href="/docs/cloud-browser-api/cdp-reference/IO#method-read">IO.read</a>` | `handle` (required), `size` (maximum bytes for this chunk), `offset` (seek before reading) | `data`, `base64Encoded`, `eof` |
| `<a href="/docs/cloud-browser-api/cdp-reference/IO#method-close">IO.close</a>` | `handle` | Nothing. Discards the backing storage. |
| `<a href="/docs/cloud-browser-api/cdp-reference/IO#method-resolveBlob">IO.resolveBlob</a>` | `objectId` of a `Blob` | `uuid`, readable as handle `blob:<uuid>` |

Four rules that cover every mistake we see:

1. **Loop until `eof` is true.** An empty `data` without `eof` is not the end of the stream.
2. **`size` is a maximum, not a promise.** A read may return fewer bytes than requested and still not be at the end. Ask for 10 MB on a large body and size your buffers from what actually came back.
3. **Decode per chunk when `base64Encoded` is true**, then concatenate the decoded bytes. Concatenating base64 text and decoding once only works if every chunk happens to be 4-character aligned.
4. **Always `IO.close`**, including on your error path. An abandoned handle pins the payload in the browser until the session ends.

#### Recipe: fetch a file in the page and stream it out

 This never triggers a browser download and never produces a large message. The page fetches the file into a `Blob`, `IO.resolveBlob` turns it into a handle, and you read it in chunks of your choosing.

 ```
const fs = require('fs');

async function streamFile(client, url, destination) {
    // 1. Fetch inside the page and keep the bytes there as a Blob.
    const evaluated = await client.send('Runtime.evaluate', {
        expression: 'fetch(' + JSON.stringify(url) + ').then(r => r.blob())',
        awaitPromise: true,
    });

    // 2. Turn the Blob into an IO stream handle.
    const resolved = await client.send('IO.resolveBlob', {
        objectId: evaluated.result.objectId,
    });
    const handle = 'blob:' + resolved.uuid;

    // 3. Drain it. Nothing crossing the wire is larger than one chunk.
    const chunks = [];
    try {
        for (;;) {
            const read = await client.send('IO.read', { handle, size: 10 * 1024 * 1024 });
            if (read.data) {
                chunks.push(Buffer.from(read.data, read.base64Encoded ? 'base64' : 'utf8'));
            }
            if (read.eof) break;
        }
    } finally {
        // 4. Always release the handle, success or failure.
        await client.send('IO.close', { handle }).catch(() => {});
    }

    fs.writeFileSync(destination, Buffer.concat(chunks));
    return chunks.reduce((total, chunk) => total + chunk.length, 0);
}
```

 ```
import base64

CHUNK = 10 * 1024 * 1024  # ~13.4 MB per reply once base64-encoded

def stream_file(cdp_session, url, destination):
    # 1. Fetch inside the page and keep the bytes there as a Blob.
    evaluated = cdp_session.send('Runtime.evaluate', {
        'expression': f'fetch({url!r}).then(r => r.blob())',
        'awaitPromise': True,
    })

    # 2. Turn the Blob into an IO stream handle.
    resolved = cdp_session.send('IO.resolveBlob', {
        'objectId': evaluated['result']['objectId'],
    })
    handle = 'blob:' + resolved['uuid']

    # 3. Drain it, decoding each chunk on its own.
    written = 0
    try:
        with open(destination, 'wb') as out:
            while True:
                read = cdp_session.send('IO.read', {'handle': handle, 'size': CHUNK})
                data = read.get('data') or ''
                if data:
                    raw = base64.b64decode(data) if read.get('base64Encoded') else data.encode()
                    out.write(raw)
                    written += len(raw)
                if read.get('eof'):
                    break
    finally:
        # 4. Always release the handle, success or failure.
        try:
            cdp_session.send('IO.close', {'handle': handle})
        except Exception:
            pass

    return written
```

#### Recipe: a large PDF without a large message

 `Page.printToPDF` accepts `transferMode: 'ReturnAsStream'`, which returns a handle instead of an inline base64 document. Drain it with the same loop.

 ```
const printed = await client.send('Page.printToPDF', {
    transferMode: 'ReturnAsStream',
    printBackground: true,
});

// printed.stream is an IO handle: read it in chunks, then IO.close it.
const handle = printed.stream;
```

### Stream the downloads themselves

 The two recipes above avoid the browser's download machinery entirely. When you do need a real download (a click behind a login, an anti-bot flow) and the file is large, add `downloads_stream=1` to your Cloud Browser WebSocket URL. The session then answers `getDownloads` with a different, smaller shape:

- each download's content **once**, under `files_by_guid`, with `guid_names` mapping every GUID to its filename
- anything above ~8 MB of base64 handed back under `streams` as `{stream, size, name}` instead of inline content

 ```
// Opt in on the connection; nothing later in the session can turn it on.
const browser = await puppeteer.connect({
    browserWSEndpoint: 'wss://cloud-browser.scrapfly.io/devtools/browser?key=' + KEY + '&downloads_stream=1',
});
```

 ```
{
  "files": {},
  "files_by_guid": {
    "9F2C1A4E-...": "JVBERi0xLjQKJe..."
  },
  "guid_names": {
    "9F2C1A4E-...": "invoice.pdf",
    "B71D0E88-...": "export.csv"
  },
  "streams": {
    "B71D0E88-...": { "stream": "scrapfly-download:1", "size": 91250688, "name": "export.csv" }
  }
}
```

 A `stream` value is an ordinary IO handle: drain it with the same `IO.read` / `IO.close` loop as above. Three differences from a Chrome-owned stream are worth knowing:

- **Reads are sequential.** Passing `offset` is an error, because the bytes are released as they are read rather than kept for a seek.
- **`size` counts base64 characters here, not raw bytes.** On a Chrome-owned handle `size` is raw bytes and the reply inflates to ~1.34x; on a download handle the content is already base64, so a 10 MB read is a ~10 MB message carrying ~7.5 MB of file. Chunks are capped at 32 MB and default to 4 MB when you send no `size`. The cut is aligned to a base64 group, so chunks concatenate without decoding each one.
- **An abandoned handle is reclaimed after 120 seconds**, and reaching `eof` releases it. Both make a forgotten `IO.close` survivable, but sending it is still the contract.

 Download handles are namespaced `scrapfly-download:<n>` and only those are served by the session. `blob:` handles, `Page.printToPDF` streams and every other IO handle reach Chromium untouched, so the opt-in never changes how the rest of the IO domain behaves.

 **Which path to choose:** use `getDownloads` when the file is small and arrived through a real browser download you needed anyway (a click behind a login, an anti-bot flow). Use `downloads_stream=1` when that same flow produces a large file. Use `IO.resolveBlob` when you know the URL and only need the bytes, without a download at all.

## Common Use Cases

##### PDF Downloads

 Download PDFs generated by web applications, such as invoices, reports, or tickets that are created dynamically after form submission or authentication.

##### Export Files

 Retrieve data exports (CSV, Excel, JSON) from dashboards and analytics platforms that require browser interaction to generate.

##### Generated Images

 Download images that are generated on-demand, such as charts, QR codes, or dynamically created graphics.

##### Protected Documents

 Access documents behind authentication or CAPTCHA protection that can only be downloaded through a real browser session.

## Best Practices

##### Wait for Downloads to Complete

 Always wait for the download to complete before calling `getDownloads`. You can either use a fixed timeout, listen for the `Browser.downloadProgress` event with state `completed`, or poll `getDownloadsMetadatas` until files appear.

##### Automatic Cleanup

 Downloaded files are **automatically cleaned up** when the browser session ends (either when you close the CDP connection or when the session times out with `auto_close=true`). You don't need to manually delete files unless you want to clear downloads during a long-running session.

##### Check File Size First

 For large files, call `getDownloadsMetadatas` first to check the file size. Budget **2.7x the file size** for the reply: base64 adds a third, and the content is returned under both `files` and `files_by_guid`. A 40 MB download answers with about 107 MB in one WebSocket message, so past a few tens of megabytes stream it instead. See [Large Downloads and the IO Domain](#large-downloads).

##### Handle Download Failures

 Downloads can fail or be canceled. Listen to the `Browser.downloadProgress` event and check for `state: 'canceled'` or `state: 'failed'` to handle errors gracefully.

## Limitations

- **Maximum file size:** 500 MB per file. A larger download stays on disk and is listed by `getDownloadsMetadatas`, but retrieving it fails with `download too large: "name" (N bytes) exceeds the 524288000 byte limit                    per file`. One oversized file fails the whole `getDownloads` call, so pull the others by name with `getDownload`
- **Client message limit binds first:** 500 MB of file is ~1.34 GB of reply in the default shape, far above what any client accepts, so use `downloads_stream=1` well before the cap
- **Session-scoped:** Downloads are only available during the session that triggered them
- **Auto-cleanup:** All downloads are deleted when the browser session ends
- **Transfer overhead:** a `getDownloads` reply is ~2.7x the file size (base64 adds a third, and the content is returned under both `files` and `files_by_guid`)
- **One message per reply:** the whole payload arrives in a single WebSocket message, so your client's message limit is what a large download hits first. See [Large Downloads and the IO Domain](#large-downloads)

## Related Documentation

- [Cloud Browser Getting Started](https://scrapfly.io/docs/cloud-browser-api/getting-started) - Connection and basic usage
- [Puppeteer Integration](https://scrapfly.io/docs/cloud-browser-api/puppeteer) - Full Puppeteer setup guide
- [Playwright Integration](https://scrapfly.io/docs/cloud-browser-api/playwright) - Full Playwright setup guide
- [IO Domain Reference](https://scrapfly.io/docs/cloud-browser-api/cdp-reference/IO) - `read`, `close` and `resolveBlob` schemas for streaming large payloads
- [Cloud Browser Billing](https://scrapfly.io/docs/cloud-browser-api/billing) - Pricing and cost optimization
- [Error Reference](https://scrapfly.io/docs/cloud-browser-api/errors) - Troubleshooting guide
