     [Blog](https://scrapfly.io/blog)   /  [hidden-api](https://scrapfly.io/blog/tag/hidden-api)   /  [How to Scrape Goodreads for Book Data, Ratings, and Reviews](https://scrapfly.io/blog/posts/how-to-scrape-goodreads)   # How to Scrape Goodreads for Book Data, Ratings, and Reviews

 by [Ziad Shamndy](https://scrapfly.io/blog/author/ziad) Sep 29, 2026 26 min read [\#hidden-api](https://scrapfly.io/blog/tag/hidden-api) [\#python](https://scrapfly.io/blog/tag/python) [\#scrapeguide](https://scrapfly.io/blog/tag/scrapeguide) 

 [  ](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-goodreads "Share on LinkedIn") [  ](https://x.com/intent/tweet?url=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-goodreads&text=How%20to%20Scrape%20Goodreads%20for%20Book%20Data%2C%20Ratings%2C%20and%20Reviews "Share on X") [  ](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-goodreads "Share on Facebook")    

 

 

Summarize this article with

 [  ](https://chat.openai.com/?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-goodreads) [  ](https://claude.ai/new?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-goodreads) [  ](https://x.com/i/grok?text=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-goodreads) [  ](https://www.perplexity.ai/search/new?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-goodreads) [  ](https://www.google.com/search?udm=50&aep=11&q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-goodreads) 



         

A Goodreads book page can show the right title and still leave a scraper with incomplete reviews if it depends on one CSS selector. Goodreads exposes the same record through three layers: JSON-LD, an embedded Next.js payload, and the visible page markup.

This guide builds a Goodreads scraper in Python using the open source Scrapfly Goodreads scraper. It covers how to fetch pages with Scrapfly, parse the JSON-LD `Book` block, read reviews from the Apollo state inside `__NEXT_DATA__`, and discover books from Goodreads list and search pages.



[**Scrapfly Goodreads scraper**github.com/scrapfly/scrapfly-scrapers/tree/main/goodreads-scraper](https://github.com/scrapfly/scrapfly-scrapers/tree/main/goodreads-scraper)

## Key Takeaways

- Goodreads book pages expose the same core record through JSON-LD, an embedded `__NEXT_DATA__` payload, and visible HTML, so a resilient scraper reads structured data first and uses markup for the fields JSON-LD leaves out.
- A live test on 2026-07-31 found that Goodreads book, author, and list pages returned a normal response to a plain HTTP request, while the `/search` endpoint returned an AWS WAF challenge instead of HTML.
- Review data on a book page is a sample, not the full archive. A single page load carries around 30 reviews, while the real total sits in the Apollo `getReviews.totalCount` field.
- Goodreads list pages and search results share the same server rendered `tr[itemtype="http://schema.org/Book"]` row markup, so one parser handles both.
- Goodreads stopped issuing new public API developer keys in December 2020, which is the main reason scraping remains the practical path to structured book data.

**Get web scraping tips in your inbox**Trusted by 100K+ developers and 30K+ enterprises. Unsubscribe anytime.







## What Data Can a Goodreads Scraper Extract?

A Goodreads scraper can pull book title, author, average rating, ratings count, review count, and a sample of reviews from a single book page request. Genre tags, awards, ISBN, and format details arrive with that same request.

The table below maps each field to the source the scraper in this guide reads it from:

| Field | Source used by the scraper | Fallback source | Page type |
|---|---|---|---|
| Title | `h1[data-testid="bookTitle"]` | JSON-LD `name` | Book page |
| Author | JSON-LD `author[0].name` | Not applicable | Book page |
| Average rating | JSON-LD `aggregateRating.ratingValue` | Not applicable | Book page |
| Ratings count | JSON-LD `aggregateRating.ratingCount` | Not applicable | Book page |
| Review count | JSON-LD `aggregateRating.reviewCount` | Not applicable | Book page |
| Genres | `div[data-testid="genresList"]` markup | Not applicable | Book page |
| Review text | Apollo `Review.text` | Not applicable | Book page |
| List row rating | `span.minirating` text | Not applicable | List and search pages |

Fetching [The Great Gatsby](https://www.goodreads.com/book/show/4671.The_Great_Gatsby) for the repository's `results/book.json` returned a title of "The Great Gatsby", an author of F. Scott Fitzgerald, an average rating of 3.93, a ratings count of 6,123,188, and a review count of 138,742.

Every count above is a point in time sample, since Goodreads ratings and reviews change continuously.

The same JSON-LD block also carried the ISBN, page count, format, language, awards, and cover image. The embedded `__NEXT_DATA__` payload went further, exposing normalized records for the book, its parent work, the primary contributor, and a batch of reviews.

[Web Scraping with PythonIntroduction tutorial to web scraping with Python. How to collect and parse public data. Challenges, best practices and an example project.](https://scrapfly.io/blog/posts/web-scraping-with-python)

With the field map in place, the next step is getting that HTML in the first place, since Goodreads treats different page types differently.



## How Do You Fetch a Goodreads Book Page with Scrapfly?



ScrapFly's [Web Scraping API](https://scrapfly.io/products/web-scraping-api) is a single HTTP endpoint for collecting web data at scale, with a **98% success rate** across **130M+ proxies in 120+ countries**.

- [Anti-Scraping Protection bypass](https://scrapfly.io/docs/scrape-api/unblocker) - automatically defeats Cloudflare, DataDome, PerimeterX, Akamai, and 90+ other bot systems.
- [Smart proxy rotation](https://scrapfly.io/docs/scrape-api/proxy) - residential and datacenter pools with country and ASN level geo-targeting.
- [JavaScript rendering](https://scrapfly.io/docs/scrape-api/javascript-rendering) - render SPAs and dynamic pages through real cloud browsers.
- [Browser automation scenarios](https://scrapfly.io/docs/scrape-api/javascript-scenario) - scroll, click, fill forms, and wait for elements without managing a browser fleet.
- [Format conversion](https://scrapfly.io/docs/scrape-api/getting-started#api_param_format) - return pages as HTML, JSON, clean text, or LLM ready Markdown.
- [Session management](https://scrapfly.io/docs/scrape-api/session) - keep cookies, headers, and IPs consistent across multi step flows.
- [Smart caching](https://scrapfly.io/docs/scrape-api/getting-started#api_param_cache) - cache successful responses to cut cost on repeat scraping jobs.
- [Python](https://scrapfly.io/docs/sdk/python), [TypeScript](https://scrapfly.io/docs/sdk/typescript), [Scrapy](https://scrapfly.io/docs/sdk/scrapy), and [no-code integrations](https://scrapfly.io/docs/integration/getting-started) including Make, n8n, Zapier, LangChain, and LlamaIndex.

Every code sample in this guide comes from the `goodreads.py` module in the scraper repository. It relies on `scrapfly-sdk`, which sends the requests and returns a ready to use [parsel](https://pypi.org/project/parsel/) selector, and [loguru](https://pypi.org/project/loguru/) for logging. Install both with one command:

shell```shell
pip install "scrapfly-sdk[all]" loguru
```



The module starts by creating one shared Scrapfly client and a base configuration that every request reuses.



python```python
import json
import os
import re
from datetime import datetime, timedelta, timezone
from typing import Dict, List, Optional
from urllib.parse import parse_qsl, quote_plus, urlencode, urljoin, urlparse, urlunparse

from loguru import logger as log
from parsel import Selector
from scrapfly import ScrapeConfig, ScrapflyClient, ScrapeApiResponse

SCRAPFLY = ScrapflyClient(key=os.environ["SCRAPFLY_KEY"])
BASE_CONFIG = {
    # goodreads.com requires Anti Scraping Protection bypass feature.
    "asp": True,
}

# Goodreads timestamps are epoch milliseconds, and go negative for pre 1970 publications
EPOCH = datetime(1970, 1, 1, tzinfo=timezone.utc)
```



### How Should `ScrapeConfig` Classify Goodreads Responses?

The scraper uses no custom classifier. It enables `asp` on every request so challenge pages are resolved before parsing, and it treats missing structured data as a failure.

In a 2026-07-31 test, a plain HTTP client got a 200 response on book, author, and list pages, but `/search` returned HTTP 202 with an `x-amzn-waf-action: challenge` header and an empty body. Setting `asp=True` once in `BASE_CONFIG` covers that case on every page type.

A 200 response can still be the wrong page. `_find_book_ld_json` returns an empty dictionary when no `Book` block exists, and `_apollo` logs a warning when `__NEXT_DATA__` is missing.

Fetching a book page then takes a single call.

python```python
async def scrape_book(url: str) -> Dict:
    """scrape a single book page and return parsed book data"""
    log.info("scraping book {}", url)
    result = await SCRAPFLY.async_scrape(ScrapeConfig(url, **BASE_CONFIG))
    return parse_book(result)
```



`async_scrape` returns a `ScrapeApiResponse`, and its `selector` property gives the parsing functions in the next section a parsel selector over the returned HTML.

[Scrapfly's getting started guide](https://scrapfly.io/docs/scrape-api/getting-started) covers the rest of the `ScrapeConfig` options, including proxy country selection and JavaScript rendering.

With a response in hand, the next step is turning that HTML into structured fields, starting with the structured data Goodreads embeds directly in the page.



## How Do You Parse Goodreads JSON-LD and `__NEXT_DATA__`?

Parsing is done in a specific order. JSON-LD provides the base structure for the book, visual markup populates the fields left out by JSON-LD, and the Apollo state stored in `__NEXT_DATA__` provides reviews. Ratings and counts come from structured data because it changes less often than the page markup.

### How Do You Extract Goodreads `Book` and `aggregateRating` JSON-LD?

A Goodreads book page carries a `script[type="application/ld+json"]` tag with a `Book` object. That object nests an `aggregateRating` block with the numbers a scraper usually needs first.

python```python
def _find_book_ld_json(sel: Selector) -> Dict:
    """find the Book JSON-LD block embedded in the page"""
    for script in sel.xpath('//script[@type="application/ld+json"]/text()').getall():
        try:
            data = json.loads(script)
        except json.JSONDecodeError:
            continue
        if data.get("@type") == "Book":
            return data
    return {}
```



This helper walks every JSON-LD block on the page and returns the first one typed as `Book`. A malformed block is skipped instead of crashing the parse, and an empty dictionary comes back when no `Book` block exists.

The book parser then combines that JSON-LD with a handful of scoped CSS selectors.



python```python
def parse_book(response: ScrapeApiResponse) -> Dict:
    """parse book page and return book data"""
    sel = response.selector
    ld = _find_book_ld_json(sel)
    rating = ld.get("aggregateRating") or {}
    authors = ld.get("author") or []

    description = " ".join(
        t.strip() for t in sel.css('div[data-testid="description"] .Formatted::text').getall() if t.strip()
    )
    genres = [
        g.strip()
        for g in sel.css(
            'div[data-testid="genresList"] span.BookPageMetadataSection__genreButton a ' "span.Button__labelItem::text"
        ).getall()
        if g.strip()
    ]
    awards = [a.strip() for a in (ld.get("awards") or "").split(",") if a.strip()]

    return {
        "url": sel.xpath('//link[@rel="canonical"]/@href').get(),
        "title": sel.css('h1[data-testid="bookTitle"]::text').get() or ld.get("name"),
        "author": {
            "name": authors[0].get("name") if authors else None,
            "url": authors[0].get("url") if authors else None,
        },
        "description": description or None,
        "image_url": ld.get("image") or sel.xpath('//meta[@property="og:image"]/@content').get(),
        "genres": genres or None,
        "num_pages": ld.get("numberOfPages"),
        "format": ld.get("bookFormat"),
        "language": ld.get("inLanguage"),
        "isbn": ld.get("isbn"),
        "awards": awards or None,
        "first_published": sel.css('p[data-testid="publicationInfo"]::text').get(),
        "rating": {
            "average": rating.get("ratingValue"),
            "ratings_count": rating.get("ratingCount"),
            "reviews_count": rating.get("reviewCount"),
        },
    }
```



Ratings, page count, author, ISBN, book format, and language are taken from JSON-LD directly. Awards, in turn, come from JSON-LD as a string separated by commas, and the parser transforms them into an array.

For title and image, we have two sources each. Title is taken from JSON-LD `name`, if the `bookTitle` heading is not present. Image, if missing in JSON-LD, is retrieved from `og:image` meta.

The `results/book.json` output file from the repository illustrates the structure of this function's result for "The Great Gatsby":



Output Examplejson```json
{
  "url": "https://www.goodreads.com/book/show/41733839-the-great-gatsby",
  "title": "The Great Gatsby",
  "author": {
    "name": "F. Scott Fitzgerald",
    "url": "https://www.goodreads.com/author/show/3190.F_Scott_Fitzgerald"
  },
  "genres": [
    "Classics",
    "Fiction",
    "School",
    "Historical Fiction",
    "Romance",
    "Literature",
    "Novels"
  ],
  "num_pages": 180,
  "format": "Paperback",
  "language": "English",
  "isbn": "9780743273565",
  "first_published": "First published April 10, 1925",
  "rating": {
    "average": 3.93,
    "ratings_count": 6123188,
    "reviews_count": 138742
  }
}
```







The description, image, and awards fields are trimmed from this excerpt to keep it short. The canonical URL points at a different edition ID than the one requested, because Goodreads canonicalizes the legacy `/book/show/4671` path to its preferred edition; the legacy URL itself returns 200 without redirecting.

[How to Scrape Hidden Web DataThe visible HTML doesn't always represent the whole dataset available on the page. In this article, we'll be taking a look at scraping of hidden web data. What is it and how can we scrape it using Python?](https://scrapfly.io/blog/posts/how-to-scrape-hidden-web-data)

### How Do You Read Goodreads Apollo State From `__NEXT_DATA__`?

The `script#__NEXT_DATA__` tag holds a full Apollo cache under `props.pageProps.apolloState`, keyed by typed references such as `Book:kca://book/...` and `Review:kca://review/...`. Records point at each other through `__ref` pointers rather than nesting.



python```python
def _apollo(response: ScrapeApiResponse) -> Dict:
    """parse the normalized Apollo cache out of __NEXT_DATA__

    Goodreads exposes no standalone __APOLLO_STATE__ global, the cache only reaches the page
    through the Next.js payload, so a missing payload means the Apollo layer is unavailable.
    """
    raw = response.selector.css("script#__NEXT_DATA__::text").get()
    if not raw:
        log.warning("no __NEXT_DATA__ on {}", response.context["url"])
        return {}
    return json.loads(raw).get("props", {}).get("pageProps", {}).get("apolloState", {}) or {}


def _resolve(apollo: Dict, node) -> Dict:
    """follow an Apollo __ref pointer to the record it names"""
    if isinstance(node, dict) and "__ref" in node:
        return apollo.get(node["__ref"]) or {}
    return node if isinstance(node, dict) else {}
```



`_apollo` pulls the cache out of the Next.js payload and logs a warning instead of raising when the payload is missing. `_resolve` follows a single `__ref` pointer to the record it names, and passes plain dictionaries through untouched.

A standalone `__APOLLO_STATE__` global variable was not present on the tested page, only the state nested inside `__NEXT_DATA__`. That is why `_apollo` reads the Next.js payload directly.

### When Should Goodreads CSS Selectors Be the Fallback?

CSS selectors should only fill in what structured data cannot. Visible markup on Goodreads changes more often than JSON-LD, so `parse_book` uses CSS selectors for fields JSON-LD does not carry, such as the description, genres, and first published line.

The title is the one field read from both sources. The parser takes the `bookTitle` heading first and falls back to the JSON-LD `name` when the heading is missing, so a markup change there does not blank the field.

Each CSS selector the parser uses is anchored to a `data-testid` container. That anchoring matters. Goodreads reuses generic attributes like `data-testid="name"` for both the author link near the top of the page and every reviewer name further down.

A selector without a scoped parent risks grabbing a reviewer name or a review excerpt instead of the book field it was meant for. Reading the author from JSON-LD sidesteps that collision entirely.

[Guide to Parsel - the Best HTML Parsing in PythonLearn to extract data from websites with Parsel, a Python library for HTML parsing using CSS selectors and XPath.](https://scrapfly.io/blog/posts/guide-to-html-parsing-with-parsel-python)

With the book record covered, reviews need their own extraction logic, since a page wide selector risks blending fields from two different reviews together.



## How Do You Extract Goodreads Reviews Without Mixing Records?

The reliable way to extract Goodreads reviews is to skip the review cards and read the `Review` records from the Apollo state. Each record carries its own rating, text, and a pointer to its reviewer, so a review's text can never end up attached to someone else's name.

Review texts come in as HTML fragments and time stamps come as epoch millisecond values, so a couple of helper functions process them.



python```python
def _plain_text(html: Optional[str]) -> Optional[str]:
    """flatten a Goodreads review body, which arrives as an HTML fragment, into plain text"""
    if not html:
        return None
    return " ".join(Selector(text=html).xpath("string()").get("").split()) or None


def _iso(epoch_ms: Optional[int]) -> Optional[str]:
    """convert a Goodreads epoch milliseconds timestamp to an ISO date string"""
    if not isinstance(epoch_ms, int):
        return None
    return (EPOCH + timedelta(milliseconds=epoch_ms)).isoformat()
```



`_iso` adds a `timedelta` to the `EPOCH` constant rather than calling `datetime.fromtimestamp`. That keeps negative timestamps working, which Goodreads uses for publications dated before 1970.

### Which Goodreads Review Card Selectors Were Observed Live?

The tested page rendered 30 `article.ReviewCard` containers, with reviewer names under `div[data-testid="name"]` and text under `div.TruncatedContent__text`. Unscoped card selectors risk pairing one reviewer's name with another's text.

The scraper skips the cards and reads `ROOT_QUERY.getReviews` instead, resolving each edge to its `Review` record and creator.



python```python
def parse_reviews(response: ScrapeApiResponse) -> List[Dict]:
    """parse the review sample of a book page from its Apollo Review records

    Reviews are resolved through the getReviews root query, so every review belongs to the
    requested book and each reviewer stays attached to their own review. The page ships a
    sample of the reviews rather than all of them: the review URL exposes no working
    pagination parameter, and the real total is reported as getReviews.totalCount.
    """
    apollo = _apollo(response)
    connection = (apollo.get("ROOT_QUERY") or {}).get("getReviews") or {}
    reviews = []
    for edge in connection.get("edges") or []:
        review = _resolve(apollo, edge.get("node"))
        if not review.get("id"):
            continue
        creator = _resolve(apollo, review.get("creator"))
        reviews.append(
            {
                "review_id": review["id"],
                "reviewer": creator.get("name"),
                "reviewer_url": creator.get("webUrl"),
                # a text review left without stars is reported as 0, which is not a rating
                "rating": review.get("rating") or None,
                "text": _plain_text(review.get("text")),
                "created_at": _iso(review.get("createdAt")),
                "updated_at": _iso(review.get("updatedAt")),
                "likes": review.get("likeCount"),
                "comments": review.get("commentCount"),
                "spoiler": review.get("spoilerStatus"),
            }
        )
    log.success(f"parsed {len(reviews)} reviews of {connection.get('totalCount')} total")
    return reviews


async def scrape_reviews(url: str) -> List[Dict]:
    """scrape the review sample a Goodreads book page ships with"""
    log.info("scraping reviews of {}", url)
    result = await SCRAPFLY.async_scrape(ScrapeConfig(url, **BASE_CONFIG))
    return parse_reviews(result)
```



A `0` rating becomes `None`, since Goodreads uses zero for reviews left without stars. The log line reports `totalCount` so the gap between sample and archive stays visible.

Scraping The Hunger Games reviews page returned 30 records. Here is one from `results/reviews.json`, trimmed:

json```json
{
  "review_id": "kca://review:goodreads/amzn1.gr.review:goodreads.v1.z_A1Ex53N4FbaAFWT0Tgfg",
  "reviewer": "Nataliya",
  "reviewer_url": "https://www.goodreads.com/user/show/3672777-nataliya",
  "rating": 4,
  "text": "Suzanne Collins has balls ovaries of steel to make us willingly cheer for a teenage girl to kill other children. In a YA book. ...",
  "created_at": "2010-12-08T00:00:07+00:00",
  "updated_at": "2023-05-12T05:50:30.478000+00:00",
  "likes": 1455,
  "comments": 167,
  "spoiler": false
}
```



Full review text belongs to the reviewer, so a production scraper should store it rather than republish it in bulk.

Thirty reviews is a sample, not the archive. The review URL exposed no working pagination parameter during testing, so reaching further reviews is outside the scope of this guide.

With single book pages covered end to end, the next step is finding book URLs to feed that parser at scale, starting with Goodreads list pages.



Scrapfly

#### Scale your web scraping effortlessly

Scrapfly handles proxies, browsers, and anti-bot bypass — so you can focus on data.

[Try Free →](https://scrapfly.io/register)## How Do You Scrape Goodreads Lists and Discover Book URLs?

A Goodreads list page is a fast way to discover book URLs in bulk, since one page can carry up to 100 book rows in a single server rendered table.

The workflow runs in two stages, parsing the list into book stubs first, then optionally enriching each URL through the book scraper already covered above.

### How Do Goodreads `tableList` Rows Expose Titles, Authors, and Ratings?

Every row of a list page is marked up by `itemtype="http://schema.org/Book"`. The title, author, ranking, and rating data all lie within that row; however, no data is supplied as JSON data, which means that numbers need to be extracted from the text using regular expressions.

Unlike book pages, the list page does not contain any JSON-LD or `__NEXT_DATA__` block, so that row markup is the only available source.

There are two functions that deal with cleanup work.



python```python
def _to_int(value: Optional[str]) -> Optional[int]:
    """convert a string like '6,103,353' to an int, stripping non-digit characters"""
    if not value:
        return None
    digits = re.sub(r"[^\d]", "", value)
    return int(digits) if digits else None


def _clean_url(href: Optional[str]) -> Optional[str]:
    """absolutize a Goodreads href and drop its tracking query

    Search result links carry from_search, qid and rank parameters that belong to one search
    response, so keeping them would make the same book look like a different URL.
    """
    if not href:
        return None
    return urljoin("https://www.goodreads.com", urlparse(href)._replace(query="", fragment="").geturl())
```



`_clean_url` does more than absolutize a relative link. Dropping the query string means the same book always maps to the same URL, which is what makes deduplication work later.

The list parser then reads every row with scoped selectors and a few regular expressions.



python```python
def parse_list(response: ScrapeApiResponse) -> List[Dict]:
    """parse a list page and return book stubs (title, url, author, rating, etc.) found on it"""
    sel = response.selector
    books = []
    for row in sel.css("tr[itemtype='http://schema.org/Book']"):
        row_html = row.get()
        rating_text = row.css("span.minirating::text").get() or ""
        rating_m = re.search(r"([\d.]+)\s*avg rating", rating_text)
        ratings_m = re.search(r"([\d,]+)\s*ratings", rating_text)
        score_m = re.search(r"score:\s*([\d,]+)", row_html)
        votes_m = re.search(r"([\d,]+)\s*people voted", row_html)
        book_path = row.css("a.bookTitle::attr(href)").get()

        books.append(
            {
                "rank": _to_int(row.css("td.number::text").get()),
                "title": row.css("a.bookTitle span[itemprop='name']::text").get(),
                "url": _clean_url(book_path),
                "author": row.css("span[itemprop='author'] span[itemprop='name']::text").get(),
                "author_url": _clean_url(row.css("a.authorName::attr(href)").get()),
                "image_url": row.css("img.bookCover::attr(src)").get(),
                "avg_rating": float(rating_m.group(1)) if rating_m else None,
                "ratings_count": _to_int(ratings_m.group(1)) if ratings_m else None,
                "score": _to_int(score_m.group(1)) if score_m else None,
                "votes": _to_int(votes_m.group(1)) if votes_m else None,
            }
        )
    log.success(f"parsed {len(books)} books from the list")
    return books
```



The `span.minirating` text reads like "4.26 avg rating, 7,102,528 ratings", so two patterns split it into a float and an integer. List score and vote counts sit elsewhere in the row, which is why those patterns run against the full row HTML.

Fetching the [Books That Everyone Should Read At Least Once list](https://www.goodreads.com/list/show/264.Books_That_Everyone_Should_Read_At_Least_Once) returned 100 rows, with "To Kill a Mockingbird" by Harper Lee at rank 1.

A list row is only a stub, so `scrape_list` can optionally feed every row URL back into `scrape_book` for the full record.

python```python
async def scrape_list(url: str, enrich: bool = False) -> List[Dict]:
    """
    scrape a book list page and return book stubs found on it.
    if enrich=True, additionally scrape each book's page for full details.
    """
    log.info("scraping list {}", url)
    result = await SCRAPFLY.async_scrape(ScrapeConfig(url, **BASE_CONFIG))
    stubs = parse_list(result)

    if not enrich:
        return stubs

    books = []
    for stub in stubs:
        book_url = stub.get("url")
        if not book_url:
            continue
        books.append(await scrape_book(book_url))
    return books
```



With `enrich=False` the function returns stubs from a single request. With `enrich=True` it makes one extra request per row, so a 100 row list costs 101 requests.

### How Should Goodreads Author, Search, and Shelf Pages Be Added?

Each extra page type should only be added once it has been tested, and in the repository scraper only search has been. Search results render with the same `schema.org/Book` table rows as list pages, so `parse_list` reads them without changes. Rank, score, and votes are list page columns and stay `None` on search rows.

Search results are paginated through plain `page=` links, so two helpers build page URLs and read the last page number.

python```python
def _page_url(url: str, page: int) -> str:
    """set the page query parameter of a Goodreads search URL"""
    parts = urlparse(url)
    query = [(key, value) for key, value in parse_qsl(parts.query) if key != "page"] + [("page", str(page))]
    return urlunparse(parts._replace(query=urlencode(query)))

def _total_pages(response: ScrapeApiResponse) -> int:
    """read the last page number out of the plain link list Goodreads paginates with"""
    pages = [int(text) for text in response.selector.css('a[href*="page="]::text').getall() if text.strip().isdigit()]
    return max(pages) if pages else 1
```



The search scraper fetches the first page, reads the page count from it, then fetches the remaining pages concurrently.



python```python
async def scrape_search(query: str, max_pages: int = 2) -> List[Dict]:
    """scrape book stubs from Goodreads search results

    Search results are rendered with the same table markup as list pages, so the rows are
    read by parse_list. Rank, score and votes are list page columns and stay None here.
    """
    base_url = f"https://www.goodreads.com/search?q={quote_plus(query)}"

    log.info(f"scraping the first search page for '{query}'")
    first_page = await SCRAPFLY.async_scrape(ScrapeConfig(base_url, **BASE_CONFIG))
    books = parse_list(first_page)
    total_pages = min(_total_pages(first_page), max_pages)

    log.info(f"scraping search pagination, remaining ({total_pages - 1}) more pages")
    to_scrape = [ScrapeConfig(_page_url(base_url, page), **BASE_CONFIG) for page in range(2, total_pages + 1)]
    async for result in SCRAPFLY.concurrent_scrape(to_scrape):
        if not isinstance(result, ScrapeApiResponse):
            continue
        try:
            books.extend(parse_list(result))
        except Exception as e:
            log.error(f"failed to scrape search page: {e}")

    # the same book can show up on more than one search page
    unique = list({book["url"]: book for book in books if book.get("url")}.values())
    log.success(f"scraped {len(unique)} books from search")
    return unique
```



The failure to retrieve a page results in skipping it instead of halting the whole process, and the final command filters out duplicated books from different pages. A search for "dune" returned 40 unique books across two pages, starting with "Dune (Dune, #1)" by Frank Herbert.

The repository scraper does not cover author or shelf pages, so this guide skips them rather than publish untested selectors.

With book, review, list, and search parsing in place, the last step is keeping the pipeline running when Goodreads changes.



## How Do You Build a Resilient Goodreads Scraper Pipeline?

A resilient Goodreads scraper needs bounded pagination, failures that skip a page instead of stopping the run, and typed validation on every parsed record. The repository covers these with a `run.py` script that bounds every crawl and a `test.py` suite that validates each output against a schema.

### How Should Goodreads Pagination, Retries, and Rate Limits Be Handled?

Pagination should be explicitly limited, and any failures in scraping should be skipped, not endlessly retried. The function `scrape_search` explicitly limits the number of pages via `max_pages` and skips any non-`ScrapeApiResponse` result in the `concurrent_scrape` loop.

The scraper itself doesn't implement any retry mechanism. Retries live in the test suite instead: every test is marked `flaky(reruns=3, reruns_delay=30)`, so a failing test is retried up to three times, 30 seconds apart, as the next section's example shows.

`run.py` demonstrates the expected way to invoke the scrapers, without any enrichment and limited to two pages. It also enables debug mode of Scrapfly and saves the results in `results` folder.



python```python
from pathlib import Path
import asyncio
import json
import goodreads

output = Path(__file__).parent / "results"
output.mkdir(exist_ok=True)


async def run():
    goodreads.BASE_CONFIG["debug"] = True

    print("running Goodreads.com scrape and saving results to ./results directory")

    url = "https://www.goodreads.com/book/show/4671.The_Great_Gatsby"
    book = await goodreads.scrape_book(url)
    output.joinpath("book.json").write_text(json.dumps(book, indent=2, ensure_ascii=False), encoding="utf-8")

    url = "https://www.goodreads.com/book/show/2767052/reviews"
    reviews = await goodreads.scrape_reviews(url)
    output.joinpath("reviews.json").write_text(json.dumps(reviews, indent=2, ensure_ascii=False), encoding="utf-8")

    url = "https://www.goodreads.com/list/show/264.Books_That_Everyone_Should_Read"
    book_list = await goodreads.scrape_list(url, enrich=False)
    output.joinpath("list.json").write_text(json.dumps(book_list, indent=2, ensure_ascii=False), encoding="utf-8")

    search = await goodreads.scrape_search("dune", max_pages=2)
    output.joinpath("search.json").write_text(json.dumps(search, indent=2, ensure_ascii=False), encoding="utf-8")


if __name__ == "__main__":
    asyncio.run(run())
```



Set the `SCRAPFLY_KEY` environment variable before running it. The `debug` flag logs request details to the Scrapfly dashboard, and the saved JSON files double as reference output, so parser changes can be checked without hitting the live site.

Goodreads has not published a rate limit, so keep concurrency low and treat any figure found elsewhere as unverified.

### How Should Goodreads Schema Drift Be Tested?

Schema drift on Goodreads is a documented risk, not a hypothetical one. The independent [maria-antoniak/goodreads-scraper](https://github.com/maria-antoniak/goodreads-scraper) project posted a January 2025 notice stating that a Goodreads redesign broke large parts of its scraper, and that the project would not be updated further.

The repository's test suite guards against that kind of silent break with [Cerberus](https://pypi.org/project/Cerberus/) schemas. The book schema requires a non empty title and author name and types the rating fields as a float average and integer counts, so a missing title or a count that arrives as a string fails the test.

The review test goes further than type checks.



python```python
REVIEW_SCHEMA = {
    "review_id": {"type": "string", "minlength": 1},
    "reviewer": {"type": "string", "nullable": True},
    "reviewer_url": {"type": "string", "nullable": True},
    "rating": {"type": "integer", "nullable": True},
    "text": {"type": "string", "nullable": True},
    "created_at": {"type": "string", "nullable": True},
    "updated_at": {"type": "string", "nullable": True},
    "likes": {"type": "integer", "nullable": True},
    "comments": {"type": "integer", "nullable": True},
    "spoiler": {"type": "boolean", "nullable": True},
}


@pytest.mark.asyncio
@pytest.mark.flaky(reruns=3, reruns_delay=30)
async def test_review_scraping():
    url = "https://www.goodreads.com/book/show/2767052/reviews"
    results = await goodreads.scrape_reviews(url)
    validator = Validator(REVIEW_SCHEMA, allow_unknown=True)
    for item in results:
        validate_or_fail(item, validator)
    assert len(results) >= 10
    assert len({item["review_id"] for item in results}) == len(results)
    # each review has to keep its own reviewer and text, a mixed record shows up as a blank one
    assert all(item["reviewer"] and item["text"] for item in results)
```



The schema checks types, and the assertions check behavior. Review IDs must be unique, and every review must keep both a reviewer and text, since a mixed record shows up as a blank one.

The search test applies the same idea to URL cleanup, asserting that no discovered book URL still carries a `?` query string. That kind of check catches drift early, well before a production pipeline silently starts returning empty or duplicated fields.



## FAQ

Does Goodreads Still Issue Public API Developer Keys?Goodreads stated on its Developers Group page that it stopped issuing new developer keys for its public API on December 8, 2020, and planned to retire the existing developer tools.







Can Goodreads Book Metadata Be Scraped Without JavaScript?On the 2026-07-31 test run, a plain HTTP request with a browser style user agent returned JSON-LD, `__NEXT_DATA__`, and full review card markup directly in the HTML, with no JavaScript execution involved.

Treat that as one observed page, not a guarantee across every Goodreads route or client environment.







What Is the Difference Between a Goodreads Rating and a Review?A rating is the numeric score rolled into `aggregateRating.ratingValue` and `ratingCount`, submitted with a single click. A review is a separate text record tracked under `reviewCount` and the Apollo `Review` objects covered earlier, and not every rating includes one.









## Conclusion

A resilient Goodreads scraper reads the book record from JSON-LD, takes reviews from the Apollo state inside `__NEXT_DATA__`, and uses scoped CSS selectors for the title heading, description, genres, and first published line, with JSON-LD `name` as the title fallback.

List and search pages carry no structured data, so their parser reads the row markup directly. Author and shelf pages were not tested.

List and search rows are stubs, not full records. Route every discovered URL back through the same validated book parser rather than trusting the rating text shown in a table row.

Scrapfly's [Web Scraping API](https://scrapfly.io/products/web-scraping-api) handles the proxy rotation and anti scraping protection this guide leaned on for the challenge prone `/search` endpoint, so a pipeline built on it does not need to rebuild that layer from scratch.



Legal Disclaimer and PrecautionsThis tutorial covers popular web scraping techniques for education. Interacting with public servers requires diligence and respect:

- Do not scrape at rates that could damage the website.
- Do not scrape data that's not available publicly.
- Do not store PII of EU citizens protected by GDPR.
- Do not repurpose *entire* public datasets which can be illegal in some countries.

Scrapfly does not offer legal advice but these are good general rules to follow. For more you should consult a lawyer.

 

   [  Add as a preferred source ](https://google.com/preferences/source?q=scrapfly.io) Table of Contents















 

  Table of Contents- [Key Takeaways](#key-takeaways)
- [What Data Can a Goodreads Scraper Extract?](#what-data-can-a-goodreads-scraper-extract)
- [How Do You Fetch a Goodreads Book Page with Scrapfly?](#how-do-you-fetch-a-goodreads-book-page-with-scrapfly)
- [How Should ScrapeConfig Classify Goodreads Responses?](#how-should-scrapeconfig-classify-goodreads-responses)
- [How Do You Parse Goodreads JSON-LD and \_\_NEXT\_DATA\_\_?](#how-do-you-parse-goodreads-json-ld-and-next-data)
- [How Do You Extract Goodreads Book and aggregateRating JSON-LD?](#how-do-you-extract-goodreads-book-and-aggregaterating-json-ld)
- [How Do You Read Goodreads Apollo State From \_\_NEXT\_DATA\_\_?](#how-do-you-read-goodreads-apollo-state-from-next-data)
- [When Should Goodreads CSS Selectors Be the Fallback?](#when-should-goodreads-css-selectors-be-the-fallback)
- [How Do You Extract Goodreads Reviews Without Mixing Records?](#how-do-you-extract-goodreads-reviews-without-mixing-records)
- [Which Goodreads Review Card Selectors Were Observed Live?](#which-goodreads-review-card-selectors-were-observed-live)
- [How Do You Scrape Goodreads Lists and Discover Book URLs?](#how-do-you-scrape-goodreads-lists-and-discover-book-urls)
- [How Do Goodreads tableList Rows Expose Titles, Authors, and Ratings?](#how-do-goodreads-tablelist-rows-expose-titles-authors-and-ratings)
- [How Should Goodreads Author, Search, and Shelf Pages Be Added?](#how-should-goodreads-author-search-and-shelf-pages-be-added)
- [How Do You Build a Resilient Goodreads Scraper Pipeline?](#how-do-you-build-a-resilient-goodreads-scraper-pipeline)
- [How Should Goodreads Pagination, Retries, and Rate Limits Be Handled?](#how-should-goodreads-pagination-retries-and-rate-limits-be-handled)
- [How Should Goodreads Schema Drift Be Tested?](#how-should-goodreads-schema-drift-be-tested)
- [FAQ](#faq)
- [Conclusion](#conclusion)
 
    Join the Newsletter  Get monthly web scraping insights 

 

  



Scale Your Web Scraping

Anti-bot bypass, browser rendering, and rotating proxies, all in one API. Start with 1,000 free credits.

  No credit card required  1,000 free API credits  Anti-bot bypass included 

 [Start Free](https://scrapfly.io/register) [View Docs](https://scrapfly.io/docs/onboarding) 

 Not ready? Get our newsletter instead. 

 

 ## Related Articles

 [  

 data-parsing css-selectors 

### Parsing HTML with CSS Selectors

Introduction to using CSS selectors to parse web-scraped content. Best practices, available tools and common challenges ...

 

 ](https://scrapfly.io/blog/posts/parsing-html-with-css) [  

 python crawling 

### Guide to List Crawling: Everything You Need to Know

Complete list crawling tutorial assess site defenses, bypass anti-bot systems, choose tools (Beautiful Soup, Playwright,...

 

 ](https://scrapfly.io/blog/posts/guide-to-list-crawling) [     

 python hidden-api 

### How to Scrape Google Play App Reviews and Data

Scrape Google Play app metadata, ratings, and the full review set with Python, past the few-hundred-review ceiling the f...

 

 ](https://scrapfly.io/blog/posts/how-to-scrape-google-play-app-reviews-and-data) 

  



   



 Scale your web scraping effortlessly, **1,000 free credits** [Start Free](https://scrapfly.io/register)