     [Blog](https://scrapfly.io/blog)   /  [python](https://scrapfly.io/blog/tag/python)   /  [How to Scrape News Articles from Any Website](https://scrapfly.io/blog/posts/how-to-scrape-news-articles)   # How to Scrape News Articles from Any Website

 by [Mohab Yousry](https://scrapfly.io/blog/author/mohab-yousry-9396552a) Sep 29, 2026 19 min read [\#python](https://scrapfly.io/blog/tag/python) [\#scrapeguide](https://scrapfly.io/blog/tag/scrapeguide) 

 [  ](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-news-articles "Share on LinkedIn") [  ](https://x.com/intent/tweet?url=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-news-articles&text=How%20to%20Scrape%20News%20Articles%20from%20Any%20Website "Share on X") [  ](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-news-articles "Share on Facebook")    

 

 

Summarize this article with

 [  ](https://chat.openai.com/?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-news-articles) [  ](https://claude.ai/new?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-news-articles) [  ](https://x.com/i/grok?text=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-news-articles) [  ](https://www.perplexity.ai/search/new?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-news-articles) [  ](https://www.google.com/search?udm=50&aep=11&q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fhow-to-scrape-news-articles) 



         

   **Web Scraping API — Format Conversion**Get results in data formats that suit you — HTML, Markdown, JSON and more.

 

 [ Learn More  ](https://scrapfly.io/products/web-scraping-api#features) [  Docs ](https://scrapfly.io/docs/scrape-api/getting-started#features) 

 

 

Scraping one news article is a five-line script. The pain starts at the hundredth article, on ten different sites, half of them behind Cloudflare or a metered paywall, where the neat per-site selector you wrote against the first outlet quietly breaks on the second.

This guide covers extracting article fields from any news site, finding article URLs at scale with RSS and sitemaps, getting clean text for analysis, and getting through JavaScript rendering and anti-bot systems on the outlets that actually give scrapers trouble, plus where a paywall is a hard stop.



## Key Takeaways

- Extract headline, author, publish date, body text, and section from a single article with `requests` and Newspaper4k, the actively maintained fork of the abandoned Newspaper3k
- Drop to targeted BeautifulSoup selectors only for the specific fields Newspaper4k's generic parser misses on a given page
- Discover article URLs at scale through RSS feeds (`feedparser`) and XML news sitemaps rather than crawling category pages
- Get clean article text, free of navigation, ads, and related-story blocks, with a boilerplate-removal library such as trafilatura instead of a raw `.get_text()` call
- Route JavaScript-rendered pages and Cloudflare or DataDome protected outlets through a headless browser or a managed anti-bot layer like Scrapfly's Web Scraping API with `render_js` and `asp`
- Replace per-site selectors with a generic or AI extraction model when scraping across many sites, since N sites written as N selectors means N things that break on the next redesign
- Respect robots.txt, rate limit your requests, and treat paywalled article text as off-limits for republishing, since news copy is copyrighted even when the page itself is public

**Get web scraping tips in your inbox**Trusted by 100K+ developers and 30K+ enterprises. Unsubscribe anytime.







## What Data Can You Extract From a News Article?

The standard fields on a news article are headline, author, publish date, body text, and section or tags, plus a canonical URL, a lead image, and any related-article links the page carries.

Each field earns its place for a specific downstream reason. The publish date drives time-series analysis, like tracking coverage volume around an event. Author and section support filtering a large corpus down to a specific beat or byline. The canonical URL is what you deduplicate on when the same story gets syndicated to three different paths on the same domain.

Body text deserves a second note. Presence is not the same as quality. A field can be technically populated and still be useless if it is stitched together from three paragraphs of navigation, a newsletter signup block, and two related-story teasers along with the actual article. Getting a clean body is its own problem, covered later in this guide, separate from just locating where the text lives on the page.

If the goal is pulling recent headlines from a provider's index rather than scraping article pages directly, that is a different job with a different tool. The guide below covers querying an aggregator for news items instead of parsing publisher pages.

[Guide to Google News API and AlternativesIn a world of endless information, accessing news data efficiently can be vital for many businesses. Google News has been a trusted news aggregator for a while, though the discontinuation of the official Google News API has many searching for new ways to tap into its potential. Many are left...](https://scrapfly.io/blog/posts/guide-to-google-news-api-and-alternatives)

With the target fields defined, the next decision is how to get them using a script built for one site or a generic extractor built for many.



## Should You Use a Scraping Script or AI Extraction for News?

A hand-written per-site script gives precision and full control but needs one script per site and breaks when a layout changes. Generic or AI extraction works across arbitrary sites and tolerates layout drift at a higher per-page cost and with output that needs validation. Most real projects end up using both.

Per-site selectors, meaning `requests` plus BeautifulSoup tuned to one site's DOM are cheap and precise once written. Their weakness is durability. A frontend redesign can rename a class or restructure a container, and the selector breaks loudly or worse silently grabs the wrong element.

Generic parsers and AI extraction take the opposite tradeoff. Newspaper4k, trafilatura, readability, and LLM-prompt or auto-extraction models all locate the body and metadata without a hand-tuned selector, at a higher per-page cost and with output worth spot-checking, especially from LLM-based extraction.

The honest crossover point is scale and volatility. A handful of stable sites favor scripts. Many sites or sites that redesign often favor generic or AI extraction. A [r/webscraping thread](https://www.reddit.com/r/webscraping/comments/1ikkr8f/best_way_to_extract_clean_news_articles_around_100/) on pulling clean text from around 100 articles landed on the same split in 2025, noting that Newspaper4k auto-parses most metadata, but for a handful of sites, working out each one's CSS selectors is often just as fast.

For a tool-by-tool comparison of the AI extraction options mentioned here, see

[Best AI Web Scraping Tools for LLM and RAG Pipelines in 2026A by-job ranking of the best AI web scraping tools for 2026, from prompt-based extraction to MCP servers and open-source crawlers for LLM pipelines.](https://scrapfly.io/blog/posts/best-tools-for-ai-webscraping)

Neither approach wins outright. The next section starts with the script side, since it is the fastest way to get one article's fields onto the screen and see exactly what a parser does and does not catch.



## How Do You Scrape a Single News Article in Python?

Fetch the page with `requests`, then parse the fields with Newspaper4k in a handful of lines. Drop to BeautifulSoup only for a field the parser misses.

### Fetch the Page With Requests

A plain GET request with a browser-like `User-Agent` header is the baseline for any site that does not sit behind JavaScript rendering or an anti-bot layer, and it is the fastest way to confirm you can even reach the page before adding a parsing library on top.

### Parse Fields With Newspaper4k

Newspaper4k is the actively maintained fork of Newspaper3k which has not shipped a release since 0.2.8 in September 2018. Newspaper4k shipped 0.9.6 on July 19, 2026, per its [GitHub repo](https://github.com/AndyTheFactory/newspaper4k). The core `Article` API is unchanged from 3k, so existing code ports over with an import swap.



python```python
import json
import requests
from newspaper import Article

url = "https://science.nasa.gov/image-article/apod-2026-august-6-new-sharpest-image-of-the-sun-uncovers-instability/"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"}

response = requests.get(url, headers=headers, timeout=15)

article = Article(url)
article.download(input_html=response.text)
article.parse()

data = {
    "title": article.title,
    "authors": article.authors,
    "publish_date": str(article.publish_date),
    "top_image": article.top_image,
    "text": article.text[:280] + "...",
}

print(json.dumps(data, indent=2, ensure_ascii=False))
```



Running this against a NASA science article returns:



json```json
{
  "title": "New Sharpest Image of the Sun Uncovers Instability",
  "authors": [
    "Jerry Bonnell",
    "Robert Nemiroff",
    "Cecilia Chirenti",
    "Keighley Rockcliffe"
  ],
  "publish_date": "2026-08-06 00:05:00-04:00",
  "top_image": "https://assets.science.nasa.gov/content/dam/science/cds/apod/apod/2026/august/SunFlowers_NSO_1901.jpg/jcr:content/renditions/cq5dam.web.1280.1280.jpeg",
  "text": "Explanation: What does the new sharpest image of our Sun show? Instability. To be clear, a certain kind of interactive process called the Kelvin-Helmholtz instability (KHI). This instability can create waves and swirls when two streams flow past each other -- in this case variabl..."
}
```



Passing the already-fetched HTML into `download(input_html=response.text)` avoids a second network request. Newspaper4k's own downloader works too if you skip the manual fetch, but reusing the response you already have is cheaper and lets you control headers, timeouts, and retries with `requests` directly.

### When to Drop to BeautifulSoup Selectors Instead

Newspaper4k's heuristics are generic and generic heuristics miss fields tied to a specific site's markup. The article above carries a section value in an Open Graph `<meta property="article:section">` tag, but Newspaper4k's parser does not surface `article:` namespaced meta tags as a field, so `article.meta_data` comes back without it.

python```python
from bs4 import BeautifulSoup

soup = BeautifulSoup(response.text, "html.parser")
section_tag = soup.find("meta", attrs={"property": "article:section"})
section = section_tag["content"] if section_tag else None
print(section)  # "APOD"
```



This is the general pattern for the fallback. Find the one attribute or element the generic parser skips, and target it directly rather than replacing the whole extraction with hand-written selectors. For the full mechanics of writing CSS and XPath selectors against arbitrary HTML, see

[How to Parse Web Data with Python and BeautifulsoupBeautifulsoup is one the most popular libraries in web scraping. In this tutorial, we'll take a hand-on overview of how to use it, what is it good for and explore a real -life web scraping example.](https://scrapfly.io/blog/posts/web-scraping-with-python-beautifulsoup)

One article scraped end to end is a good sanity check but it does not tell you which URLs to scrape next.



## How Do You Find Every Article URL on a News Site?

The fast and polite way to enumerate a news site's articles is its RSS feed and XML sitemaps not crawling category pages by hand.

### RSS Feeds and Feedparser

Most news sites, from independent blogs to government science agencies, publish an RSS or Atom feed listing recent articles with a title, link, and publish timestamp. Parsing it with `feedparser` is cheaper than requesting a listing page and cuts through any pagination or infinite-scroll behavior on the site itself.

python```python
import feedparser

feed = feedparser.parse("https://www.nasa.gov/feed/")

for entry in feed.entries[:5]:
    print(entry.title)
    print(entry.link)
    print(entry.published)
    print()
```



That prints five recent entries each with a real article URL ready to hand to the Newspaper4k pipeline above:

```
APOD: 2026 August 6 - New Sharpest Image of the Sun Uncovers Instability
https://science.nasa.gov/image-article/apod-2026-august-6-new-sharpest-image-of-the-sun-uncovers-instability/
Thu, 06 Aug 2026 04:05:00 +0000

NASA's IXPE May Have Proven 90-Year-Old Theory
https://science.nasa.gov/missions/ixpe/nasas-ixpe-may-have-proven-90-year-old-theory/
Wed, 05 Aug 2026 18:49:14 +0000
```



A feed only carries the last few dozen items though. For a full backlog or a site with no feed at all, sitemaps are the next stop, a point one [r/webscraping thread](https://www.reddit.com/r/webscraping/comments/1ikkr8f/best_way_to_extract_clean_news_articles_around_100/) on clean-text extraction raised directly, pointing newcomers to read up on RSS feeds and feedparser before reaching for anything heavier.



### XML Sitemaps for News

Most publishers also expose a dedicated news sitemap, typically linked from `robots.txt`, using the `news:` XML namespace that lists each article's URL, title, and publish date in one place. The Guardian's `sitemaps/news.xml`, for example, listed 477 recent article URLs when tested, each with a `<news:title>` and `<news:publication_date>` alongside the `<loc>`.

Sitemaps cover far more ground than a feed in one request and do not require guessing a pagination scheme. The mechanics of locating and parsing them including compressed sitemap indexes on larger sites are covered in full in

[How to Scrape Sitemaps to Discover Scraping TargetsUsually to find scrape targets we look at site search or category pages but there's a better way - sitemaps! In this tutorial, we'll be taking a look at how to find and scrape sitemaps for target locations.](https://scrapfly.io/blog/posts/how-to-scrape-sitemaps)

When a site has neither a feed nor a sitemap discovery falls back to actual crawling following links from a homepage or section page outward. That is a heavier job, and the many-site scale section later in this guide covers the tooling for it.

Between RSS and sitemaps most sites give up their article inventory without a single page-by-page crawl. The next problem shows up once you fetch each URL. What comes back is the article text tangled up with everything else on the page.



Scrapfly

#### Scale your web scraping effortlessly

Scrapfly handles proxies, browsers, and anti-bot bypass — so you can focus on data.

[Try Free →](https://scrapfly.io/register)## How Do You Extract Clean Article Text Without Ads, Nav, or Boilerplate?

Use a content-extraction library such as Newspaper4k, trafilatura, or readability that isolates the main article body or request LLM-ready Markdown so the text comes back without page chrome.

A raw `soup.get_text()` call pulls every text node on the page (navigation links, cookie banners, related-story teasers, footer boilerplate, and the article itself), concatenated with no boundary between them. On the NASA article used earlier, `get_text()` returned 9,867 characters. Running the same HTML through trafilatura's `extract()` returned 1,224: a one-line site tagline, the headline, and the article body.

python```python
import trafilatura

downloaded = requests.get(url, headers=headers, timeout=15).text
clean_text = trafilatura.extract(downloaded)
print(clean_text[:300])
```



Newspaper4k's `.text` attribute does similar boilerplate removal on its own, so a site already parsed with Newspaper4k rarely needs a second library. Reach for trafilatura or readability when that extraction comes back thin, or when you want Markdown output for an LLM or RAG pipeline instead of plain text.

Automation is not always the right call either. A researcher on r/webscraping needed roughly 100 clean articles for a thesis corpus some behind cookie-consent walls and one behind a paywall. The most useful reply in the thread said to skip the scraper entirely. At only 100 articles copying each one by hand is faster than writing and debugging a pipeline to do it for you.

For a comparison of extraction tools that produce LLM-ready output, check

[Best AI Web Scraping Tools for LLM and RAG Pipelines in 2026A by-job ranking of the best AI web scraping tools for 2026, from prompt-based extraction to MCP servers and open-source crawlers for LLM pipelines.](https://scrapfly.io/blog/posts/best-tools-for-ai-webscraping)

Clean text solves the parsing side of the problem. The next section covers the sites where you never get HTML worth parsing in the first place because the body loads after JavaScript runs or the request never makes it past a bot check.



## How Do You Scrape News Sites That Block Bots or Render With JavaScript?

The big outlets are the hard part. Many render the article body with JavaScript, sit behind Cloudflare or DataDome, or gate content behind a paywall, so a plain `requests` call returns an empty shell or a challenge page instead of an article.

### JavaScript-Rendered News Sites

The symptom is consistent. A 200 response comes back, but the body text is missing, or the page is mostly empty `<div>` containers waiting for a client-side framework to fill them in. Check before reaching for a browser: count the `<p>` tags in the raw response. A plain request to a BBC News article returned the full body in the initial HTML (49 `<p>` tags, and Newspaper4k parsed the text without rendering), so BBC needs no browser. A client-rendered page comes back with the container markup and almost none of the article text.

The fix is rendering the page before parsing it with a local headless browser like Playwright or Selenium, or a managed Cloud Browser that handles rendering for you. Either way you parse the fully rendered DOM not the initial response.

### Anti-Bot Systems on Major Outlets

Plenty of major news outlets run anti-bot protection. Plain requests to the New York Times, Reuters, and Wall Street Journal homepages came back 401 or 403 from DataDome, Bloomberg answered with a PerimeterX captcha, and Politico and the Financial Times returned Cloudflare 403s. A 403 status, a JavaScript challenge page, or content that looks complete but is silently missing the body text are all signs of Cloudflare, DataDome, or a similar system standing between your request and the article.

Watch for these failure modes specifically:

- **403 with no clear cause** - the request looks identical to a browser's but gets rejected anyway, usually on TLS fingerprint or header order
- **Empty JavaScript shell** - a 200 response with none of the article content, covered above
- **Challenge page instead of content** - a "Just a moment" or CAPTCHA interstitial in place of the article
- **Silent partial extraction** - the page loads, but a script or ad-blocking heuristic on the site's end serves a truncated or teaser version of the body to detected bots

A plain `requests` call to a Politico article returns HTTP 403 with Cloudflare's `cf-mitigated: challenge` header and a "Just a moment" page. Scrapfly's [Web Scraping API](https://scrapfly.io/products/web-scraping-api) handles this layer directly, reporting a 98% success rate against Cloudflare, DataDome, Akamai, and other anti-bot vendors with JavaScript rendering and LLM-ready Markdown output built into the same request.



python```python
from scrapfly import ScrapeConfig, ScrapflyClient

scrapfly = ScrapflyClient(key="YOUR_SCRAPFLY_API_KEY")

result = scrapfly.scrape(ScrapeConfig(
    url="https://www.politico.com/news/2026/09/29/talarico-hamilton-christian-democrats-republican-attacks-01096657",
    asp=True,          # solves anti-bot challenges (Cloudflare, DataDome, etc.)
    render_js=True,    # returns the browser-rendered DOM
    country="US",
))

print(result.upstream_status_code)
print(len(result.scrape_result["content"]))  # full rendered HTML length
```



Run against the same Politico article, this returned HTTP 200 with the full rendered article HTML instead of the challenge page. `asp=True` combined with `render_js=True` is the same pattern used against any JavaScript-heavy, anti-bot-protected page, one request in and a fully rendered HTML out ready for the Newspaper4k or BeautifulSoup parsing shown earlier. SDKs exist for Python, TypeScript, Go, and Rust, plus a Scrapy extension, so this is not a Python-only path.

### Paywalls and What You Can and Cannot Do

Hard paywalls where the server never sends the full article body to an unauthenticated request are not something you bypass. Circumventing one is a terms-of-service and often a legal problem not a scraping technique and it is out of scope for this guide.

Metered or soft paywalls behave differently. Some serve the full body but hide it behind a client-side overlay after a view count, a rendering and cookie-handling problem rather than an access-control one. Either way, the rule is simple. If the server itself refuses to send the content without a valid subscription that refusal is the line.

Between JavaScript rendering, anti-bot systems, and paywalls, this section covers the wall that stops most scrapers cold on the outlets people actually want data from. The next question is what changes once you are pointing this whole pipeline at many sites instead of one.



## How Do You Scrape News Articles Across Many Sites Without a Selector per Site?

Stop writing a parser per site. Point a generic article extractor (an auto extraction model or an LLM-prompt extraction) at each URL, and let one pipeline handle layouts that have nothing in common with each other.

The brittleness problem compounds with scale. N sites written as N hand-tuned selectors means N things that can break independently on the next redesign, and nobody notices until a downstream report comes back with blank fields. A [r/webscraping thread](https://www.reddit.com/r/webscraping/comments/1el3ns4/how_to_efficiently_scrape_news_pages_from_1000/) from a developer trying to pull news pages from somewhere between 10 and 2,000 company websites got the same advice from multiple replies, use generic rules and heuristics instead of writing per-site code because per-site code does not survive contact with that many different frontends.

Scrapfly's [AI Extraction API](https://scrapfly.io/products/extraction-api) applies this idea directly to article pages. Its `article` auto-extraction model returns `headline`, `author` and `authors_list`, `date_published`, `article_body`, `main_image`, `canonical_url`, `language`, and `related_articles` (full list in the [article model reference](https://scrapfly.io/docs/extraction-api/automatic-ai/models/article)) from an arbitrary article URL with no site-specific configuration, and `extraction_prompt` lets you specify custom fields in plain language when the built-in model does not cover what you need.



python```python
from scrapfly import ScrapflyClient, ExtractionConfig

scrapfly = ScrapflyClient(key="YOUR_SCRAPFLY_API_KEY")

article_url = "https://www.theguardian.com/technology/2026/aug/06/meta-ai-smart-glasses-privacy"
html = requests.get(article_url, headers=headers, timeout=15).text

extraction_result = scrapfly.extract(ExtractionConfig(
    body=html,
    content_type="text/html",
    url=article_url,
    extraction_model="article",  # built-in auto model for article pages
))

print(extraction_result.data)  # structured article fields, no selectors written
```



The same call works unchanged against a BBC article, a Guardian article, or a small independent outlet's WordPress page, since the model reads the page content rather than a fixed DOM path. Re-extracting from cached HTML instead of re-fetching keeps iteration cheap when tuning a schema across a large corpus.

For discovery at that scale, the [Crawler API](https://scrapfly.io/products/crawler-api) handles recursive site-wide crawling, with `use_sitemaps=True` to prefer sitemaps and `include_only_paths` or `exclude_paths` to keep the crawl inside article sections instead of login pages or archives. Feed its output into the extraction call above and discovery plus extraction cover the whole pipeline from a bare domain to structured records.

If the extracted articles are feeding an LLM or agent pipeline downstream, this guide covers the hand-off in more detail:

[Web Scraping for AI Agents in 2026How AI agents consume the web, why their fetch layer breaks, and how to build agent-grade web access that holds up in production.](https://scrapfly.io/blog/posts/ai-agent-web-scraping)



## FAQ

Can you scrape news websites?Yes, for publicly accessible pages, within the limits of robots.txt, a site's terms of service, and copyright on the article text itself. Facts are not copyrightable, but the publisher's sentences are, so analyzing or summarizing coverage is fine while republishing full article text generally is not. Paywalled content that requires a subscription to view is off-limits regardless of technique.







Should I use Newspaper3k or Newspaper4k?Use Newspaper4k. It is the actively maintained fork, with releases and commits continuing through 2026, while Newspaper3k has not shipped a release since 0.2.8 in September 2018. New projects should not start on the unmaintained original.







What is the difference between scraping news and using a news API?Scraping pulls article content directly from a publisher's own pages, working against any site with full control over what you extract. A news API or aggregator, like Google News, returns items from a provider's index instead, which is easier to query but limited to whatever that provider covers. The [Google News API guide](https://scrapfly.io/blog/posts/guide-to-google-news-api-and-alternatives) covers the aggregator route in full.







How do I scrape news articles for free?`requests`, Newspaper4k, and `feedparser` cost nothing to run against a handful of stable, unprotected sites. Costs show up once you need JavaScript rendering, anti-bot bypass, or extraction across enough sites that a managed API becomes cheaper than the engineering time to maintain it yourself.







How do I get clean article text without ads and menus?Use a content-extraction library, Newspaper4k's own `.text` field, trafilatura, or readability, or request LLM-ready Markdown output so the body comes back without navigation, ads, or related-story blocks mixed in.









## Summary

Start simple. Use RSS or sitemaps for discovery, then `requests` plus Newspaper4k for extraction on one site at a time. Escalate deliberately from there, generic or AI extraction once you are past a handful of sites, and a managed anti-bot layer plus JavaScript rendering once you hit a protected outlet.

Both walls, many-site extraction and anti-bot-protected outlets, are exactly where Scrapfly's [Web Scraping API](https://scrapfly.io/products/web-scraping-api) and [AI Extraction API](https://scrapfly.io/products/extraction-api) earn their place, handling rendering, anti-bot bypass, and schema-free extraction behind a single request instead of a growing pile of per-site scripts.



### Web Scraping API

Scrape any website with our powerful API. Anti-bot bypass, JavaScript rendering, and rotating proxies built-in.



[Try Web Scraping API](https://scrapfly.io/docs/scrape-api/getting-started)



Legal Disclaimer and PrecautionsThis tutorial covers popular web scraping techniques for education. Interacting with public servers requires diligence and respect:

- Do not scrape at rates that could damage the website.
- Do not scrape data that's not available publicly.
- Do not store PII of EU citizens protected by GDPR.
- Do not repurpose *entire* public datasets which can be illegal in some countries.

Scrapfly does not offer legal advice but these are good general rules to follow. For more you should consult a lawyer.

 

   [  Add as a preferred source ](https://google.com/preferences/source?q=scrapfly.io) Table of Contents















 

  Table of Contents- [Key Takeaways](#key-takeaways)
- [What Data Can You Extract From a News Article?](#what-data-can-you-extract-from-a-news-article)
- [Should You Use a Scraping Script or AI Extraction for News?](#should-you-use-a-scraping-script-or-ai-extraction-for-news)
- [How Do You Scrape a Single News Article in Python?](#how-do-you-scrape-a-single-news-article-in-python)
- [Fetch the Page With Requests](#fetch-the-page-with-requests)
- [Parse Fields With Newspaper4k](#parse-fields-with-newspaper4k)
- [When to Drop to BeautifulSoup Selectors Instead](#when-to-drop-to-beautifulsoup-selectors-instead)
- [How Do You Find Every Article URL on a News Site?](#how-do-you-find-every-article-url-on-a-news-site)
- [RSS Feeds and Feedparser](#rss-feeds-and-feedparser)
- [XML Sitemaps for News](#xml-sitemaps-for-news)
- [How Do You Extract Clean Article Text Without Ads, Nav, or Boilerplate?](#how-do-you-extract-clean-article-text-without-ads-nav-or-boilerplate)
- [How Do You Scrape News Sites That Block Bots or Render With JavaScript?](#how-do-you-scrape-news-sites-that-block-bots-or-render-with-javascript)
- [JavaScript-Rendered News Sites](#javascript-rendered-news-sites)
- [Anti-Bot Systems on Major Outlets](#anti-bot-systems-on-major-outlets)
- [Paywalls and What You Can and Cannot Do](#paywalls-and-what-you-can-and-cannot-do)
- [How Do You Scrape News Articles Across Many Sites Without a Selector per Site?](#how-do-you-scrape-news-articles-across-many-sites-without-a-selector-per-site)
- [FAQ](#faq)
- [Summary](#summary)
 
    Join the Newsletter  Get monthly web scraping insights 

 

  



Scale Your Web Scraping

Anti-bot bypass, browser rendering, and rotating proxies, all in one API. Start with 1,000 free credits.

  No credit card required  1,000 free API credits  Anti-bot bypass included 

 [Start Free](https://scrapfly.io/register) [View Docs](https://scrapfly.io/docs/onboarding) 

 Not ready? Get our newsletter instead. 

 

 ## Related Articles

 [  

 python data-parsing 

### How to Parse Datetime Strings with Python and Dateparser

Dateparser is a popular Python package for parsing datetime strings. Here's how it can be used in web scraping and how t...

 

 ](https://scrapfly.io/blog/posts/parsing-datetime-strings-with-python-and-dateparser) [     

 python blocking 

### How to Bypass Anti-Bot Protection in 2026: All 8 Major Vendors

Identify and bypass Cloudflare, DataDome, PerimeterX, Kasada, Akamai, Incapsula, F5, and AWS WAF with Python code exampl...

 

 ](https://scrapfly.io/blog/posts/how-to-bypass-anti-bot-protection) [     

 python scrapeguide 

### How to Scrape Capterra Reviews and Software Data

Learn how to scrape Capterra software reviews and listings using Python and Scrapfly, bypassing Cloudflare Bot Managemen...

 

 ](https://scrapfly.io/blog/posts/how-to-scrape-capterra) 

  



   



 Scale your web scraping effortlessly, **1,000 free credits** [Start Free](https://scrapfly.io/register)