EMZETT.
Login

Scraper

In short: A program that automatically reads content from websites (e.g. fetches HTML pages and extracts the relevant data), instead of a human copying it manually.

In more detail: Web scraping moves in a legal and ethical grey area: many websites explicitly forbid it in their terms of use or block it technically (e.g. via robots.txt, rate limiting or bot detection), since aggressive scraping can act like a small DDoS attack. Legitimate uses include, for example, price-comparison portals or search engine indexing; it becomes problematic when copying copyrighted content or bypassing access protection.

In Depth

Technically, a simple scraper usually consists of two steps: an HTTP request to the target page (often with a “real” user-agent header, to avoid being immediately detected as a bot) and a parser that filters the desired values out of the returned HTML:

import requests
from bs4 import BeautifulSoup
 
response = requests.get("https://example-shop.com/product/123")
soup = BeautifulSoup(response.text, "html.parser")
price = soup.select_one(".price").text

Websites that want to prevent scraping use several measures at once: robots.txt as a purely voluntary courtesy agreement (a scraper can technically ignore it, but well-behaved bots like Google respect it), rate limiting per IP, CAPTCHAs when a suspicious pattern appears, and increasingly client-side JavaScript rendering, which leaves a simple HTTP scraper without a real browser (e.g. headless Chrome via Playwright/Puppeteer) empty-handed.

The legal assessment depends heavily on context: scraping publicly accessible, non-copyrighted facts (e.g. prices) is generally allowed in many jurisdictions, as long as no access barriers are bypassed and the terms of use aren’t contractually and bindingly rejected — but once copyrighted content (text, images) is copied at scale, or technical protection measures are actively circumvented, it becomes legally risky.

Scraper vs. crawler vs. API

Three related terms are often mixed up: a crawler (as used by search engines) systematically follows links to discover and index as many pages as possible, without necessarily being interested in specific individual data points. A scraper is more targeted — it calls known pages and extracts specific data points (prices, product names, reviews). An official API is the legitimate, and usually more stable, way provided by the target site itself to get the same data in a structured form without having to parse HTML — many companies offer APIs precisely FOR this reason, to reduce uncontrolled scraping of their website.

Countermeasures from an operator’s perspective

From the perspective of a website operator wanting to curb unwanted scraping, there are several levels of countermeasures: at the network level, rate limiting or a web application firewall can block suspicious request patterns (many requests in a short time from the same IP). At the application level, CAPTCHAs or behavioural analysis help (real users move the mouse, scroll irregularly — a simple scraper doesn’t). Advanced bot-detection services (e.g. Cloudflare Bot Management) additionally analyse TLS fingerprints and browser properties to distinguish even well-disguised automated requests from real users — a constant arms race, since scraper developers in turn try to appear like a real browser (e.g. by using real headless browsers instead of pure HTTP libraries).

Even where scraping is legally permitted, ethical questions remain: a scraper that queries a website at high frequency can cause real infrastructure costs for the operator (server load, bandwidth) — a “good” scraper therefore throttles itself (e.g. a one-second pause between requests) and respects robots.txt even if it isn’t technically forced to. With personal data (e.g. systematically scraping public social media profiles), the GDPR additionally comes into play: even publicly viewable data remains personal data, and its automated mass processing needs its own legal basis — one reason why larger scraping projects in Europe should be carefully reviewed legally before they’re implemented.

See also: DDoS