EMZETT.
Login

Web Scraping

In short: Automatically extracting data from websites with a script — either from raw HTML or via an official API, if one exists.

In more detail: For simple pages without their own API, HTTP requests + an HTML parser are enough (Python: requests + BeautifulSoup; Node.js: fetch/axios + cheerio). If the page needs JavaScript to render (many modern admin interfaces), that’s no longer enough — then you need a real (headless) browser such as Playwright or Puppeteer. Login-protected pages: fetch the session cookie/token once via the login form, then reuse it. If an official API exists, it’s almost always the better choice (more stable, no terms-of-service risk) — scraping is usually just the fallback.

Our context: Came up as a short digression — a general question of how you could automatically query your own devices/services on the home network (router, NAS, smart home) or your own accounts on platforms, independent of the Emzett project.

In Depth

Two technical hurdles usually decide how much work a scraping project becomes. First: static vs. dynamic HTML. If the server delivers the complete HTML with all content directly (View Source shows the text you also see in the browser), a simple HTTP request is enough. If the page only assembles its content via JavaScript in the browser (View Source only shows an empty skeleton), the scraper has to actually run a real browser to get at the finished content — considerably slower and more resource-intensive.

Second: anti-scraping measures. Many sites recognise automated access by patterns such as missing/atypical browser headers, overly regular request intervals, or missing mouse behaviour, and then block it (often with CAPTCHAs or by silently serving false data). Legally, scraping is in a grey area: reading publicly accessible data is usually allowed, but terms of service can explicitly forbid it, and aggressive scraping (many requests in a short time) can be regarded as an attack on server capacity — another reason to prefer an official API if one exists.

import requests
from bs4 import BeautifulSoup
 
response = requests.get(url, headers={"User-Agent": "..."})
soup = BeautifulSoup(response.text, "html.parser")

robots.txt and etiquette

Most websites publish machine-readable rules under /robots.txt stating which areas automated programs (crawlers/scrapers) may visit and which not — not directly legally binding, but a clear signal from the site operator that reputable scrapers should respect. Fair scraping also includes: artificially delaying requests (not firing dozens of parallel requests at once), setting an honest User-Agent header instead of posing as a normal browser, and caching when data is needed repeatedly instead of fetching the same page over and over.

When a real browser is worth it

Playwright/Puppeteer remote-control a real (usually invisible, “headless”) browser — this makes them considerably slower and more resource-hungry than a plain HTTP request, but in return they can handle JavaScript-heavy pages, logins via real forms, and interactions such as scrolling or clicking, which a plain HTML parser can’t reproduce at all:

from playwright.sync_api import sync_playwright
 
with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto(url)
    page.wait_for_selector(".content")  # wait until JS has rendered the content
    html = page.content()

See also: HTML, API, Browser