Scraping has a reputation problem. On the one hand you have people running 50-thread scrapers on a competitor’s site to undercut them by a penny, on the other you have legitimate owners who’d pay a fair price for public pricing data — if they knew the rules. The legal reality is more boring and more workable than the headlines suggest.

What scraping is actually good for

Three use cases consistently pay for themselves at small-business scale:

  • Price monitoring. Watching 20-50 SKUs across two or three competitors, daily, and re-pricing a handful of items a week. At typical retail margins this pays back days, not months. Rates can double or halve based on your position — your margin is what you do with the data, not the scraping itself.
  • Competitor catalogs. A weekly snapshot of what products a competitor carries, their stockouts, and their promo structure. This is research your sales team already does by hand; scraping just makes it comprehensive.
  • Lead generation. Pulling publicly listed business directories (chambers of commerce, government registries, industry listing sites) and enriching them into a CRM. This one has the most legal surface area of the three — see below.

What doesn’t pay: scraping social media feeds, private login-walled data, or anything behind a CAPTCHA. If a site fights you, the effort is wasted — the data is rarely worth the fight.

Three layers matter, and people mix them up:

  1. Terms of Service (ToS). A contract between you and the site. Violating it can get your account banned and can support a breach-of-contract claim. Scraping isn’t always forbidden by ToS, but most commercial sites’ ToS have an anti-scraping clause these days.
  2. The law. In the US, hiQ Labs v. LinkedIn (2022) was a big deal: the Ninth Circuit held that scraping public data generally doesn’t violate the Computer Fraud and Abuse Act (CFAA) — the “authorization” question is about the website’s technical access controls, not its ToS. But that’s a specific ruling in one circuit, not a blank check. The EU’s Database Directive and sui generis database rights can bite if you extract substantial parts of a database.
  3. Personal data. This is where small businesses actually get in trouble. GDPR (Europe) and the PDP Law (Indonesia’s UU 27/2022) apply when you collect personal data — names, emails, phone numbers — of identifiable individuals. Lead-gen scraping that collects personal data without a legal basis is a compliance problem, full stop. Fines under the Indonesian PDP Law can reach 2% of annual revenue.

The practical summary: public business data, lightly scraped at low frequency, is a gray-to-acceptable zone in most jurisdictions. Personal data at scale is where you cross the line.

The robots.txt rule: respect it. It’s not law, but it’s the site owner’s stated preference, and courts and scraping services treat it as the baseline of “reasonable behavior.” If a site blocks you in robots.txt or via rate limiting, stop or back off — don’t rotate proxies to get around it. Circumvention is what turns a gray area into a lawsuit.

A minimal, polite scraper in Python

Here’s a scraper that gets the job done without hammering anyone. requests fetches, BeautifulSoup parses — that’s all you need for 95% of small-business targets.

import requests
from bs4 import BeautifulSoup
import time
import csv

headers = {
    "User-Agent": "MyCompanyPriceChecker/1.0 (+https://mycompany.example)",
}

def fetch_prices(url):
    resp = requests.get(url, headers=headers, timeout=10)
    resp.raise_for_status()
    soup = BeautifulSoup(resp.text, "html.parser")
    results = []
    for item in soup.select(".product"):
        name = item.select_one(".name").text.strip()
        price = item.select_one(".price").text.strip()
        results.append((name, price))
    return results

with open("prices.csv", "w", newline="") as f:
    writer = csv.writer(f)
    writer.writerow(["name", "price"])
    for url in ["https://competitor.example/products/a",
                "https://competitor.example/products/b"]:
        writer.writerows(fetch_prices(url))
        time.sleep(2)   # 2 seconds between requests — be a guest, not a burglar

Three things matter here beyond the code: a real User-Agent identifying you, a 2-second delay between requests, and checking robots.txt before you start (usually https://competitor.example/robots.txt). That last one takes ten seconds and tells you what paths are off-limits.

Three rules to stay out of trouble

  1. Identify yourself. A real User-Agent with your domain. Scrapers that hide are the ones that get sued.
  2. Go slow. A few requests per minute, one run a day. If the site has an API, use it. If the price is ~$50-100/month for a data API (many competitor-data services charge this), compare that against your server costs — sometimes paying is cheaper than maintaining a scraper that breaks every time the site redesigns.
  3. No personal data. If the data identifies an individual, you’re in GDPR/PDP Law territory and need a legal basis. Scrape product prices and business listings; leave consumer contact data alone.

Scraping is a tool, like a drill. Used on public data, politely, at low volume, it’s a legitimate business edge. Used to hoover personal data or to defeat a site’s protections, it’s a liability. Stay in the boring lane and it keeps working.