The short version
Collecting publicly visible business information is usually fine. Collecting anything behind a login, ignoring a site’s stated refusal, or harvesting personal data without a lawful basis is not. The law involved varies by country, and this is general information rather than legal advice — but the practical tests below are the same ones we apply before quoting.
Four tests before you scrape
- Is the data behind a login or paywall? If you cannot see it without credentials, do not scrape it.
- Does robots.txt disallow it? Treat a disallow as a no. It is not a law, but crawling against it is the clearest evidence of disregard you can hand a lawyer.
- Does the site’s terms forbid it? This is where most briefs go wrong. Terms of service are a contract you accept by using the site, and “scraped anyway” is a bad position to be in.
- Is it personal data? Names, emails, phone numbers attached to a person are personal data almost everywhere now. You need a lawful basis, and for marketing you usually need consent or a legitimate-interest case you can actually document.
What it costs
- $200–$500 one-off: one source site, defined fields, delivered as CSV or Sheet.
- $49–$99/month scheduled feed: the same scrape on a cadence, delivered where you need it, with a note when the page structure changes and the data goes wrong.
- $300–$400 PDF or invoice data extraction where the source is not structured at all.
These match our Why a scraper breaks, and why that matters
Every scraper breaks eventually, because someone changes the page. The honest question
is who finds out. A silent scraper feeding a sheet is worse than no scraper, because your
decisions are made on data that quietly stopped arriving. Any feed we build says so when
the structure changes, rather than sending you empty rows that look like “no business
this week”. Tell us the source and the fields you need. We will tell you honestly whether it should
be scraped at all, then quote a fixed price:
Which dataset do you currently wish you could just query?Better options that are not scraping
Next step
