No law we read bans web scraping as such. What makes a given scrape unlawful is one of four things: getting past a login or a block, breaking terms you accepted, collecting personal data without a lawful basis, or copying protected content.
This page gives no legal advice. It quotes the courts, the laws and the regulators. And on 11 October 2026 we read what 20 well-known sites say about scraping in their terms, their robots.txt files and their API pages.
Four questions that decide it
Each question has its own law, and a scrape can pass one and fail another. In hiQ, the court of appeals sided with the scraper on the first question, and the case ended on the second.
1. Is it behind a login, or were you told to stop?
In the US, the federal anti-hacking law is the Computer Fraud and Abuse Act. It punishes whoever "intentionally accesses a computer without authorization or exceeds authorized access". In Van Buren, in 2021, the Supreme Court read that as "a gates-up-or-down inquiry": either you may enter a system or an area of it, or you may not.
In 2022 the Ninth Circuit applied it to scraping in hiQ v. LinkedIn, ruling on a preliminary injunction: "the concept of “without authorization” does not apply to public websites." A login is a gate. The court added in a footnote: "Accessing those areas “without authorization” would violate the CFAA."
Being told to stop matters too. The court described an earlier case, Power Ventures: "After Facebook sent a cease-and-desist letter, Power Ventures continued to circumvent IP barriers and gain access to password-protected Facebook member profiles." That was held to be access without authorization.
Sites do build those barriers. The same opinion records that LinkedIn blocked about 95 million automated attempts to scrape it every day.
2. Did you accept the site's terms?
Terms you accept are a contract, and the CFAA ruling did not touch them. The Ninth Circuit itself listed what stays open: "copyright infringement, misappropriation, unjust enrichment, conversion, breach of contract, or breach of privacy".
hiQ shows how that ends. On 8 December 2022 the district court entered a consent judgment: "Judgment in the amount of $500,000 USD is hereby entered against hiQ and in favor of LinkedIn." hiQ was barred from using LinkedIn in breach of its User Agreement, "whether directly or indirectly through a third party, intermediary, or proxy". It also had to destroy the data. Our page on proxies for LinkedIn has the full timeline.
Terms reach only what they cover, though. In January 2024 a court read Meta's terms of the time and found that they "do not bar logged-off scraping of public data". Facebook's current terms go further. They forbid automated collection "regardless of whether such automated access or collection is undertaken while logged-in to a Facebook account".
3. Is there personal data in it?
Then privacy law applies, wherever the page is. Under the GDPR, "processing" means any operation on personal data, "such as collection". Collecting it needs a lawful basis.
Consent is one of six bases. Another is "the legitimate interests pursued by the controller or by a third party", unless the rights of the people in the data weigh more. So a scraper does not always need consent. It does need a basis.
Outside the EU the message is the same. In 2023, 12 privacy regulators signed a joint statement, among them the UK's ICO and the authorities of Canada and Australia: "Personal information that is publicly accessible is still subject to data protection and privacy laws in most jurisdictions."
4. Are you copying content or a database?
Copyright covers what you copy, not how you reach it. In the EU, text and data mining has its own exception. It applies only if the use "has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online".
The directive's recital names what counts as machine-readable: "metadata and terms and conditions of a website or a service". So in the EU, a no-scraping line in the terms can also be an opt-out from text and data mining.
What robots.txt is, and is not
robots.txt is a request. The standard for it, RFC 9309, says so plainly: "These rules are not a form of access authorization."
Terms can change that. Six of the 20 sites we read write robots.txt into their terms. Reddit's allow crawling only "in accordance with the parameters set forth in our robots.txt file", and add that "scraping the Services without Reddit’s prior written consent is prohibited". Google forbids automated access only "in violation of the machine-readable instructions on our web pages".
What 20 well-known sites say
We picked 20 well-known sites, across shops, search, video, jobs, property, travel, reviews, social media and reference. For each, we parsed robots.txt the way RFC 9309 groups it, read the terms in a browser on our server, and opened the developer pages:
| Site | Terms on scraping | robots.txt, unnamed bots | robots.txt, GPTBot | API |
|---|---|---|---|---|
| Amazon | Forbids | Paths only | Shut out | Yes |
| eBay | Forbids | Paths only | Shut out | Yes |
| Walmart | Forbids | Paths only | Paths only | Yes |
| Etsy | Forbids | Paths only | Paths only | Yes |
| Google Search | Only where robots.txt says no | Paths only | Paths only | Yes |
| YouTube | Forbids, except search engines | Paths only | Paths only | Yes |
| Forbids | Shut out | Shut out | Yes | |
| Indeed | Forbids | Paths only | Paths only | Yes |
| Zillow | Forbids | Paths only | Paths only | Yes |
| Booking.com | Forbids | Paths only | Paths only | Yes |
| Airbnb | Forbids | Paths only | Shut out | None found |
| Tripadvisor | Forbids | Paths only | Shut out | Refused us |
| Yelp | Forbids | Refused us | Refused us | Refused us |
| Forbids | Shut out | Shut out | Yes | |
| X | Forbids | Shut out | Shut out | Yes |
| Forbids | Shut out | Paths only | Yes | |
| Wikipedia | Only abusive or unapproved use | Paths only | Paths only | Yes |
| GitHub | Allows research and archiving | Paths only | Paths only | Yes |
| Stack Overflow | Forbids | Paths only | Paths only | Yes |
| IMDb | Forbids | Shut out | Shut out | Yes |
Three things stand out. First, the terms are stricter than robots.txt. Of the 17 sites whose terms forbid scraping, 11 let unnamed bots in through robots.txt and close only listed paths. A robots.txt that lets you in is not permission.
Second, three sites are more open in their terms. Google forbids only what robots.txt forbids. Wikipedia rules out automated use that is "abusive or disruptive" or that has not been approved by its community. GitHub lets researchers use public, non-personal information if their publications are open access, and archivists use public information.
Third, AI crawlers get their own rules. Eight files shut GPTBot out completely. Facebook does the opposite of what you might expect: it shuts out every unnamed bot, but gives GPTBot its own group with only some paths closed.
Where a proxy fits
A proxy changes the address a site sees. It does not change any of the four answers above. The hiQ order banned access "directly or indirectly through a third party, intermediary, or proxy". And the access the court held unauthorized in Power Ventures continued around Facebook's IP barriers, after a cease-and-desist letter.
Where a scrape is allowed, a proxy is ordinary infrastructure. An API, or a site you have permission to read, may limit requests per address. Our guide to proxies for web scraping covers that use, and our page on avoiding IP bans covers polite request rates.
What this page could not check
- This is not legal advice, and it cannot tell you whether your project is lawful.
- It covers US federal law, the EU's GDPR and copyright directive, and one statement of 12 regulators. It does not cover US state laws, national copyright acts, or countries such as India or China.
- The 20 sites are our pick, not a sample of the web.
- We searched the terms for words about scraping and read every match. A clause in a separate document could still be missed.
- Yelp's robots.txt and two developer pages refused our server.
- Terms and robots.txt files change often. We will read all 20 again by 11 January 2027.
Sources
- Van Buren v. United States, Supreme Court, decided 3 June 2021: supremecourt.gov.
- hiQ Labs v. LinkedIn, Ninth Circuit, filed 18 April 2022: ca9.uscourts.gov.
- hiQ Labs v. LinkedIn, consent judgment and permanent injunction, N.D. Cal., 8 December 2022: courtlistener.com.
- Meta Platforms v. Bright Data, order of 23 January 2024, N.D. Cal.: courtlistener.com.
- 18 U.S.C. 1030, the Computer Fraud and Abuse Act: law.cornell.edu.
- RFC 9309, Robots Exclusion Protocol, September 2022: rfc-editor.org.
- Regulation (EU) 2016/679 (GDPR), Articles 4 and 6: eur-lex.europa.eu.
- Directive (EU) 2019/790, Article 4 and recital 18: eur-lex.europa.eu.
- Joint statement on data scraping and the protection of privacy, 24 August 2023: priv.gc.ca.
- The terms of the 20 sites, read 11 October 2026; Walmart's, Etsy's and Zillow's from Internet Archive copies of 5 October, 7 October and 1 February 2026.
- Our own tests on our server, 11 October 2026: robots.txt at 19:40 UTC, terms at 19:45 to 19:47 UTC, developer pages at 20:03 UTC.



