Skip to content

Is web scraping legal? What courts, laws and site terms actually say

Scraping public pages is not hacking under US law, but terms, personal data and copyright still apply. We read 20 big sites' terms: 17 forbid scraping.

HProxy TeamOctober 11, 2026Updated October 11, 20268 min read

Is web scraping legal? What courts, laws and site terms actually say

No law we read bans web scraping as such. What makes a given scrape unlawful is one of four things: getting past a login or a block, breaking terms you accepted, collecting personal data without a lawful basis, or copying protected content.

This page gives no legal advice. It quotes the courts, the laws and the regulators. And on 11 October 2026 we read what 20 well-known sites say about scraping in their terms, their robots.txt files and their API pages.

Four questions that decide it

Diagram: four questions that decide whether a scrape is lawful. One, is it behind a login or were you told to stop, under anti-hacking law. Two, did you accept the terms, under contract law. Three, is there personal data, under the GDPR. Four, are you copying content or a database, under copyright. robots.txt alone is not access authorization, yet 6 of 20 sites put it in their terms
Diagram: HProxy, from the court records, laws and regulators' statement named below, read 11 October 2026

Each question has its own law, and a scrape can pass one and fail another. In hiQ, the court of appeals sided with the scraper on the first question, and the case ended on the second.

1. Is it behind a login, or were you told to stop?

In the US, the federal anti-hacking law is the Computer Fraud and Abuse Act. It punishes whoever "intentionally accesses a computer without authorization or exceeds authorized access". In Van Buren, in 2021, the Supreme Court read that as "a gates-up-or-down inquiry": either you may enter a system or an area of it, or you may not.

In 2022 the Ninth Circuit applied it to scraping in hiQ v. LinkedIn, ruling on a preliminary injunction: "the concept of “without authorization” does not apply to public websites." A login is a gate. The court added in a footnote: "Accessing those areas “without authorization” would violate the CFAA."

Being told to stop matters too. The court described an earlier case, Power Ventures: "After Facebook sent a cease-and-desist letter, Power Ventures continued to circumvent IP barriers and gain access to password-protected Facebook member profiles." That was held to be access without authorization.

Sites do build those barriers. The same opinion records that LinkedIn blocked about 95 million automated attempts to scrape it every day.

2. Did you accept the site's terms?

Terms you accept are a contract, and the CFAA ruling did not touch them. The Ninth Circuit itself listed what stays open: "copyright infringement, misappropriation, unjust enrichment, conversion, breach of contract, or breach of privacy".

hiQ shows how that ends. On 8 December 2022 the district court entered a consent judgment: "Judgment in the amount of $500,000 USD is hereby entered against hiQ and in favor of LinkedIn." hiQ was barred from using LinkedIn in breach of its User Agreement, "whether directly or indirectly through a third party, intermediary, or proxy". It also had to destroy the data. Our page on proxies for LinkedIn has the full timeline.

Terms reach only what they cover, though. In January 2024 a court read Meta's terms of the time and found that they "do not bar logged-off scraping of public data". Facebook's current terms go further. They forbid automated collection "regardless of whether such automated access or collection is undertaken while logged-in to a Facebook account".

3. Is there personal data in it?

Then privacy law applies, wherever the page is. Under the GDPR, "processing" means any operation on personal data, "such as collection". Collecting it needs a lawful basis.

Consent is one of six bases. Another is "the legitimate interests pursued by the controller or by a third party", unless the rights of the people in the data weigh more. So a scraper does not always need consent. It does need a basis.

Outside the EU the message is the same. In 2023, 12 privacy regulators signed a joint statement, among them the UK's ICO and the authorities of Canada and Australia: "Personal information that is publicly accessible is still subject to data protection and privacy laws in most jurisdictions."

4. Are you copying content or a database?

Copyright covers what you copy, not how you reach it. In the EU, text and data mining has its own exception. It applies only if the use "has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online".

The directive's recital names what counts as machine-readable: "metadata and terms and conditions of a website or a service". So in the EU, a no-scraping line in the terms can also be an opt-out from text and data mining.

What robots.txt is, and is not

robots.txt is a request. The standard for it, RFC 9309, says so plainly: "These rules are not a form of access authorization."

Terms can change that. Six of the 20 sites we read write robots.txt into their terms. Reddit's allow crawling only "in accordance with the parameters set forth in our robots.txt file", and add that "scraping the Services without Reddit’s prior written consent is prohibited". Google forbids automated access only "in violation of the machine-readable instructions on our web pages".

What 20 well-known sites say

We picked 20 well-known sites, across shops, search, video, jobs, property, travel, reviews, social media and reference. For each, we parsed robots.txt the way RFC 9309 groups it, read the terms in a browser on our server, and opened the developer pages:

What 20 well-known sites say about scraping
Terms forbid it without permissionof 20
17
Publish an official API we could openof 20
17
robots.txt shuts out GPTBotof 19 we could read
8
Terms write robots.txt into the dealof 20
6
robots.txt shuts out every unnamed botof 19 we could read
5
Source: HProxy reading of 20 sites on 11 October 2026, on our server. Walmart's, Etsy's and Zillow's terms and Stack Overflow's robots.txt came from Internet Archive copies; Yelp's robots.txt refused our server. · hproxy.comProxy.
SiteTerms on scrapingrobots.txt, unnamed botsrobots.txt, GPTBotAPI
AmazonForbidsPaths onlyShut outYes
eBayForbidsPaths onlyShut outYes
WalmartForbidsPaths onlyPaths onlyYes
EtsyForbidsPaths onlyPaths onlyYes
Google SearchOnly where robots.txt says noPaths onlyPaths onlyYes
YouTubeForbids, except search enginesPaths onlyPaths onlyYes
LinkedInForbidsShut outShut outYes
IndeedForbidsPaths onlyPaths onlyYes
ZillowForbidsPaths onlyPaths onlyYes
Booking.comForbidsPaths onlyPaths onlyYes
AirbnbForbidsPaths onlyShut outNone found
TripadvisorForbidsPaths onlyShut outRefused us
YelpForbidsRefused usRefused usRefused us
RedditForbidsShut outShut outYes
XForbidsShut outShut outYes
FacebookForbidsShut outPaths onlyYes
WikipediaOnly abusive or unapproved usePaths onlyPaths onlyYes
GitHubAllows research and archivingPaths onlyPaths onlyYes
Stack OverflowForbidsPaths onlyPaths onlyYes
IMDbForbidsShut outShut outYes

Three things stand out. First, the terms are stricter than robots.txt. Of the 17 sites whose terms forbid scraping, 11 let unnamed bots in through robots.txt and close only listed paths. A robots.txt that lets you in is not permission.

Second, three sites are more open in their terms. Google forbids only what robots.txt forbids. Wikipedia rules out automated use that is "abusive or disruptive" or that has not been approved by its community. GitHub lets researchers use public, non-personal information if their publications are open access, and archivists use public information.

Third, AI crawlers get their own rules. Eight files shut GPTBot out completely. Facebook does the opposite of what you might expect: it shuts out every unnamed bot, but gives GPTBot its own group with only some paths closed.

Where a proxy fits

A proxy changes the address a site sees. It does not change any of the four answers above. The hiQ order banned access "directly or indirectly through a third party, intermediary, or proxy". And the access the court held unauthorized in Power Ventures continued around Facebook's IP barriers, after a cease-and-desist letter.

Where a scrape is allowed, a proxy is ordinary infrastructure. An API, or a site you have permission to read, may limit requests per address. Our guide to proxies for web scraping covers that use, and our page on avoiding IP bans covers polite request rates.

What this page could not check

  • This is not legal advice, and it cannot tell you whether your project is lawful.
  • It covers US federal law, the EU's GDPR and copyright directive, and one statement of 12 regulators. It does not cover US state laws, national copyright acts, or countries such as India or China.
  • The 20 sites are our pick, not a sample of the web.
  • We searched the terms for words about scraping and read every match. A clause in a separate document could still be missed.
  • Yelp's robots.txt and two developer pages refused our server.
  • Terms and robots.txt files change often. We will read all 20 again by 11 January 2027.

Sources

  • Van Buren v. United States, Supreme Court, decided 3 June 2021: supremecourt.gov.
  • hiQ Labs v. LinkedIn, Ninth Circuit, filed 18 April 2022: ca9.uscourts.gov.
  • hiQ Labs v. LinkedIn, consent judgment and permanent injunction, N.D. Cal., 8 December 2022: courtlistener.com.
  • Meta Platforms v. Bright Data, order of 23 January 2024, N.D. Cal.: courtlistener.com.
  • 18 U.S.C. 1030, the Computer Fraud and Abuse Act: law.cornell.edu.
  • RFC 9309, Robots Exclusion Protocol, September 2022: rfc-editor.org.
  • Regulation (EU) 2016/679 (GDPR), Articles 4 and 6: eur-lex.europa.eu.
  • Directive (EU) 2019/790, Article 4 and recital 18: eur-lex.europa.eu.
  • Joint statement on data scraping and the protection of privacy, 24 August 2023: priv.gc.ca.
  • The terms of the 20 sites, read 11 October 2026; Walmart's, Etsy's and Zillow's from Internet Archive copies of 5 October, 7 October and 1 February 2026.
  • Our own tests on our server, 11 October 2026: robots.txt at 19:40 UTC, terms at 19:45 to 19:47 UTC, developer pages at 20:03 UTC.

Frequently asked questions

Is web scraping legal in the US?

Reading public pages with a program is not access without authorization under the federal anti-hacking law, the Ninth Circuit held in hiQ v. LinkedIn in 2022. Contract, copyright and privacy claims still apply. hiQ itself ended with a $500,000 judgment for LinkedIn under its User Agreement.

Is web scraping legal in the EU and the UK?

No law we read bans it as such. If the data identifies people, the GDPR applies: collecting it is processing, and it needs a lawful basis such as legitimate interests. The UK's ICO was one of 12 regulators that said in 2023 that public personal data is still protected. For copying content, EU law lets owners opt out of text and data mining.

Is it legal to scrape public data?

Public is not the same as free to use. Public pages fall outside the US anti-hacking law, but personal data stays protected, a site's terms can bind you, and copyright still covers the content.

Is robots.txt legally binding?

Not by itself. RFC 9309 says its rules are not a form of access authorization. But 6 of the 20 big sites we read write robots.txt into their terms, so ignoring it there can breach the contract.

Can I scrape Amazon, LinkedIn or Facebook?

Their terms forbid it without permission, as do the terms of 17 of the 20 sites we read. LinkedIn and Facebook also shut every unnamed bot out in robots.txt. All three publish an official API, which comes with its own terms.

Does a proxy make scraping legal?

No. A proxy changes the address a site sees, not the law or the contract. The court order that ended hiQ banned it from using LinkedIn directly or indirectly through a third party, intermediary, or proxy.

Get proxies that are alive right now

Our free proxy list re-checks every exit every few minutes across 100+ countries, with a live last-checked time, so you copy IPs that worked moments ago, not a stale text dump. When the location has to survive a real check, the paid network holds up.

129M+ proxy checks run · 100+ countries · HTTP / HTTPS / SOCKS · re-checked every few minutes · no signup

HProxy.

Honest guides and comparisons on proxies, scraping and staying unblocked, from the team that runs the network.

RSS feed