Skip to content

HTTP cookies explained: how they work and why they matter when scraping

How HTTP cookies work, what Secure, HttpOnly and SameSite mean, what 100 top sites set on a first visit, and how to keep cookies in curl and Python.

HProxy TeamOctober 11, 2026Updated October 11, 20267 min read

HTTP cookies explained: how they work and why they matter when scraping

An HTTP cookie is a small piece of data that a website stores in your browser. The standard behind it, RFC 6265, says cookies let servers "maintain a stateful session over the mostly stateless HTTP protocol". The server sets a cookie with a Set-Cookie header. The browser sends it back with later requests, in a Cookie header. That is how a site keeps you logged in, keeps your cart, or knows you came back.

We opened 100 of the most visited websites on 11 October 2026. 57 of them set at least one cookie in their very first answer.

It takes two headers. The site's answer carries the cookie. Here is the one httpbin.org set in our test:

set-cookie: session=hproxy-demo-1; Path=/

From then on, every request to that site carries it back:

Cookie: session=hproxy-demo-1

The part before the first semicolon is the name and the value. Everything after it is an attribute. The attributes tell the browser how long to keep the cookie, and which requests it may go with.

RFC 6265, from April 2011, is still the standard. It does not define SameSite or Partitioned, though. Browsers added those two, and MDN documents them.

The share in the last column comes from our survey of 313 cookies, further down.

AttributeWhat it doesShare of the 313 cookies
Expires or Max-AgeHow long the cookie lives. Max-Age counts seconds and wins over Expires85%
DomainWhich hosts get the cookie. With it, subdomains get it too78%
PathWhich paths on the site get the cookienot counted
SecureOnly over a secure connection, such as HTTPS75%
HttpOnlyScripts on the page cannot read it42%
SameSiteWhether the cookie goes with requests from other sites61%
PartitionedKept apart for each top-level site. Needs Secure6%

Two of these are often mixed up. Secure "limits the scope of the cookie to 'secure' channels". HttpOnly "limits the scope of the cookie to HTTP requests", which keeps it away from scripts. RFC 6265 adds: "the HttpOnly attribute is independent of the Secure attribute". A cookie can have one, both or neither.

MDN sums SameSite up in one line: "Controls whether or not a cookie is sent with cross-site requests". It takes three values:

  • Strict: "Send the cookie only for requests originating from the same site that set the cookie."
  • Lax: the same, plus a click from another site, as long as it loads the whole page with a safe method such as GET.
  • None: "Send the cookie with both cross-site and same-site requests." It needs Secure as well.

When a site leaves it out, browsers decide. MDN notes: "Some browsers use Lax as the default value if SameSite is not specified".

Session and persistent cookies

A cookie with an end date is persistent. Max-Age gives the lifetime in seconds, Expires a date. MDN settles a clash between the two: "If both Expires and Max-Age are set, Max-Age has precedence." A cookie with neither is a session cookie. RFC 6265 says the browser keeps it until "the current session is over". MDN adds a warning: browsers that restore your tabs restore session cookies too.

What cookies do big websites set?

We took the latest Tranco list, a research ranking of the most visited sites, and opened the first 250 domains once each. We kept the first 100 that answered with a web page, so the last one sat at rank 199. Then we read every Set-Cookie header in their answers, including the redirects.

57 of the 100 sites set a cookie right away, 313 in all. The median site set 1, the busiest set 23. 75% of the cookies were Secure, 42% HttpOnly, and 6% Partitioned.

SameSite on the 313 cookies that 100 top sites set on a first visit (%)
SameSite=None
45%
Not set
39%
SameSite=Lax
13%
SameSite=Strict
3%
Most cookies either opt into every cross-site request or leave the choice to the browser.
Source: HProxy survey of Tranco list 8PXYV on 11 October 2026: the first HTTP answers of 100 sites, read on our server · hproxy.comProxy.

The cookies were built to last. About a third asked to live more than a year, and another 26% up to a year. Only 15% were session cookies. 4% had an end date in the past, which is how a site deletes a cookie.

One detail matters if you parse cookies yourself. RFC 6265 asks servers for the standard HTTP date format. Two of the 100 sites wrote their dates in JavaScript's format instead, such as Thu Nov 11 2027 07:35:42 GMT+0000 (Coordinated Universal Time). Browsers cope, because the standard gives them a forgiving way to read dates. Python's own date parser read those four dates without a time zone, and our first count went wrong. Parse cookie dates the way a browser does, or use a library that does.

Why do cookies matter when scraping?

A scraper is a client like a browser, so the same rules apply. It has to keep the cookies a site sets, or every request starts from zero.

In curl, a cookie jar does that. -c jar.txt writes every cookie "after a completed operation", in "the Netscape cookie file format". -b jar.txt sends them back.

In Python, a requests Session does it. Its documentation says it "persists cookies across all requests made from the Session instance". We ran this on our server:

import requests

s = requests.Session()
s.get("https://httpbin.org/cookies/set?session=hproxy-demo-2")
print(s.cookies.get_dict())
print(s.get("https://httpbin.org/cookies").json())
print(requests.get("https://httpbin.org/cookies").json())
{'session': 'hproxy-demo-2'}
{'cookies': {'session': 'hproxy-demo-2'}}
{'cookies': {}}

The Session sent the cookie back. The plain requests.get sent nothing.

Cookies follow you through every proxy

A cookie lives in your scraper, not in the proxy. Change the proxy, and the cookie still goes out. We set one cookie through a proxy, then sent the same cookie jar through a second proxy:

Terminal output: a site sets session=hproxy-demo-1 through proxy A; curl's cookie jar stores it; the same Cookie header then reaches the site through proxy A from ADDRESS-1 and through proxy B from ADDRESS-2
Our own capture, 11 October 2026: cookie-two-proxies.sh on our server, addresses masked

The site saw two different addresses with the same session cookie. Any site can link those two requests to one visitor, whatever the proxies did. So if you rotate proxies to look like separate visitors, give each one its own empty cookie jar. If you need one login, keep one jar on one exit, as we explain in sticky vs rotating proxy sessions.

Bot protection uses cookies as well. Cloudflare places a __cf_bm cookie on sites with its bot products. It "expires after 30 minutes of continuous inactivity" and holds data for Cloudflare's bot score. Akamai's version is covered in how the _abck cookie works.

If a login needs one address for a while, our residential proxies offer sticky sessions.

What this page could not check

  • The survey read only the HTTP headers of one page load per site. Cookies that JavaScript or other sites set later are missing.
  • It ran once, from one server in the United States, with curl and a browser's user agent. Other countries, consent banners and real browsers can bring other cookies.
  • 100 sites from the first 199 of one ranking, on one morning.
  • The two-proxy test used two free proxies and httpbin.org, which does not tie a session to an address. Real sites may.
  • Browsers change their cookie defaults with new versions. We will run the survey again by 11 January 2027.

Sources

  • RFC 6265, HTTP State Management Mechanism, IETF, April 2011: rfc-editor.org.
  • MDN Web Docs, Set-Cookie header, read 11 October 2026: developer.mozilla.org.
  • Requests documentation, Session Objects, read 11 October 2026: requests.readthedocs.io.
  • curl, man page, --cookie-jar and --cookie, read 11 October 2026: curl.se.
  • Cloudflare, Cloudflare Cookies, updated 5 May 2026, read 11 October 2026: developers.cloudflare.com.
  • Tranco, list 8PXYV, created 10 October 2026: tranco-list.eu.
  • Our own tests on our server, 11 October 2026: the cookie survey at 07:35 UTC, one cookie through two proxies at 07:38 UTC, the requests Session at 07:40 UTC.

Frequently asked questions

What is an HTTP cookie?

A small piece of data that a website stores in your browser with a Set-Cookie header. The browser sends it back with later requests in a Cookie header. That lets the site keep a session, because HTTP itself keeps no state between requests.

What is the difference between a session cookie and a persistent cookie?

A persistent cookie has an Expires or Max-Age attribute and lives until then. A session cookie has neither and is kept until the browser's session ends. Browsers that restore sessions can keep it longer than that.

What does HttpOnly mean on a cookie?

Scripts on the page cannot read the cookie, so it only travels with HTTP requests. It is independent of Secure, which keeps a cookie to HTTPS. In our survey, 42% of the cookies that top sites set on a first visit were HttpOnly.

What is the difference between SameSite Lax and Strict?

Strict sends the cookie only with requests from the same site. Lax also sends it when you follow a link to the site from elsewhere, as long as the request is a safe one such as GET. Some browsers treat a cookie without SameSite as Lax.

Do cookies work through a proxy?

Yes, and they follow you. The cookie lives in your browser or scraper, not in the proxy, so it goes out through every proxy you use. In our test, the same session cookie reached a site from two different addresses.

How do I keep cookies when scraping with Python?

Use a requests Session. It stores the cookies from each answer and sends them with the next request. A plain requests.get sends none of them, as our test showed.

Get proxies that are alive right now

That list re-checks every exit every few minutes across 100+ countries, with a live last-checked time, so you copy IPs that worked moments ago, not a stale text dump. When the location has to survive a real check, the paid network holds up.

129M+ proxy checks run · 100+ countries · HTTP / HTTPS / SOCKS · re-checked every few minutes · no signup

HProxy.

Honest guides and comparisons on proxies, scraping and staying unblocked, from the team that runs the network.

RSS feed