Skip to content

XPath vs CSS selector: which to use for scraping

We ran six scraping tasks with XPath and CSS in lxml, parsel, BeautifulSoup and Chrome. CSS can now select parents; speed flips between engines.

HProxy TeamOctober 11, 2026Updated October 11, 20266 min read

XPath vs CSS selector: which to use for scraping

Both pick elements out of a page. CSS selectors are shorter and match classes as words. XPath can also find an element by its text, step to a sibling or a parent, and return attributes and text.

We ran the same six scraping tasks both ways in lxml, parsel, BeautifulSoup and Chrome on 11 October 2026. The test page was our own: 2,000 product cards, made up for the test.

The short answer

  • Use CSS for classes, IDs and structure: div.product > h2.name. It is shorter, and in a browser it was 27 times faster in our test.
  • Use XPath to find an element by its text, to read the value next to a label, or to get an attribute or text node directly.
  • Going up is no longer XPath-only. CSS :has() selects an element by what it contains.
  • In lxml and parsel, CSS is XPath underneath. Choose by what is easier to read, not by power.

Six tasks in four parsers

Terminal output: six scraping tasks run with XPath and CSS in lxml, parsel, BeautifulSoup and Chrome. Exact-class XPath finds 1960 names where CSS finds 2000; :has() works everywhere; text and label tasks work in XPath and in BeautifulSoup's own extension only
Our own capture, 11 October 2026: xp_test.py and browser_test.mjs on our server

Read it task by task:

  1. Every product name. CSS found 2,000. The common XPath //div[@class='product'] found 1,960.
  2. Prices of products on sale. Both found 285, with :has() on the CSS side.
  3. Names containing "Lamp". Only XPath, with contains(). Standard CSS cannot match text.
  4. The card above each Sale badge. Both found 285: XPath goes up with .., CSS with :has(> span.sale).
  5. The weight next to the "Weight" label. Only XPath, with following-sibling. CSS can find siblings, but not by the label's text.
  6. The third product. Both found it.

The class trap

Task 1 shows a mistake that even published guides make. In XPath, @class='product' compares the whole attribute, so a card with class="product featured" does not match. CSS .product matches the class as one word among several:

from lxml import html

# page holds the HTML of the test page, as a string
tree = html.fromstring(page)
len(tree.xpath("//div[@class='product']"))
# 1960
len(tree.xpath("//div[contains(concat(' ', normalize-space(@class), ' '), ' product ')]"))
# 2000
len(tree.cssselect("div.product"))
# 2000

Forty cards were missing, and nothing raised an error. The long form in the middle is what CSS does for you.

In Python, CSS runs as XPath

lxml and parsel do not have a CSS engine of their own. They hand CSS to cssselect, a library that "translates CSS selectors to XPath 1.0 expressions", in the words of lxml's documentation. You can see the translation:

from cssselect import GenericTranslator

GenericTranslator().css_to_xpath("div.product > h2.name")
# descendant-or-self::div[@class and contains(@class, 'product') and contains(concat(' ',
# normalize-space(@class), ' '), ' product ')]/h2[@class and contains(@class, 'name') and
# contains(concat(' ', normalize-space(@class), ' '), ' name ')]

So in lxml, CSS cannot do anything XPath cannot. It also pays for the translation: 9.90 ms against 5.37 ms in our test. parsel takes both and adds a handy ::text pseudo-element to CSS:

from parsel import Selector

sel = Selector(text=page)
sel.css("div.product:has(span.sale) span.price::text").getall()[:3]
# ['254.30', '529.81', '931.70']
sel.xpath("//h2[contains(., 'Lamp')]/text()").getall()[:3]
# ['Floor Lamp 2', 'Desk Lamp 6', 'Desk Lamp 7']
sel.xpath("//th[.='Weight']/following-sibling::td[1]/text()").getall()[:3]
# ['29 kg', '40 kg', '20 kg']
sel.xpath("//span[contains(@class, 'sale')]/../@data-id").getall()[:3]
# ['7', '14', '21']

Which is faster?

It depends on the engine. We timed the same selection of 2,000 product names, taking the median of 50 runs:

Time to select 2,000 product names (median of 50 runs)
Chrome, CSS (querySelectorAll)
0.2 ms
lxml, XPath
5.37 ms
Chrome, XPath (document.evaluate)
5.5 ms
parsel, XPath
9 ms
lxml, CSS
9.9 ms
parsel, CSS
11.03 ms
BeautifulSoup, CSS
49.4 ms
Source: HProxy test on our server, 11 October 2026: Chrome 143 headless, Python 3.12.3 with lxml 6.1.3, cssselect 1.6.0, parsel 1.12.1, beautifulsoup4 4.15.0; one test page of 2,000 cards · hproxy.comProxy.

In Chrome, CSS won by a wide margin. In lxml and parsel, XPath won, because their CSS runs as XPath. BeautifulSoup was the slowest at selecting, and slower still at parsing: 423.3 ms with the lxml parser against 14.7 ms for lxml itself.

Text, parents and siblings

These are the three jobs that decide it.

Text. XPath 1.0 has contains(), which "returns true if the first argument string contains the second argument string". Standard CSS has nothing like it: Chrome threw a SyntaxError. Two libraries add their own. BeautifulSoup's Soup Sieve has :-soup-contains(), and its documentation notes that a :contains() pseudo-class was once in a CSS draft but "was dropped from the draft in the end":

from bs4 import BeautifulSoup

soup = BeautifulSoup(page, "lxml")
[td.get_text() for td in soup.select('th:-soup-contains("Weight") + td')][:3]
# ['29 kg', '40 kg', '20 kg']

That selector works in BeautifulSoup only. lxml and parsel raised an ExpressionError on it, and Chrome a SyntaxError.

Parents. CSS has :has() now. MDN describes it as "a way of selecting a parent element or a previous sibling element with respect to a reference element", and says it has been "available across browsers since December 2023". In the browser:

document.querySelectorAll("div.product:has(> span.sale)").length;
// 285
document.evaluate("//span[contains(@class, 'sale')]/..", document, null,
  XPathResult.ORDERED_NODE_SNAPSHOT_TYPE, null).snapshotLength;
// 285

Siblings by label. Pages often put values in tables of labels. XPath's following-sibling axis "contains all the following siblings of the context node", so //th[.='Weight']/following-sibling::td[1] reads exactly the right cell. CSS th + td finds the cell after every label, with no way to say which label.

In Selenium and Playwright

Both take both. Selenium lists "css selector" and "xpath" among its locators. Playwright says its "XPath locators are equivalent to calling Document.evaluate", so the browser timings above apply.

Playwright adds one difference that matters on modern sites: "CSS selectors pierce open shadow DOM", while "XPath does not pierce shadow roots". Its documentation also recommends neither as a first choice: "We recommend prioritizing user-visible locators like text or accessible role instead of using CSS that is tied to the implementation and could break when the page changes."

Where a proxy fits

Selectors read a page you already have. A proxy matters one step earlier, when the site limits how many pages one address may fetch. Our guides for Scrapy, Playwright and Selenium show the setup, and our page on whether web scraping is legal covers what a site's terms allow.

What this page could not check

  • One test page, made up for the test, on one server. Real pages are messier.
  • Timings are medians of 50 runs and moved by up to a quarter between two runs. The order stayed the same.
  • We tested Chrome 143 only, not Firefox or Safari.
  • We did not test XPath 2.0 or later. Browsers and lxml run XPath 1.0.
  • We will rerun the tests with current versions by 11 January 2027.

Sources

  • W3C, XML Path Language (XPath) Version 1.0, Recommendation of 16 November 1999: w3.org.
  • W3C, Selectors Level 4, Working Draft of 22 January 2026: w3.org.
  • MDN, :has() and Document.evaluate(), read 11 October 2026: developer.mozilla.org.
  • cssselect documentation, version 1.6.0: cssselect.readthedocs.io. lxml, cssselect: lxml.de.
  • Soup Sieve, pseudo-classes: facelessuser.github.io.
  • parsel documentation: parsel.readthedocs.io.
  • Playwright, other locators: playwright.dev. Selenium, locator strategies: selenium.dev.
  • Our own tests on our server, 11 October 2026: xp_test.py at 21:05 UTC and browser_test.mjs at 21:05 UTC, with the versions in the chart's source line.

Frequently asked questions

What is better, XPath or CSS selector?

For most scraping, CSS: it is shorter and matches classes as words. Use XPath when you need an element by its text, a value next to a label, or an attribute as the result. In our test, standard CSS could not do two of six tasks; XPath did all six.

Which is faster, XPath or CSS selectors?

It depends on the engine. In Chrome, querySelectorAll took 0.20 ms and document.evaluate 5.50 ms for the same 2,000 elements. In lxml, XPath took 5.37 ms and CSS 9.90 ms, because lxml turns CSS into XPath first. BeautifulSoup took 49.40 ms.

Can a CSS selector select a parent element?

Yes, with :has(), which works in all major browsers since December 2023 and in Python's cssselect and Soup Sieve. In our test, div.product:has(> span.sale) found the same 285 cards as the XPath parent step.

Can CSS select an element by its text?

Not standard CSS: Chrome threw a SyntaxError. BeautifulSoup's Soup Sieve adds :-soup-contains(), and Playwright adds :has-text(). XPath does it with contains().

XPath or CSS in Selenium and Playwright?

Both work in both, because the browser runs them. Playwright says its XPath locators are equivalent to calling Document.evaluate. Its CSS pierces open shadow DOM and its XPath does not, and it recommends user-visible locators over either.

Get proxies that are alive right now

Our free proxy list re-checks every exit every few minutes across 100+ countries, with a live last-checked time, so you copy IPs that worked moments ago, not a stale text dump. When the location has to survive a real check, the paid network holds up.

129M+ proxy checks run · 100+ countries · HTTP / HTTPS / SOCKS · re-checked every few minutes · no signup

HProxy.

Honest guides and comparisons on proxies, scraping and staying unblocked, from the team that runs the network.

RSS feed