Skip to main content
https://badge.fury.io/py/htmlement.svg https://readthedocs.org/projects/python-htmlement/badge/?version=stable https://github.com/willforde/python-htmlement/actions/workflows/tests.yml/badge.svg?branch=master&event=push https://codecov.io/gh/willforde/python-htmlement/branch/master/graph/badge.svg?token=D5EKKLIVBP Maintainability

HTMLement

HTMLement is a pure Python HTML Parser.

The object of this project is to be a “pure-python HTML parser” which is also “faster” than “beautifulsoup”. And like “beautifulsoup”, will also parse invalid html.

The most simple way to do this is to use ElementTree XPath expressions. Python does support a simple (read limited) XPath engine inside its “ElementTree” module. A benefit of using “ElementTree” is that it can use a “C implementation” whenever available.

This “HTML Parser” extends html.parser.HTMLParser to build a tree of ElementTree.Element instances.

Install

Run

pip install htmlement

-or-

pip install git+https://github.com/willforde/python-htmlement.git

Parsing HTML

Here I’ll be using a sample “HTML document” that will be “parsed” using “htmlement”:

html = """
<html>
  <head>
    <title>GitHub</title>
  </head>
  <body>
    <a href="https://github.com/marmelo">GitHub</a>
    <a href="https://github.com/marmelo/python-htmlparser">GitHub Project</a>
  </body>
</html>
"""

# Parse the document
import htmlement
root = htmlement.fromstring(html)

Root is an ElementTree.Element and supports the ElementTree API with XPath expressions. With this I’m easily able to get both the title and all anchors in the document.

# Get title
title = root.find("head/title").text
print("Parsing: %s" % title)

# Get all anchors
for a in root.iterfind(".//a"):
    print(a.get("href"))

And the output is as follows:

Parsing: GitHub
https://github.com/willforde
https://github.com/willforde/python-htmlement

Parsing HTML with a filter

Here I’ll be using a slightly more complex “HTML document” that will be “parsed” using “htmlement with a filter” to fetch only the menu items. This can be very useful when dealing with large “HTML documents” since it can be a lot faster to only “parse the required section” and to ignore everything else.

html = """
<html>
  <head>
    <title>Coffee shop</title>
  </head>
  <body>
    <ul class="menu">
      <li>Coffee</li>
      <li>Tea</li>
      <li>Milk</li>
    </ul>
    <ul class="extras">
      <li>Sugar</li>
      <li>Cream</li>
    </ul>
  </body>
</html>
"""

# Parse the document
import htmlement
root = htmlement.fromstring(html, "ul", attrs={"class": "menu"})

In this case I’m not unable to get the title, since all elements outside the filter were ignored. But this allows me to be able to extract all “list_item elements” within the menu list and nothing else.

# Get all listitems
for item in root.iterfind(".//li"):
    # Get text from listitem
    print(item.text)

And the output is as follows:

Coffee
Tea
Milk

Metadata

Release files for htmlement 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for htmlement 2.0.0
File Size Uploaded
htmlement-2.0.0.tar.gz 8.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for htmlement 2.0.0
File Interpreter ABI Platform
htmlement-2.0.0-py2.py3-none-any.whl Python 2, Python 3 none any Details

Total release size: 17.5 kB

Release files / htmlement-2.0.0.tar.gz

Download URL htmlement-2.0.0.tar.gz
Size 8.9 kB
Tags Source
SHA-256 checksum
How to use checksums
97ac10e8fcad471257f5cce2a2355f292394fda7c8266c8f9a7d1dd2445ec188
BLAKE2b-256 checksum
How to use checksums
d521669bb4868e2ae702e3545a4f432d4204ae10dfef8e357b026cfdc8679de0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.9.15

Release files / htmlement-2.0.0-py2.py3-none-any.whl

Download URL htmlement-2.0.0-py2.py3-none-any.whl
Size 8.6 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
122580394fa0dad576ad9566f389d2df513194de0cc6d1ddeba20a68dcfc3ffe
BLAKE2b-256 checksum
How to use checksums
17b7449393fdeaca1c9f5085c1b06f115cedbb393a114303628b29903bc59e67
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.9.15

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

1.0.1

2 release files

1.0.0

1 release file

0.2.3

2 release files

0.2.1

2 release files

0.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page