HTML Table Takeout
A fast, lightweight HTML table parser that supports rowspan, colspan, links and nested tables. No external dependencies are needed.
The input may be text, a URL or local file Path.
HTML5 logo by W3C.
Quick Start
Install the package:
pip install html-table-takeout
Pass in a URL and print out the parsed Table as CSV:
from html_table_takeout import parse_html
# start with http:// or https:// to source from a URL
tables = parse_html('https://en.wikipedia.org/wiki/List_of_S%26P_500_companies')
print(tables[0].to_csv())
# output:
# Symbol,Security,GICS Sector,GICS Sub-Industry,Headquarters Location,Date added,CIK,Founded
# MMM,3M,Industrials,Industrial Conglomerates,"Saint Paul, Minnesota",1957-03-04,0000066740,1902
# ...
Pass in HTML text and print out the parsed Table as valid HTML:
from html_table_takeout import parse_html
tables = parse_html("""
<table>
<tr>
<td rowspan='2'>1</td> <!-- rowspan will be expanded -->
<td>2</td>
</tr>
<tr>
<td>3</td>
</tr>
</table>""")
print(tables[0].to_html(indent=4))
# output:
# <table data-table-id='0'>
# <tbody>
# <tr>
# <td>1</td>
# <td>2</td>
# </tr>
# <tr>
# <td>1</td>
# <td>3</td>
# </tr>
# </tbody>
# </table>
Usage
The core parse_html() function returns a list of zero or more top-level Table. A Table is guaranteed to have this structure:
- rows: List of one or more
TRow- cells: List of zero or more
TCellresulting from rowspan and colspan expansion- elements: List of zero or more
TText,TLink,TRef
- elements: List of zero or more
- cells: List of zero or more
| Type | Description |
|---|---|
Table |
Each parsed table has an auto-assigned unique id |
TRow |
Equal to each <tr> in the original table |
TCell |
Expanded <td> or <th> cells from row/colspan |
TText |
HTML-decoded text inside <td> or <th> |
TLink |
Equal to each <a> inside <td> or <th> |
TRef |
Reference to the child Table |
All tables are guaranteed to have at least one TRow containing one TCell.
The parse_html() function also provides filtering by text or attributes to target the tables you want. Check out its docstring for all options.
Why did you make this
Most HTML table parsers require extra DOM and data processing libraries that aren't needed for my application. I need a parser that handles nesting and gives me the flexibility to process the parsed result however I want.
Now you too can take out tables to go.
Developing
Install development dependencies:
pip install build mypy pytest
Run tests:
pytest
Build the package:
python -m build
Metadata
Release files for html-table-takeout 1.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| html_table_takeout-1.1.2.tar.gz | 10.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| html_table_takeout-1.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 20.1 kB
Release files / html_table_takeout-1.1.2.tar.gz
| Download URL | html_table_takeout-1.1.2.tar.gz |
|---|---|
| Size | 10.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e2021e7c93271d2c08c1c2fb09f8408f6aac36d353e37cea0a196f3fe54e7f30
|
|
BLAKE2b-256 checksum How to use checksums |
e302e2e9ef488b6400ea8bcd3ce95e5ec67e8be3368f6014ad014017072dcec9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.13.2
|
Release files / html_table_takeout-1.1.2-py3-none-any.whl
| Download URL | html_table_takeout-1.1.2-py3-none-any.whl |
|---|---|
| Size | 9.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e590c8a41b455722b8eac89ecadb3a1560a563994659b9186f50839ef2ba3238
|
|
BLAKE2b-256 checksum How to use checksums |
eec803180b73e9928014c7ebc4f9f0b89eac5f1278813155e22f98d5b3b6f081
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.13.2
|