0.2.0: Works on Windows! See documentation for available DLL download locations. Documentation rewritten and expanded.
PyTidyLib is a Python package that wraps the HTML Tidy library. This allows you, from Python code, to “fix” invalid (X)HTML markup. Some of the library’s many capabilities include:
Clean up unclosed tags and unescaped characters such as ampersands
Output HTML 4 or XHTML, strict or transitional, and add missing doctypes
Convert named entities to numeric entities, which can then be used in XML documents without an HTML doctype.
Clean up HTML from programs such as Word (to an extent)
Indent the output, including proper (i.e. no) indenting for pre elements, which some (X)HTML indenting code overlooks.
Small example of use
The following code cleans up an invalid HTML document and sets an option:
from tidylib import tidy_document
document, errors = tidy_document('''<p>fõo <img src="bar.jpg">''',
options={'numeric-entities':1})
print document
print errors
Docs
Documentation is shipped with the source distribution and is available at the PyTidyLib web page.
Metadata
Release files for pytidylib6 0.2.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pytidylib6-0.2.2.tar.gz | 155.1 kB | Details |
Release files / pytidylib6-0.2.2.tar.gz
| Download URL | pytidylib6-0.2.2.tar.gz |
|---|---|
| Size | 155.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
421ae35f32a32610cf3e8a5c85b830c92fb8c620192b42916fb8a4a34d5bc006
|
|
BLAKE2b-256 checksum How to use checksums |
a54498ddad5e111352282b82f763aba8bf75da9adb789f473a588028edbbf145
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |