Skip to main content

html2ans

https://img.shields.io/pypi/v/html2ans.svg https://img.shields.io/pypi/pyversions/html2ans.svg https://circleci.com/gh/washingtonpost/html2ans.svg?style=shield https://img.shields.io/pypi/l/html2ans.svg

This project provides a standardized method of parsing HTML elements into ANS elements. It is mainly used by Arc Publishing’s professional services team to migrate client data into the Arc platform, but can also be used for arbitrary conversion of HTML to JSON.

html2ans is hosted on pypi.

Please use the GitHub issue tracker to submit bugs or request features.

Full documentation can be found here.

Quickstart

Generating ANS from HTML

from html2ans.default import Html2Ans

parser = Html2Ans()
content_elements = parser.generate_ans(your_html_here)

Adding Parsers

Basic Addition

If you need to parse a certain tag in a customized way, you can write your own parser class and add it to the parsers Html2Ans will use like so:

from html2ans.default import Html2Ans

parser = Html2Ans()
parser.add_parser(YourCustomImageParser())
parser.generate_ans(your_html_here)

The types of items your parser can parse should be listed in its applicable_elements attribute.

The default parser class (DefaultHtmlAnsParser or Html2Ans) has parsers for text, links, images, various social media embeds, etc.

Prioritized Addition

The parsers that can be used for each element type (e.g. img, p) are held in a list. If you want your parser to have a higher priority than the default parsers, add it like so:

from html2ans.default import Html2Ans

parser = Html2Ans()
parser.insert_parser('img', YourCustomImageParser(), 0)
parser.generate_ans(your_html_here)

Creating Custom Parsers

Missing from the snippet above is a definition of YourCustomImageParser. Before talking about how to create such a parser, let’s examine why you might need to do so.

The default image parser html2ans.parsers.image.ImageParser applies to html img tags only. Imagine you need to parse html whose images come in div tags (labelled with the class fancy-figure) that also hold a caption (labelled with the class fancy-caption). Here is a possible implementation of a parser for such images (note: this returns basic image ANS, not a reference):

from html2ans.parsers.image import ImageParser
from html2ans.parsers.base import ParseResult

class YourCustomImageParser(ImageParser):
    applicable_elements = ['div', 'figure']
    applicable_classes = ['fancy-figure']

    def parse(self, element, *args, **kwargs):
        image_tag = element.find('img')
        caption_tag = element.find('p', {"class": "fancy-caption"})
        if image_tag:
            image = self.construct_output(image_tag)
            if caption_tag:
              image["caption"] = caption_tag.text
            return ParseResult(image, True)
        return ParseResult(None, True)

Custom Parsing Tips

ANS Versions

Some ANS types require a version. You can set a version in your main parser (Html2Ans) and then automatically include that version in any element parser’s output by setting the parser’s version_required attribute to True.

Note: this doesn’t mean valid, version-compatible ANS is automatically produced!

Keeping HTML in text Output

To adjust what HTML is/isn’t left inline when parsing text, adjust the INLINE_TAGS attribute on the text parser. Every parser inherits from html2ans.parsers.utils.AbstractParserUtilities which provides a list of default INLINE_TAGS which can be used to make sure text formatters (e.g. strong, em, etc.) are left in place when text is parsed.

Removing Unnecessary Tags

Sometimes it is helpful to remove unnecessary tags (e.g. <p></p>, <div><img src="..." /></div>). By default, Html2Ans considers p and div tags with no attributes other than id, class, or style to be unnecessary “wrappers”. When these are encountered, they are ignored and their children are parsed.

The benefit of this is that <p></p> is ignored and <div><img src="..." /></div> is parsed as an image.

The downside is that sometimes you don’t want your HTML removed! There are a few options in this case. You can configure what tags can be considered wrappers via the WRAPPER_TAGS attribute on Html2Ans. So if div tags should never be removed, simply remove div from this list. If a more complicated set of rules are necessary, override the is_wrapper method on Html2Ans.

If it’s easier to modify the HTML than to modify this library, you can also add an arbitrary attribute like so: <div no_parse_flag="true">...</div>. This div will not be considered a wrapper when it is encountered.

Metadata

Release files for html2ans 3.0.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for html2ans 3.0.6
File Size Uploaded
html2ans-3.0.6.tar.gz 18.0 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for html2ans 3.0.6
File Interpreter ABI Platform
html2ans-3.0.6-py3.6.egg Legacy Egg format - - Details
html2ans-3.0.6-py2.py3-none-any.whl Python 2, Python 3 none any Details

Total release size: 85.1 kB

Release files / html2ans-3.0.6.tar.gz

Download URL html2ans-3.0.6.tar.gz
Size 18.0 kB
Tags Source
SHA-256 checksum
How to use checksums
6348bf55bfbe45cc16c7614fff3cfba77c1500a9dc2cb07d76bb2e4708523ccb
BLAKE2b-256 checksum
How to use checksums
3e169b652369f28e061ef43d513fc7fcf50949a5420510895120fd6a34cc0a49
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.1.1 pkginfo/1.5.0.1 requests/2.22.0 setuptools/42.0.2 requests-toolbelt/0.9.1 tqdm/4.40.2 CPython/3.6.9

Release files / html2ans-3.0.6-py3.6.egg

Download URL html2ans-3.0.6-py3.6.egg
Size 44.9 kB
Tags Egg
SHA-256 checksum
How to use checksums
7841393568c09efcc71c0b0d213091782abcc5a23dc99cb39902b2df24e5c1c0
BLAKE2b-256 checksum
How to use checksums
def2b4b00d00b12b681f490a3fc3a288795bd7cad46792335ddc132c045fc61a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.1.1 pkginfo/1.5.0.1 requests/2.22.0 setuptools/42.0.2 requests-toolbelt/0.9.1 tqdm/4.40.2 CPython/3.6.9

Release files / html2ans-3.0.6-py2.py3-none-any.whl

Download URL html2ans-3.0.6-py2.py3-none-any.whl
Size 22.2 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
6f9d770e693de409d54766bf407ffdbe366e864bb2c0e06886cb9c123630574e
BLAKE2b-256 checksum
How to use checksums
de51b147d7c2dcab29257c0eab2612f7053cba06d7c5b87b4c85606ade0720cb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.1.1 pkginfo/1.5.0.1 requests/2.22.0 setuptools/42.0.2 requests-toolbelt/0.9.1 tqdm/4.40.2 CPython/3.6.9

Release history Release notifications | RSS feed

This release

3.0.6 This release

3 release files

3.0.5

3 release files

3.0.4

3 release files

3.0.3

3 release files

3.0.2

3 release files

3.0.1

3 release files

3.0.0

3 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page