Skip to main content

boilerpot

Templating content from HTML. A Python do-alike to boilerpipe.

boilerpipe (http://code.google.com/p/boilerpipe/) is a Java program that looks at HTML tags and tries to deduce where the actual content is sans navigation, headers & footers, etc.

This is a rough rewrite of that written in Python and should be considered super duper alpha. I did it during a 2-day company (Curata.com) hackathon. I haven’t even run a comparison of its output against its step-father let alone done any corpus comparisons against commoncrawl.org. Consider yourself warned.

The only advantages over boilerpipe is that it is easier to interface with Python and the code is much more accessible: 500 lines of Python in one module versus 9000 lines of Java scattered accross a bazillion files and directories (I hate me some directories).

Release files for boilerpot 0.92

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for boilerpot 0.92
File Size Uploaded
boilerpot-0.92.tar.gz 9.4 kB Details

Release files / boilerpot-0.92.tar.gz

Download URL boilerpot-0.92.tar.gz
Size 9.4 kB
Tags Source
SHA-256 checksum
How to use checksums
1a495f3f428c28898261c704bb777ea7df75abdf86789356c79c9be252e25119
BLAKE2b-256 checksum
How to use checksums
b99027e7b4bf2d47ca1689746fd120cc5ffa87826c4af2a29ca0a98f249dc461
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No

Release history Release notifications | RSS feed

This release

0.92 This release

1 release file

0.91

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page