Skip to main content

Trafilatura Core

Trafilatura Core PyPI version PyPI downloads license

Also available as:

Online playground | npm package CLI & lib | Source code on GitHub

Docs

Getting started | Python library help

Social

Star us on GitHub | Follow us on GitHub

Trafilatura Core extracts main content by removing boilerplate from HTML documents.

  • Two language versions — TypeScript and Python: available as a TypeScript library on npm and a Python library on PyPI.

  • This is Trafilatura Core's open-source native Python implementation. It translates the TypeScript port of Python Trafilatura's fast=True extraction path. go-trafilatura served as a DOM translation aid for the TypeScript port.

  • The Core in its name means it is reduced to one task: extracting main content by removing boilerplate. An optional Source URL supplies metadata and image-resolution context; it is never fetched. Use Turndown for Markdown conversion and Markdownee for crawling live websites.

  • Compared with Mozilla Readability, Trafilatura and Trafilatura Core use layered structural and content heuristics with fallback and recall escalation, rather than centering extraction on the candidate scoring inherited from Arc90’s original readability.js article extractor; Trafilatura Core also offers configuration options for boilerplate removal.

This package provides the native Python library, with library APIs only. It returns cleaned HTML, diagnostics, and available page metadata; see the Python reference for language differences.

Install and use

Requires Python 3.10 or newer:

pip install trafilaturacore

Save this as clean.py:

from trafilaturacore import clean

source = """<nav>Home</nav><article><h1>Reading saved pages</h1>
<p>Save the original HTML before cleaning a page. A local copy lets
you compare the extracted article with its navigation and footer.</p>
<p>Keep the source address beside the snapshot. It provides context
when relative image links need to be resolved after extraction.</p>
</article>"""
result = clean(source)
print(result.html)
python clean.py

The result prints the article heading and paragraphs without navigation. clean() accepts a string or UTF-8 bytes. Its CleanResult has html, messages, and optional metadata; aclean() accepts the same options for async callers. Cancelling its await does not stop an already running worker.

Configure cleanup

Select precision, balanced (default), or recall for extraction; keep cleans the whole document. Image, link, table, and user-comment handling are independent. The Python reference covers options, custom policies, async usage, resource limits, and errors.

Python uses lxml and nh3. Parsing, serialization, date/URL handling, custom configuration, CSS preservation, diagnostics, and limits can differ from the TypeScript library. Dependencies install separately; platforms without compatible dependency wheels need their build prerequisites.

Extraction is not a security boundary. Before rendering untrusted output, apply sanitization appropriate to your output context and a Content Security Policy. Remote links and images may remain.

Acknowledgements

  • Trafilatura — original Python implementation by Adrien Barbaresi.
  • go-trafilatura — Go port by Markus Mobius, used as a DOM translation aid.

See About for extraction lineage and scope.

Support

Report problems in the issue tracker.

License

Licensed under Apache-2.0. See the compact third-party notices.

Metadata

Release files for trafilaturacore 0.8.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for trafilaturacore 0.8.9
File Interpreter ABI Platform
trafilaturacore-0.8.9-py3-none-any.whl Python 3 none any Details

Release files / trafilaturacore-0.8.9-py3-none-any.whl

Download URL trafilaturacore-0.8.9-py3-none-any.whl
Size 111.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3cbb9bfa5ded1e22549faba500f576fc5887e0d0621108d5dd588fc458aea1b2
BLAKE2b-256 checksum
How to use checksums
01f355883efd01b8ce0a9a059eb7a416b20c8e003d5a1cd32487e3ba2429992c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.8.10

1 release file

This release

0.8.9 This release

1 release file

0.8.6

1 release file

0.8.5

1 release file

0.8.4

1 release file

0.8.2

1 release file

0.8.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page