Skip to main content

web-article-extractor

A small, dependency-light toolkit for pulling readable content off the web. It extracts article text from any URL using a two-stage strategy — trafilatura first (fast, no browser), then a headless Playwright Chromium fallback for JavaScript-heavy pages — and fetches transcripts from YouTube videos through a manual-then-automatic subtitle cascade.

Naming. The PyPI distribution is harnais-web-extractor (this repository keeps the name web-article-extractor). The Python import package is web_article_extractor.

Install

From PyPI:

pip install harnais-web-extractor
# Playwright also needs a browser binary the first time:
playwright install chromium
# YouTube transcript extraction requires yt-dlp (optional extra):
pip install "harnais-web-extractor[youtube]"

Or from source (GitHub):

pip install git+https://github.com/JohnLinotte/web-article-extractor.git

Usage

Command line

# Extract an article as Markdown (default):
python -m web_article_extractor https://example.com/some-article

# As JSON:
python -m web_article_extractor https://example.com/some-article --format json

# A YouTube URL fetches the transcript instead:
python -m web_article_extractor "https://www.youtube.com/watch?v=dQw4w9WgXcQ"

YouTube transcripts require yt-dlp. Install it via the youtube extra (pip install "harnais-web-extractor[youtube]") or provide any yt-dlp binary on your PATH. Without it, fetch_transcript() cannot run.

yt-dlp tuning (optional env vars, all empty by default so the package works on any machine):

  • YT_DLP_COOKIES_FROM_BROWSER=firefox — pass --cookies-from-browser firefox to yt-dlp (needed for age-restricted or members-only videos).
  • YT_DLP_JS_RUNTIME=node — pass --js-runtime node.
  • YT_DLP_BIN=/path/to/yt-dlp — override the binary location.

Python API

from web_article_extractor import extract_article, fetch_transcript, is_youtube_url

result = extract_article("https://example.com/some-article")
if result:
    print(result["title"])
    print(result["content"])      # Markdown
    print(result["word_count"])

if is_youtube_url(url):
    transcript = fetch_transcript(url)
    if transcript:
        print(transcript["text"])

extract_article returns a dict with url, title, content, source_method, extracted_at and word_count, or None when extraction fails entirely.

License

MIT — see LICENSE.

Metadata

Release files for harnais-web-extractor 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for harnais-web-extractor 0.1.2
File Size Uploaded
harnais_web_extractor-0.1.2.tar.gz 13.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for harnais-web-extractor 0.1.2
File Interpreter ABI Platform
harnais_web_extractor-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 27.3 kB

Release files / harnais_web_extractor-0.1.2.tar.gz

Download URL harnais_web_extractor-0.1.2.tar.gz
Size 13.4 kB
Tags Source
SHA-256 checksum
How to use checksums
362388526a0cadee01168e5465ba4bbe389449ee5d85184807db02387bc83965
BLAKE2b-256 checksum
How to use checksums
de32c8fe7eacb006c0dc2ed186282c91571a05597a9b2e883f39cad412ef27c2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.13

Release files / harnais_web_extractor-0.1.2-py3-none-any.whl

Download URL harnais_web_extractor-0.1.2-py3-none-any.whl
Size 13.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cbbc7b358c8f41598a04df179dd6346d64554965ab62dc82c162a4b826565d09
BLAKE2b-256 checksum
How to use checksums
421bc3781c34406e2d808e388cc130dace41680cd03a4c6e31884fbcd28d37a9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.13

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page