web-article-extractor
A small, dependency-light toolkit for pulling readable content off the web. It
extracts article text from any URL using a two-stage strategy — trafilatura
first (fast, no browser), then a headless Playwright Chromium fallback for
JavaScript-heavy pages — and fetches transcripts from YouTube videos through a
manual-then-automatic subtitle cascade.
Naming. The PyPI distribution is
harnais-web-extractor(this repository keeps the nameweb-article-extractor). The Python import package isweb_article_extractor.
Install
From PyPI:
pip install harnais-web-extractor
# Playwright also needs a browser binary the first time:
playwright install chromium
# YouTube transcript extraction requires yt-dlp (optional extra):
pip install "harnais-web-extractor[youtube]"
Or from source (GitHub):
pip install git+https://github.com/JohnLinotte/web-article-extractor.git
Usage
Command line
# Extract an article as Markdown (default):
python -m web_article_extractor https://example.com/some-article
# As JSON:
python -m web_article_extractor https://example.com/some-article --format json
# A YouTube URL fetches the transcript instead:
python -m web_article_extractor "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
YouTube transcripts require
yt-dlp. Install it via theyoutubeextra (pip install "harnais-web-extractor[youtube]") or provide anyyt-dlpbinary on yourPATH. Without it,fetch_transcript()cannot run.yt-dlp tuning (optional env vars, all empty by default so the package works on any machine):
YT_DLP_COOKIES_FROM_BROWSER=firefox— pass--cookies-from-browser firefoxto yt-dlp (needed for age-restricted or members-only videos).YT_DLP_JS_RUNTIME=node— pass--js-runtime node.YT_DLP_BIN=/path/to/yt-dlp— override the binary location.
Python API
from web_article_extractor import extract_article, fetch_transcript, is_youtube_url
result = extract_article("https://example.com/some-article")
if result:
print(result["title"])
print(result["content"]) # Markdown
print(result["word_count"])
if is_youtube_url(url):
transcript = fetch_transcript(url)
if transcript:
print(transcript["text"])
extract_article returns a dict with url, title, content,
source_method, extracted_at and word_count, or None when extraction
fails entirely.
License
MIT — see LICENSE.
Metadata
Release files for harnais-web-extractor 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| harnais_web_extractor-0.1.2.tar.gz | 13.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| harnais_web_extractor-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 27.3 kB
Release files / harnais_web_extractor-0.1.2.tar.gz
| Download URL | harnais_web_extractor-0.1.2.tar.gz |
|---|---|
| Size | 13.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
362388526a0cadee01168e5465ba4bbe389449ee5d85184807db02387bc83965
|
|
BLAKE2b-256 checksum How to use checksums |
de32c8fe7eacb006c0dc2ed186282c91571a05597a9b2e883f39cad412ef27c2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.13
|
Release files / harnais_web_extractor-0.1.2-py3-none-any.whl
| Download URL | harnais_web_extractor-0.1.2-py3-none-any.whl |
|---|---|
| Size | 13.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cbbc7b358c8f41598a04df179dd6346d64554965ab62dc82c162a4b826565d09
|
|
BLAKE2b-256 checksum How to use checksums |
421bc3781c34406e2d808e388cc130dace41680cd03a4c6e31884fbcd28d37a9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.13
|