A pragmatic pipeline around trafilatura for JS-rendered pages and list discovery.
Project description
trafi-pipeline
基于 trafilatura 的“组合式”网页正文抽取管线:先抓取/渲染,再抽取;支持列表页发现详情页,再批量抽取。
特性
- 自动渲染策略:
auto | always | never,正文过短或抓取失败时可自动切换渲染 - 列表页爬取:按深度与页数限制发现详情页链接
- 代理支持:抓取/渲染分别配置代理
- 图片处理:保留图片、追加图片列表、或原位插入(Markdown)
- 元数据:标题、来源站点、耗时等
安装
pip install trafi-pipeline
可选依赖:
pip install "trafi-pipeline[http]" # 使用 httpx
pip install "trafi-pipeline[render]" # 使用 Playwright 进行渲染
首次使用 Playwright 需要安装浏览器:
python -m playwright install chromium
快速开始
from trafipipe import Pipeline, PipelineConfig
pipeline = Pipeline(PipelineConfig())
result = pipeline.extract_url("https://example.com/article")
print(result.text)
返回字段(ExtractResult):
text:正文title:标题source:来源站点images:图片 URL 列表videos:视频 URL 列表used_render:是否使用渲染status_code:抓取到的 HTTP 状态码(抓取失败时可能为空)elapsed_ms:耗时(毫秒)fetch_ms:抓取耗时(毫秒)render_ms:渲染耗时(毫秒)extract_ms:正文抽取耗时(毫秒)image_ms:图片收集耗时(毫秒)video_ms:视频收集耗时(毫秒)error:错误信息(如有)
常见配置
代理与渲染
from trafipipe import Pipeline, PipelineConfig, ProxyConfig
cfg = PipelineConfig()
cfg.fetch.proxy = ProxyConfig(http="http://user:pass@host:port", https="http://user:pass@host:port")
cfg.render.proxy = ProxyConfig(server="http://user:pass@host:port")
cfg.render.extra_headers = {
"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
"Referer": "https://mp.weixin.qq.com/",
}
cfg.render.cookies = [
{"name": "your_cookie", "value": "xxx", "domain": ".mp.weixin.qq.com", "path": "/"}
]
cfg.render.reuse_context = True # 批量渲染时复用 context 以提速
pipeline = Pipeline(cfg)
result = pipeline.extract_url("https://example.com/article")
图片保留与输出格式
cfg = PipelineConfig()
cfg.extract.keep_images = True
cfg.extract.append_images = True # 在正文末尾追加 [Images] 列表
cfg.extract.inline_images = False # 设为 True 时输出 Markdown 并原位插入图片
cfg.extract.keep_videos = True
cfg.extract.append_videos = False # 在正文末尾追加 [Videos] 列表
cfg.extract.inline_videos = False # 设为 True 时会在正文中插入 [Video] url
cfg.extract.output_format = "txt" # "txt" 或 "md"
说明:
- 当
inline_images=True且output_format="txt"时,会把转为[Image] url。 - 当
inline_videos=True时,会把 HTML 中的<video>/<source>转成[Video] url。
微信文章图片(mp.weixin.qq.com)
from trafipipe import Pipeline, PipelineConfig
cfg = PipelineConfig()
cfg.extract.keep_images = True
cfg.extract.append_images = True
cfg.render.mode = "auto" # 如图片仍缺失可改为 "always"
cfg.render.extra_headers = {"Referer": "https://mp.weixin.qq.com/"}
cfg.render.cookies = [
{"name": "your_cookie", "value": "xxx", "domain": ".mp.weixin.qq.com", "path": "/"}
]
result = Pipeline(cfg).extract_url("https://mp.weixin.qq.com/s/xxxxxx")
print(result.images)
列表页发现链接并抽取
from trafipipe import Pipeline, PipelineConfig
cfg = PipelineConfig()
cfg.crawl.max_pages = 50
cfg.crawl.max_depth = 2
cfg.crawl.max_workers = 4 # 并发抓取列表页
pipeline = Pipeline(cfg)
urls = pipeline.crawl(["https://example.com/list"])
results = pipeline.crawl_and_extract(urls, max_workers=4)
CLI
trafipipe extract https://example.com/article
trafipipe crawl https://example.com/list --max-pages 50 --max-depth 2 --workers 4
开发
pip install -e ".[dev]"
pytest
ruff check .
性能基准
python doc/benchmark.py --file doc/urls.txt --render auto --repeat 1
python doc/benchmark.py --file doc/urls.txt --format csv --summary > report.csv
python doc/benchmark.py --file doc/urls.txt --format json --summary > report.json
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
trafi_pipeline-0.1.0.tar.gz
(108.5 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file trafi_pipeline-0.1.0.tar.gz.
File metadata
- Download URL: trafi_pipeline-0.1.0.tar.gz
- Upload date:
- Size: 108.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cac47eddbd7e17ae88599b5ebce60d026b118f377a40751474ee7e05b5a65073
|
|
| MD5 |
19700e3deaee398229959d728b3669d1
|
|
| BLAKE2b-256 |
74d21d0b382d97751f0e9dcf5f268352d6f3221746fe917d50dce46fb429d8cb
|
File details
Details for the file trafi_pipeline-0.1.0-py3-none-any.whl.
File metadata
- Download URL: trafi_pipeline-0.1.0-py3-none-any.whl
- Upload date:
- Size: 19.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8f6428ec74ef44b4b013a2eba0c5bb6208a17d8baba6495b70fe357c87082c6a
|
|
| MD5 |
9030307a41251c782f65a548085e9feb
|
|
| BLAKE2b-256 |
9c957a1a53719a2e481ed939d31ab16401d080c2ff3885649c471c228bd88869
|