A pragmatic pipeline around trafilatura for JS-rendered pages and list discovery.
Project description
trafi-pipeline
基于 trafilatura 的“组合式”网页正文抽取管线:先抓取/渲染,再抽取;支持列表页发现详情页,再批量抽取。
特性
- 自动渲染策略:
auto | always | never,正文过短或抓取失败时可自动切换渲染 auto模式还会根据页面“展开/更多/阅读全文”等标记触发渲染,避免只抓到摘要- 列表页爬取:按深度与页数限制发现详情页链接
- 代理支持:抓取/渲染分别配置代理
- 图片处理:保留图片、追加图片列表、或原位插入(Markdown)
- 元数据:标题、来源站点、耗时等
安装
pip install trafi-pipeline
建议完整安装(抓取 + 渲染能力都具备):
pip install "trafi-pipeline[http,render]"
python -m playwright install chromium
分开安装(更细粒度控制):
pip install "trafi-pipeline[http]" # 使用 httpx
pip install "trafi-pipeline[render]" # 使用 Playwright 进行渲染
说明:
- 默认配置
render.mode="auto",可能会触发渲染;若未安装render依赖或未安装浏览器,将导致结果为空或报错。 - 如果不需要渲染,请显式设置
render.mode="never",避免依赖缺失导致失败。
快速开始
from trafipipe import Pipeline, PipelineConfig
pipeline = Pipeline(PipelineConfig())
result = pipeline.extract_url("https://example.com/article")
print(result.text)
返回字段(ExtractResult):
text:正文title:标题source:来源站点images:图片 URL 列表videos:视频 URL 列表used_render:是否使用渲染status_code:抓取到的 HTTP 状态码(抓取失败时可能为空)elapsed_ms:耗时(毫秒)fetch_ms:抓取耗时(毫秒)render_ms:渲染耗时(毫秒)extract_ms:正文抽取耗时(毫秒)image_ms:图片收集耗时(毫秒)video_ms:视频收集耗时(毫秒)error:错误信息(如有)
常见配置
代理与渲染
from trafipipe import Pipeline, PipelineConfig, ProxyConfig
cfg = PipelineConfig()
cfg.fetch.proxy = ProxyConfig(http="http://user:pass@host:port", https="http://user:pass@host:port")
cfg.render.proxy = ProxyConfig(server="http://user:pass@host:port")
cfg.render.extra_headers = {
"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
"Referer": "https://mp.weixin.qq.com/",
}
cfg.render.cookies = [
{"name": "your_cookie", "value": "xxx", "domain": ".mp.weixin.qq.com", "path": "/"}
]
cfg.render.reuse_context = True # 批量渲染时复用 context 以提速
pipeline = Pipeline(cfg)
result = pipeline.extract_url("https://example.com/article")
图片保留与输出格式
cfg = PipelineConfig()
cfg.extract.keep_images = True
cfg.extract.append_images = True # 在正文末尾追加 [Images] 列表
cfg.extract.inline_images = False # 设为 True 时输出 Markdown 并原位插入图片(对所有站点生效)
cfg.extract.keep_videos = True
cfg.extract.append_videos = False # 在正文末尾追加 [Videos] 列表
cfg.extract.inline_videos = False # 设为 True 时会在正文中插入 [Video] url
cfg.extract.output_format = "txt" # "txt" 或 "md"
说明:
- 当
inline_images=True且output_format="txt"时,会把转为[Image] url。 - 当
inline_videos=True时,会把 HTML 中的<video>/<source>转成[Video] url。
微信文章图片(mp.weixin.qq.com)
from trafipipe import Pipeline, PipelineConfig
cfg = PipelineConfig()
cfg.extract.keep_images = True
cfg.extract.append_images = True
cfg.render.mode = "auto" # 如图片仍缺失可改为 "always"
cfg.render.extra_headers = {"Referer": "https://mp.weixin.qq.com/"}
cfg.render.cookies = [
{"name": "your_cookie", "value": "xxx", "domain": ".mp.weixin.qq.com", "path": "/"}
]
result = Pipeline(cfg).extract_url("https://mp.weixin.qq.com/s/xxxxxx")
print(result.images)
列表页发现链接并抽取
from trafipipe import Pipeline, PipelineConfig
cfg = PipelineConfig()
cfg.crawl.max_pages = 50
cfg.crawl.max_depth = 2
cfg.crawl.max_workers = 4 # 并发抓取列表页
pipeline = Pipeline(cfg)
urls = pipeline.crawl(["https://example.com/list"])
results = pipeline.crawl_and_extract(urls, max_workers=4)
CLI
trafipipe extract https://example.com/article
trafipipe crawl https://example.com/list --max-pages 50 --max-depth 2 --workers 4
开发
pip install -e ".[dev]"
pytest
ruff check .
性能基准
python doc/benchmark.py --file doc/urls.txt --render auto --repeat 1
python doc/benchmark.py --file doc/urls.txt --format csv --summary > report.csv
python doc/benchmark.py --file doc/urls.txt --format json --summary > report.json
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
trafi_pipeline-0.1.5.tar.gz
(110.0 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file trafi_pipeline-0.1.5.tar.gz.
File metadata
- Download URL: trafi_pipeline-0.1.5.tar.gz
- Upload date:
- Size: 110.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
23d11483b4081627c272f68545f1fd50529cdec0efb194f0475fb435e90e3c63
|
|
| MD5 |
8e89648abf1e87c3580de3551e5d6882
|
|
| BLAKE2b-256 |
71fd3591e085158be375eee5f426b0ba2985466aca47e8d0c1aeb6b6de74dd4a
|
File details
Details for the file trafi_pipeline-0.1.5-py3-none-any.whl.
File metadata
- Download URL: trafi_pipeline-0.1.5-py3-none-any.whl
- Upload date:
- Size: 20.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a77b161ea062295c20378781727d95ea862ae7db3c6634f65f82ae6cc78f8583
|
|
| MD5 |
a80f2704ef6a4e31dfebe232f8a25a74
|
|
| BLAKE2b-256 |
6c1c1edbfcd36c2f43abe9933b1a20afe007208c693f80f5d641b85f8ea13c1e
|