A pragmatic pipeline around trafilatura for JS-rendered pages and list discovery.
Project description
trafi-pipeline
基于 trafilatura 的“组合式”网页正文抽取管线:先抓取/渲染,再抽取;支持列表页发现详情页,再批量抽取。
特性
- 自动渲染策略:
auto | always | never,正文过短或抓取失败时可自动切换渲染 auto模式还会根据页面“展开/更多/阅读全文”等标记触发渲染,避免只抓到摘要- 列表页爬取:按深度与页数限制发现详情页链接
- 代理支持:抓取/渲染分别配置代理
- 图片处理:保留图片、追加图片列表、或原位插入(Markdown)
- 元数据:标题、来源站点、耗时等
- SVG 文本抽取:对 PDF/SVG 转换页面可直接提取文字
安装
pip install trafi-pipeline
建议完整安装(抓取 + 渲染能力都具备):
pip install "trafi-pipeline[http,render]"
python -m playwright install chromium
分开安装(更细粒度控制):
pip install "trafi-pipeline[http]" # 使用 httpx
pip install "trafi-pipeline[render]" # 使用 Playwright 进行渲染
说明:
- 默认配置
render.mode="auto",可能会触发渲染;若未安装render依赖或未安装浏览器,将导致结果为空或报错。 - 如果不需要渲染,请显式设置
render.mode="never",避免依赖缺失导致失败。
快速开始
from trafipipe import Pipeline, PipelineConfig
pipeline = Pipeline(PipelineConfig())
result = pipeline.extract_url("https://example.com/article")
print(result.text)
返回字段(ExtractResult):
text:正文title:标题source:来源站点images:图片 URL 列表videos:视频 URL 列表used_render:是否使用渲染status_code:抓取到的 HTTP 状态码(抓取失败时可能为空)elapsed_ms:耗时(毫秒)fetch_ms:抓取耗时(毫秒)render_ms:渲染耗时(毫秒)extract_ms:正文抽取耗时(毫秒)image_ms:图片收集耗时(毫秒)video_ms:视频收集耗时(毫秒)error:错误信息(如有)
说明:
- 如遇到验证码/人机验证页面,会直接返回
error="captcha_detected",避免误判为正文。
常见配置
代理与渲染
from trafipipe import Pipeline, PipelineConfig, ProxyConfig
cfg = PipelineConfig()
cfg.fetch.proxy = ProxyConfig(http="http://user:pass@host:port", https="http://user:pass@host:port")
cfg.render.proxy = ProxyConfig(server="http://user:pass@host:port")
cfg.render.extra_headers = {
"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
"Referer": "https://mp.weixin.qq.com/",
}
cfg.render.wait_selector = "article, .article, .article-content" # 支持多选择器
cfg.render.cookies = [
{"name": "your_cookie", "value": "xxx", "domain": ".mp.weixin.qq.com", "path": "/"}
]
cfg.render.reuse_context = True # 批量渲染时复用 context 以提速
pipeline = Pipeline(cfg)
result = pipeline.extract_url("https://example.com/article")
说明:
- 未设置
wait_selector时,会按域名与 HTML 结构自动挑选常见正文容器(多选择器)用于等待。
图片保留与输出格式
cfg = PipelineConfig()
cfg.extract.keep_images = True
cfg.extract.append_images = True # 在正文末尾追加图片列表(HTML 模式下为 <ul><img>)
cfg.extract.inline_images = False # 设为 True 时输出 Markdown 并原位插入图片(对所有站点生效)
cfg.extract.keep_videos = True
cfg.extract.append_videos = False # 在正文末尾追加视频列表(HTML 模式下为 <ul><video>)
cfg.extract.inline_videos = False # 设为 True 时会在正文中插入 [Video] url
cfg.extract.output_format = "txt" # "txt" / "md" / "html"
说明:
- 当
inline_images=True且output_format="txt"时,会把转为[Image] url。 - 当
output_format="html"时,将保留 HTML 并输出<img>标签(会自动修正懒加载 src)。 - HTML 输出会自动附带一份基础样式(居中排版、图片/视频自适应、表格样式等)。
- 当
inline_videos=True时,会把 HTML 中的<video>/<source>转成[Video] url;若output_format="html"则输出<video>标签。
微信文章图片(mp.weixin.qq.com)
from trafipipe import Pipeline, PipelineConfig
cfg = PipelineConfig()
cfg.extract.keep_images = True
cfg.extract.append_images = True
cfg.render.mode = "auto" # 如图片仍缺失可改为 "always"
cfg.render.extra_headers = {"Referer": "https://mp.weixin.qq.com/"}
cfg.render.cookies = [
{"name": "your_cookie", "value": "xxx", "domain": ".mp.weixin.qq.com", "path": "/"}
]
result = Pipeline(cfg).extract_url("https://mp.weixin.qq.com/s/xxxxxx")
print(result.images)
列表页发现链接并抽取
from trafipipe import Pipeline, PipelineConfig
cfg = PipelineConfig()
cfg.crawl.max_pages = 50
cfg.crawl.max_depth = 2
cfg.crawl.max_workers = 4 # 并发抓取列表页
pipeline = Pipeline(cfg)
urls = pipeline.crawl(["https://example.com/list"])
results = pipeline.crawl_and_extract(urls, max_workers=4)
CLI
trafipipe extract https://example.com/article
trafipipe crawl https://example.com/list --max-pages 50 --max-depth 2 --workers 4
开发
pip install -e ".[dev]"
pytest
ruff check .
性能基准
python doc/benchmark.py --file doc/urls.txt --render auto --repeat 1
python doc/benchmark.py --file doc/urls.txt --format csv --summary > report.csv
python doc/benchmark.py --file doc/urls.txt --format json --summary > report.json
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
trafi_pipeline-0.1.7.tar.gz
(431.1 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file trafi_pipeline-0.1.7.tar.gz.
File metadata
- Download URL: trafi_pipeline-0.1.7.tar.gz
- Upload date:
- Size: 431.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a52b709c412ade5a1b9c83cd9b72364d013ea4e88d3f8440a38d0e3c69417e97
|
|
| MD5 |
8f4947c8b04dfe35c45bb30370e3249e
|
|
| BLAKE2b-256 |
c18410c836bb47b7feb6cf3d4deb63cbfd131ed587491173c483e42ecb664a25
|
File details
Details for the file trafi_pipeline-0.1.7-py3-none-any.whl.
File metadata
- Download URL: trafi_pipeline-0.1.7-py3-none-any.whl
- Upload date:
- Size: 24.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b330085cd44757f3684ecd1a715917382852452a5979fbb3160bc1d98fba5b34
|
|
| MD5 |
935be8aa3509af8baf0405a864040d23
|
|
| BLAKE2b-256 |
3232ae3ac219e287ef95e71fa3551e2c9f31328b65043f49fd4fd9b3a82c9eb2
|