Convert web pages to clean Markdown format
Project description
blog2md
Convert web pages to clean Markdown format.
English
Use Cases
| Scenario | Command | Description |
|---|---|---|
| Save blog articles offline | blog2md <url> |
Fetch blog and convert to Markdown |
| Backup technical docs | blog2md <doc-url> --force-js |
Render JS then fetch |
| Batch process websites | blog2md --batch urls.txt |
Concurrent URL processing |
| Subscribe to blog updates | blog2md --rss <feed-url> |
Auto-discover articles from RSS |
| Build knowledge base | blog2md <url> --crawl |
Recursively crawl links |
| Extract video subtitles | blog2md <url> |
Get page text (不含视频) |
| Migrate to Obsidian | blog2md <url> --target obsidian --vault-path /path/to/vault |
Write directly to Obsidian vault |
Installation
# Basic (CLI + SDK)
pip install blog2md
# With JavaScript rendering (for SPA/React/Vue sites)
pip install blog2md[js]
playwright install chromium
# From source
git clone https://github.com/nestedcat/blog2md.git
cd blog2md
pip install -e ".[js]"
CLI Examples
# 1. Basic extraction (most common)
blog2md https://example.com/article
# 2. Specify output directory
blog2md https://example.com/article --output ./docs
# 3. Output to stdout
blog2md https://example.com/article --stdout
# 4. Keep original image URLs (no download)
blog2md https://example.com/article --images keep
# 5. Skip YAML frontmatter
blog2md https://example.com/article --no-frontmatter
# 6. Batch process URL list
blog2md --batch urls.txt --concurrency 5
# 7. RSS subscription
blog2md --rss https://example.com/feed.xml --output ./articles
# 8. JavaScript rendering (for SPA/React/Vue)
blog2md https://example.com/spa-page --force-js
# 9. Crawl with depth 1 (same domain)
blog2md https://example.com --crawl --crawl-depth 1
# 10. Crawl cross-domain with depth 2
blog2md https://example.com --crawl --crawl-depth 2 --crawl-cross-domain
# 11. Crawl with delay (reduce server load)
blog2md https://example.com --crawl --crawl-depth 2 --crawl-delay 2.0
# 12. Force refresh (ignore cache)
blog2md https://example.com/article --force
# 13. Write to GitHub repo
blog2md https://example.com/article \
--target github \
--github-token ghp_xxx \
--github-repo owner/repo \
--github-path content/posts/
# 14. Write to Obsidian vault
blog2md https://example.com/article \
--target obsidian \
--vault-path /path/to/your/vault
Python API
import blog2md
# One-liner to extract
result = blog2md.extract("https://example.com/article")
print(result.markdown)
print(result.title) # Title from frontmatter
print(result.author) # Author
print(result.date) # Publication date
# Save to file
result.save("/path/to/file.md")
CLI Options
| Option | Description |
|---|---|
-o, --output |
Output directory |
--batch <file> |
Batch file (one URL per line or JSON) |
--rss <url> |
RSS/Atom feed URL |
--concurrency N |
Concurrent requests (1-10, default 3) |
-f, --filename |
Custom output filename |
--stdout |
Output to stdout |
--no-frontmatter |
Skip YAML frontmatter |
--force-js |
Force JavaScript rendering |
--images |
Image mode: local/keep/inline/skip (详细说明见下) |
--images-dir |
Image directory (default _images) |
--crawl |
Recursively follow links |
--crawl-depth N |
Max crawl depth (default 1) |
--crawl-cross-domain |
Allow cross-domain crawling |
--crawl-delay N |
Delay between requests in seconds (default 1.0) |
--target |
Output target: file/github/obsidian |
--cache/--no-cache |
Enable/disable cache (default enabled) |
--force |
Force reprocess ignoring cache |
-v, --verbose |
Verbose logging |
Image Modes Explained
| Mode | Description | Use Case |
|---|---|---|
local |
Download images to _images/ folder, convert URLs to relative paths |
Default - Best for offline reading |
keep |
Keep original image URLs unchanged | When images are already hosted reliably |
inline |
(Not fully implemented) Convert to base64 inline | For single-file portability |
skip |
Replace images with placeholder [image] |
When you only want text content |
Features
- trafilatura: High-precision content extraction, removes ads/nav
- BeautifulSoup: HTML parsing and content detection
- Format preservation: Headings, code blocks, tables, bold/italic
- Image handling: Download to local
_images/, convert URLs to relative - Sidebar cleaning: Auto-detect and remove nav elements
- Caching: Content hash detection, ETag/Last-Modified support
中文
使用场景
| 场景 | 命令 | 说明 |
|---|---|---|
| 保存博客文章离线阅读 | blog2md <url> |
抓取博客,转为 Markdown |
| 备份技术文档 | blog2md <doc-url> --force-js |
渲染 JS 后抓取 |
| 批量抓取网站 | blog2md --batch urls.txt |
并发处理多个 URL |
| 订阅博客更新 | blog2md --rss <feed-url> |
从 RSS 自动发现文章 |
| 构建知识库 | blog2md <url> --crawl |
循环抓取链接建立知识库 |
| 提取视频字幕 | blog2md <url> |
获取页面文字(不含视频) |
| 迁移到 Obsidian | blog2md <url> --target obsidian --vault-path /path/to/vault |
直接写入 Obsidian 库 |
安装
# 基本安装 (CLI + SDK)
pip install blog2md
# JavaScript 渲染支持 (适用于 SPA/React/Vue)
pip install blog2md[js]
playwright install chromium
# 从源码安装
git clone https://github.com/nestedcat/blog2md.git
cd blog2md
pip install -e ".[js]"
CLI 示例
# 1. 基本抓取(最常用)
blog2md https://example.com/article
# 2. 指定输出目录
blog2md https://example.com/article --output ./docs
# 3. 输出到标准输出
blog2md https://example.com/article --stdout
# 4. 保留原始图片 URL(不下载)
blog2md https://example.com/article --images keep
# 5. 跳过 YAML frontmatter
blog2md https://example.com/article --no-frontmatter
# 6. 批量处理 URL 列表
blog2md --batch urls.txt --concurrency 5
# 7. RSS 订阅
blog2md --rss https://example.com/feed.xml --output ./articles
# 8. JavaScript 渲染(适合 React/Vue/SPA)
blog2md https://example.com/spa-page --force-js
# 9. 循环抓取(深度 1,站内链接)
blog2md https://example.com --crawl --crawl-depth 1
# 10. 循环抓取(跨域,深度 2)
blog2md https://example.com --crawl --crawl-depth 2 --crawl-cross-domain
# 11. 循环抓取(带延迟,减少对服务器的压力)
blog2md https://example.com --crawl --crawl-depth 2 --crawl-delay 2.0
# 12. 强制刷新(忽略缓存)
blog2md https://example.com/article --force
# 13. 写入 GitHub 仓库
blog2md https://example.com/article \
--target github \
--github-token ghp_xxx \
--github-repo owner/repo \
--github-path content/posts/
# 14. 写入 Obsidian 库
blog2md https://example.com/article \
--target obsidian \
--vault-path /path/to/your/vault
Python API
import blog2md
# 一行代码抓取
result = blog2md.extract("https://example.com/article")
print(result.markdown)
print(result.title) # frontmatter 中的标题
print(result.author) # 作者
print(result.date) # 发布日期
# 保存到文件
result.save("/path/to/file.md")
CLI 选项
| 选项 | 说明 |
|---|---|
-o, --output |
输出目录 |
--batch <file> |
批量处理文件(每行一个 URL 或 JSON) |
--rss <url> |
RSS/Atom 订阅地址 |
--concurrency N |
并发数(1-10,默认 3) |
-f, --filename |
自定义输出文件名 |
--stdout |
输出到标准输出 |
--no-frontmatter |
跳过 YAML 头 |
--force-js |
强制 JavaScript 渲染 |
--images |
图片模式:local/keep/inline/skip |
--crawl |
循环抓取页面内链接 |
--crawl-depth N |
抓取深度(默认 1) |
--crawl-cross-domain |
允许跨域抓取 |
--crawl-delay N |
请求间隔(秒,默认 1.0) |
--target |
输出目标:file/github/obsidian |
--cache/--no-cache |
启用/禁用缓存(默认启用) |
--force |
强制重新抓取(忽略缓存) |
-v, --verbose |
详细日志 |
图片模式说明
| 模式 | 说明 | 适用场景 |
|---|---|---|
local |
下载图片到 _images/ 文件夹,URL 转为相对路径 |
默认 - 离线阅读最佳 |
keep |
保留原始图片 URL | 图片已可靠托管时使用 |
inline |
(未完全实现)转为 base64 内联 | 单文件便携性 |
skip |
用 [image] 占位符替换图片 |
只需文本内容时 |
功能特点
- trafilatura: 高精度正文提取,去除广告/导航
- BeautifulSoup: HTML 解析与内容识别
- 格式保留: 标题层级、代码块、表格、加粗斜体
- 图片处理: 下载到本地
_images/,自动转换 URL - 侧边栏清理: 自动识别并移除导航元素
- 缓存机制: 内容哈希检测变化,支持 ETag/Last-Modified
License
MIT License - see LICENSE for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
blog2md-1.0.3.tar.gz
(71.3 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
blog2md-1.0.3-py3-none-any.whl
(86.6 kB
view details)
File details
Details for the file blog2md-1.0.3.tar.gz.
File metadata
- Download URL: blog2md-1.0.3.tar.gz
- Upload date:
- Size: 71.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5a3bcd103c522ecfaf1bb5f3e591c2ad3be244f72e0c05120428bd37ba546731
|
|
| MD5 |
c539f2fcb3c98ae3765675f1ebb6151f
|
|
| BLAKE2b-256 |
c507b6abf3fc3b313158bbf3abd706c573159404f905e61c676f58edb55ede38
|
File details
Details for the file blog2md-1.0.3-py3-none-any.whl.
File metadata
- Download URL: blog2md-1.0.3-py3-none-any.whl
- Upload date:
- Size: 86.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f748cc267516d44e6b20efb05f76db863e008711f229ff3015bb395b808749a2
|
|
| MD5 |
f6d4a183309a63c01152cf3c78dbe6fc
|
|
| BLAKE2b-256 |
b9061d49ff7d50d2a5fe05d145da0d72d1756dfc5af800564b22fb0c1f1d9190
|