zytools-fs
zytools-fs is a small collection of Python utilities for personal crawling,
FTP transfer, task heartbeats, persistent URL de-duplication, and Bilibili video
downloads.
The package is intentionally lightweight: every helper can be imported and used directly in scripts without a framework.
Features
- FTP recursive download and upload with retry support.
- Bilibili video downloader based on
requests. - LMDB-backed URL filter for large crawl de-duplication.
- Article page detector and simple same-domain crawler.
- Task heartbeat helper for reporting script status to a
/tasksendpoint. zytoolscommand for a quick installation check.
Installation
pip install zytools-fs
Python 3.9 or newer is required.
For Bilibili DASH video merging, install ffmpeg and make sure it is available
in PATH. If ffmpeg is not found, audio and video streams are kept as
separate files.
Quick Check
zytools
Expected output:
zytools installed successfully.
FTP Transfer
from zytools.utils import FTPDownloader
with FTPDownloader(
host="127.0.0.1",
user="user",
password="password",
) as ftp:
ftp.download("/remote/path", "./downloads")
ftp.upload("./reports", "/remote/reports")
Useful options:
port: FTP port, default21.encoding: FTP filename encoding, defaultutf-8.passive: whether to use passive mode, defaultTrue.download_retriesandupload_retries: retry count.show_progress: print transfer progress throughloguru.
Bilibili Video Download
from zytools.video import download_bili_video
ok = download_bili_video(
"https://www.bilibili.com/video/BVxxxx",
output_dir="./downloads",
quality="max",
page="all",
filename="Bilibili_{BV}_{Date}_{Page}_{PartTitle}",
cookie={
"SESSDATA": "your_sessdata",
"bili_jct": "your_bili_jct",
},
)
print(ok)
Parameters:
quality:"max"for the highest available stream,"min"for the lowest.page:"all", a single page such as"1", or a range/list such as"1,3-5".force: overwrite existing output files when set toTrue.cookie: optional Bilibili cookies for videos that require login.
Filename template fields:
{Title}: video title.{BV}: BV id.{Date}: publish date inYYYYMMDDformat.{Page}: page number.{Part}: same as page number.{Duration}: duration in seconds.{PartTitle}: page title.
Only download content that you own or are allowed to download.
URL Filter
UrlFilter stores compact MD5 fingerprints in LMDB. Unlike a Bloom filter, it
does not intentionally produce false positives.
from zytools.utils import UrlFilter
with UrlFilter(file_path="url_seen.lmdb") as url_filter:
url = "https://example.com/video?id=1"
if url_filter.add(url):
print("new url")
else:
print("seen before")
print(len(url_filter))
Batch import and export:
from zytools.utils import UrlFilter
with UrlFilter("url_seen.lmdb") as url_filter:
added = url_filter.add_many(
[
"https://example.com/a",
"https://example.com/b",
]
)
url_filter.to_csv("url_seen.csv")
print(f"added {added} urls")
UrlFilter.to_lmdb("url_seen.csv", file_path="url_seen_copy.lmdb")
Article Detection
Use check_response to request one URL and classify it as an article, other
HTML page, binary resource, or fetch error.
from zytools.artice import check_response
result = check_response("https://example.com/news/1.html")
if result["type"] == "article":
print(result["title"])
print(result["date"])
print(result["text"][:300])
else:
print(result["type"], result.get("reason"))
Return type values:
article: article page with extractedtitle,date,author, andtext.other: HTML page that does not look like an article.binary: image, PDF, JavaScript, CSS, video, archive, or other non-HTML file.fetch_error: request failed or returned a bad HTTP status.
Simple URL Crawler
UrlCrawler starts from one URL, follows links breadth-first, and yields article
items. It can optionally use UrlFilter to avoid saving the same article URL
across runs.
from zytools.artice import UrlCrawler
from zytools.utils import UrlFilter
with UrlFilter("article_urls.lmdb") as url_filter:
crawler = UrlCrawler(
start_url="https://example.com/",
max_saved_urls=20,
same_domain=True,
max_depth=5,
url_fp=url_filter,
)
for item in crawler.crawl():
print(item["title"], item["url"])
crawler.save_url_filter()
Each yielded item has:
title: extracted article title.creat_date: extracted article publish date.content: extracted article text.url: final article URL.get_date: crawl batch date.
Task Heartbeat
Use update_task for a single heartbeat request, or TaskUpdater when a script
needs repeated updates with a minimum interval.
from zytools.utils import TaskUpdater, update_task
result = update_task(
name="daily job",
machine_id="machine-1",
script_path="/path/to/script.py",
server="http://127.0.0.1:8001",
)
print(result)
task = TaskUpdater(
name="daily job",
machine_id="machine-1",
script_path="/path/to/script.py",
server="http://127.0.0.1:8001",
min_interval=60,
)
task.update()
task.update(force=True)
The server is expected to accept POST /tasks with a JSON body containing
name, machine_id, script_path, enabled, and timeout_seconds.
Development
Build the package locally:
python -m build
Check the distribution metadata:
python -m twine check dist/*
Publish to PyPI:
python -m twine upload dist/*
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file zytools_fs-0.0.10.tar.gz.
File metadata
- Download URL: zytools_fs-0.0.10.tar.gz
- Upload date:
- Size: 24.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a6fb647fc41de26fe3b83998f5fc7d447a2e24fba81a1ad9b17f267b34068a12
|
|
| MD5 |
36c8e05b4de21a9c66676e876748f514
|
|
| BLAKE2b-256 |
9aeb0bd364e3240526de9862b33962ab9bf4a0090ff1e2bf1d56f7a1fdc6170e
|
File details
Details for the file zytools_fs-0.0.10-py3-none-any.whl.
File metadata
- Download URL: zytools_fs-0.0.10-py3-none-any.whl
- Upload date:
- Size: 24.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ba327a191d71a31ad5834f6555d229b5d60aea9847505f892afa460df5ee303a
|
|
| MD5 |
430c5fe47f7077161e8c53c5bf34e994
|
|
| BLAKE2b-256 |
946c97097990983be175a3bf78208368ca104503b1102ec0771820ee50e50ede
|