Skip to main content

zytools-fs

zytools-fs is a small collection of Python utilities for personal crawling, FTP transfer, task heartbeats, persistent URL de-duplication, and Bilibili video downloads.

The package is intentionally lightweight: every helper can be imported and used directly in scripts without a framework.

Features

  • FTP recursive download and upload with retry support.
  • Bilibili video downloader based on requests.
  • LMDB-backed URL filter for large crawl de-duplication.
  • Article page detector and simple same-domain crawler.
  • Task heartbeat helper for reporting script status to a /tasks endpoint.
  • zytools command for a quick installation check.

Installation

pip install zytools-fs

Python 3.9 or newer is required.

For Bilibili DASH video merging, install ffmpeg and make sure it is available in PATH. If ffmpeg is not found, audio and video streams are kept as separate files.

Quick Check

zytools

Expected output:

zytools installed successfully.

FTP Transfer

from zytools.utils import FTPDownloader

with FTPDownloader(
    host="127.0.0.1",
    user="user",
    password="password",
) as ftp:
    ftp.download("/remote/path", "./downloads")
    ftp.upload("./reports", "/remote/reports")

Useful options:

  • port: FTP port, default 21.
  • encoding: FTP filename encoding, default utf-8.
  • passive: whether to use passive mode, default True.
  • download_retries and upload_retries: retry count.
  • show_progress: print transfer progress through loguru.

Bilibili Video Download

from zytools.video import download_bili_video

ok = download_bili_video(
    "https://www.bilibili.com/video/BVxxxx",
    output_dir="./downloads",
    quality="max",
    page="all",
    filename="Bilibili_{BV}_{Date}_{Page}_{PartTitle}",
    cookie={
        "SESSDATA": "your_sessdata",
        "bili_jct": "your_bili_jct",
    },
)

print(ok)

Parameters:

  • quality: "max" for the highest available stream, "min" for the lowest.
  • page: "all", a single page such as "1", or a range/list such as "1,3-5".
  • force: overwrite existing output files when set to True.
  • cookie: optional Bilibili cookies for videos that require login.

Filename template fields:

  • {Title}: video title.
  • {BV}: BV id.
  • {Date}: publish date in YYYYMMDD format.
  • {Page}: page number.
  • {Part}: same as page number.
  • {Duration}: duration in seconds.
  • {PartTitle}: page title.

Only download content that you own or are allowed to download.

URL Filter

UrlFilter stores compact MD5 fingerprints in LMDB. Unlike a Bloom filter, it does not intentionally produce false positives.

from zytools.utils import UrlFilter

with UrlFilter(file_path="url_seen.lmdb") as url_filter:
    url = "https://example.com/video?id=1"

    if url_filter.add(url):
        print("new url")
    else:
        print("seen before")

    print(len(url_filter))

Batch import and export:

from zytools.utils import UrlFilter

with UrlFilter("url_seen.lmdb") as url_filter:
    added = url_filter.add_many(
        [
            "https://example.com/a",
            "https://example.com/b",
        ]
    )
    url_filter.to_csv("url_seen.csv")

print(f"added {added} urls")

UrlFilter.to_lmdb("url_seen.csv", file_path="url_seen_copy.lmdb")

Article Detection

Use check_response to request one URL and classify it as an article, other HTML page, binary resource, or fetch error.

from zytools.artice import check_response

result = check_response("https://example.com/news/1.html")

if result["type"] == "article":
    print(result["title"])
    print(result["date"])
    print(result["text"][:300])
else:
    print(result["type"], result.get("reason"))

Return type values:

  • article: article page with extracted title, date, author, and text.
  • other: HTML page that does not look like an article.
  • binary: image, PDF, JavaScript, CSS, video, archive, or other non-HTML file.
  • fetch_error: request failed or returned a bad HTTP status.

Simple URL Crawler

UrlCrawler starts from one URL, follows links breadth-first, and yields article items. It can optionally use UrlFilter to avoid saving the same article URL across runs.

from zytools.artice import UrlCrawler
from zytools.utils import UrlFilter

with UrlFilter("article_urls.lmdb") as url_filter:
    crawler = UrlCrawler(
        start_url="https://example.com/",
        max_saved_urls=20,
        same_domain=True,
        max_depth=5,
        url_fp=url_filter,
    )

    for item in crawler.crawl():
        print(item["title"], item["url"])

    crawler.save_url_filter()

Each yielded item has:

  • title: extracted article title.
  • creat_date: extracted article publish date.
  • content: extracted article text.
  • url: final article URL.
  • get_date: crawl batch date.

Task Heartbeat

Use update_task for a single heartbeat request, or TaskUpdater when a script needs repeated updates with a minimum interval.

from zytools.utils import TaskUpdater, update_task

result = update_task(
    name="daily job",
    machine_id="machine-1",
    script_path="/path/to/script.py",
    server="http://127.0.0.1:8001",
)

print(result)

task = TaskUpdater(
    name="daily job",
    machine_id="machine-1",
    script_path="/path/to/script.py",
    server="http://127.0.0.1:8001",
    min_interval=60,
)

task.update()
task.update(force=True)

The server is expected to accept POST /tasks with a JSON body containing name, machine_id, script_path, enabled, and timeout_seconds.

Development

Build the package locally:

python -m build

Check the distribution metadata:

python -m twine check dist/*

Publish to PyPI:

python -m twine upload dist/*

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

zytools_fs-0.0.10.tar.gz (24.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

zytools_fs-0.0.10-py3-none-any.whl (24.3 kB view details)

Uploaded Python 3

File details

Details for the file zytools_fs-0.0.10.tar.gz.

File metadata

  • Download URL: zytools_fs-0.0.10.tar.gz
  • Upload date:
  • Size: 24.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for zytools_fs-0.0.10.tar.gz
Algorithm Hash digest
SHA256 a6fb647fc41de26fe3b83998f5fc7d447a2e24fba81a1ad9b17f267b34068a12
MD5 36c8e05b4de21a9c66676e876748f514
BLAKE2b-256 9aeb0bd364e3240526de9862b33962ab9bf4a0090ff1e2bf1d56f7a1fdc6170e

See more details on using hashes here.

File details

Details for the file zytools_fs-0.0.10-py3-none-any.whl.

File metadata

  • Download URL: zytools_fs-0.0.10-py3-none-any.whl
  • Upload date:
  • Size: 24.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for zytools_fs-0.0.10-py3-none-any.whl
Algorithm Hash digest
SHA256 ba327a191d71a31ad5834f6555d229b5d60aea9847505f892afa460df5ee303a
MD5 430c5fe47f7077161e8c53c5bf34e994
BLAKE2b-256 946c97097990983be175a3bf78208368ca104503b1102ec0771820ee50e50ede

See more details on using hashes here.

Release history Release notifications | RSS feed

0.0.43

2 files

0.0.42

1 file

0.0.41

1 file

0.0.40

1 file

0.0.39

1 file

0.0.38

1 file

0.0.37

1 file

0.0.36

2 files

0.0.35

2 files

0.0.34

2 files

0.0.33

2 files

0.0.32

2 files

0.0.31

2 files

0.0.30

2 files

0.0.29

2 files

0.0.28

2 files

0.0.27

2 files

0.0.26

2 files

0.0.25

2 files

0.0.24

2 files

0.0.23

2 files

0.0.22

2 files

0.0.21

2 files

0.0.20

2 files

0.0.19

2 files

0.0.17

2 files

0.0.16

2 files

0.0.14

2 files

0.0.13

2 files

0.0.12

2 files

This release

0.0.10 This release

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page