7 projects
domdown
extracts the main content from web pages and returns cleaned HTML, optional markdown, and structured metadata.
crawlsmith
Crawlsmith helps you craft reliable web crawlers in Python, combining page fetching, HTML parsing, link discovery, and content extraction into a simple and extensible toolkit.
feedtrail
Feed Tracking and Retrieval Abstraction Interface Layer
text2ioc
Extract Indicators of Compromise (IOCs) from unstructured text.
bjorn
It loads different configurations depending on the environment variable.
ragnar
lightweight Extract-Transform-Load (ETL) framework for Python 3+
lagertha
An application that allows task scheduling