Skip to main content

markdfetch

A lightweight Python library for fetching web pages and extracting content as Markdown, plain text, or structured links.

Features

  • Fetch web pages with a simple API
  • Convert HTML to Markdown
  • Extract plain text from web pages
  • Extract links with URL and anchor text
  • Exclude unwanted HTML tags before processing
  • Include only specific HTML tags before processing
  • Support for custom request headers and timeouts
  • Automatic resolution of relative URLs
  • CSS selector support
  • Optional link deduplication
  • Automatic retry handling

Installation

pip install markdfetch

Quick Start

import markdfetch

page = markdfetch.fetch("https://example.com")

print(page.markdown())

Fetch a Page

import markdfetch

page = markdfetch.fetch("https://example.com")

print(page.status_code)
print(page.url)

Convert HTML to Markdown

page = markdfetch.fetch("https://example.com")

markdown = page.markdown()

print(markdown)

Exclude HTML Tags

Remove unwanted sections before converting to Markdown.

page = markdfetch.fetch("https://example.com")

markdown = page.markdown(
    exclude=["nav", "footer"]
)

print(markdown)

Include Specific HTML Tags

Extract content only from selected tags.

page = markdfetch.fetch("https://example.com")

markdown = page.markdown(
    include=["article"]
)

print(markdown)

Combine Include and Exclude

page = markdfetch.fetch("https://example.com")

markdown = page.markdown(
    include=["article"],
    exclude=["nav", "footer"]
)

print(markdown)

Extract Plain Text

page = markdfetch.fetch("https://example.com")

text = page.text()

print(text)

Extract Links

page = markdfetch.fetch("https://example.com")

links = page.links()

print(links)

Example output:

[
    {
        "url": "https://example.com/about",
        "text": "About Us"
    },
    {
        "url": "https://example.com/contact",
        "text": "Contact"
    }
]

Skip Empty Links

page = markdfetch.fetch("https://example.com")

links = page.links(skip_empty=True)

Extract Content Using CSS Selectors

Target specific elements using CSS selectors.

page = markdfetch.fetch("https://example.com")

markdown = page.markdown(
    selector="article"
)

print(markdown)

You can use any valid CSS selector:

page.markdown(selector=".content")
page.markdown(selector="#main")
page.markdown(selector="article.post")

Extract Text Using CSS Selectors

Extract plain text from specific sections of a page.

page = markdfetch.fetch("https://example.com")

text = page.text(
    selector=".content"
)

print(text)

Extract Unique Links

Remove duplicate URLs from the extracted links.

page = markdfetch.fetch("https://example.com")

links = page.links(
    unique=True
)

print(links)

Roadmap

Planned features:

  • Async support via httpx
  • Proxy support
  • Metadata extraction

License

MIT License

Release files for markdfetch 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for markdfetch 0.1.0
File Size Uploaded
markdfetch-0.1.0.tar.gz 4.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for markdfetch 0.1.0
File Interpreter ABI Platform
markdfetch-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 8.5 kB

Release files / markdfetch-0.1.0.tar.gz

Download URL markdfetch-0.1.0.tar.gz
Size 4.0 kB
Tags Source
SHA-256 checksum
How to use checksums
cdc84afe23d55973656d266fb6dfca591eac49c64f912151d18e3e60f07225c8
BLAKE2b-256 checksum
How to use checksums
f651aa0e05fbf4ecbf41bf337756e01fa276f66d39a709e45df57ede89b903e5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release files / markdfetch-0.1.0-py3-none-any.whl

Download URL markdfetch-0.1.0-py3-none-any.whl
Size 4.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
959f003027bb7a41205679724899cedbd296696c490d962aef00800a2b453f22
BLAKE2b-256 checksum
How to use checksums
ad13ae1a02d60cbf0311cc99b7c1a91bc8341040a521681924ecca5f387909bd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page