Skip to main content

mkdocs-htmlproofer-plugin PyPI - Version

GitHub Actions

A MkDocs plugin that validates URLs, including anchors, in rendered html files.

Installation

  1. Prerequisites
  • Python >= 3.10
  • MkDocs >= 1.4.0
  1. Install the package with pip:
pip install mkdocs-htmlproofer-plugin
  1. Enable the plugin in your mkdocs.yml:
plugins:
    - search
    - htmlproofer

Configuring

enabled

True by default, allows toggling whether the plugin is enabled. Useful for local development where you may want faster build times.

plugins:
  - htmlproofer:
      enabled: !ENV [ENABLED_HTMLPROOFER, True]

Which enables you to disable the plugin locally using:

export ENABLED_HTMLPROOFER=false
mkdocs serve

raise_error

Optionally, you may raise an error and fail the build on first bad url status. Takes precedence over raise_error_after_finish.

plugins:
  - htmlproofer:
      raise_error: True

raise_error_after_finish

Optionally, you may want to raise an error and fail the build on at least one bad url status after all links have been checked.

plugins:
  - htmlproofer:
      raise_error_after_finish: True

raise_error_excludes

When specifying raise_error: True or raise_error_after_finish: True, it is possible to ignore errors for combinations of URLs and status codes with raise_error_excludes. Each URL supports unix style wildcards *, [], ?, etc.

plugins:
  - search
  - htmlproofer:
      raise_error: True
      raise_error_excludes:
        504: ['https://www.mkdocs.org/']
        404: ['https://github.com/manuzhang/*']
        400: ['*']
        -1: ['https://flaky.example.com/*']

A URL which couldn't be requested at all, such as one whose host wouldn't resolve or which redirected too many times, is reported as -1 and excluded under that status. A request which times out is reported as 504.

ignore_urls

Avoid validating the given list of URLs by ignoring them altogether. Each URL in the list supports unix style wildcards *, [], ?, etc.

Unlike raise_error_excludes, ignored URLs will not be fetched at all.

plugins:
  - search
  - htmlproofer:
      raise_error: True
      ignore_urls:
        - https://github.com/myprivateorg/*
        - https://app.dynamic-service-of-some-kind.io*

warn_on_ignored_urls

Log a warning when ignoring URLs with ignore_urls option. Defaults to false (no warning).

plugins:
  - search
  - htmlproofer:
      raise_error: True
      ignore_urls:
        - https://github.com/myprivateorg/*
        - https://app.dynamic-service-of-some-kind.io*
      warn_on_ignored_urls: true

ignore_pages

Avoid validating the URLs on the given list of markdown pages by ignoring them altogether. Each page in the list supports unix style wildcards *, [], ?, etc.

plugins:
  - search
  - htmlproofer:
      raise_error: True
      ignore_pages:
        - path/to/file.md
        - path/to/folder/*

validate_external_urls

Avoids validating any external URLs (i.e those starting with http:// or https://). This will be faster if you just want to validate local anchors, as it does not make any network requests.

plugins:
  - htmlproofer:
      validate_external_urls: False

validate_rendered_template

Validates the entire rendered template for each page - including the navigation, header, footer, etc. This defaults to off because it is much slower and often redundant to repeat for every single page.

plugins:
  - htmlproofer:
      validate_rendered_template: True

strict_anchors

Off by default, when an anchor is accepted if either the rendered page contains it or its Markdown source provides it. Turn it on to accept only the anchors a page renders, reporting a link to one which doesn't exist, such as #heading where attr_list replaced it in ## Heading {#custom-id}, or an anchor appearing only inside a fenced code block.

It decides links written with a path to a Markdown page, page.md#anchor, including one back into the page holding it. An anchor on a page which isn't Markdown isn't checked at all. A bare #anchor reads the same either way, resolved against every id the page renders, the theme's included.

For a page.md whose heading is written as ## Renamed Heading { #custom-id }:

  • page.md#custom-id is accepted whether the option is on or off.
  • page.md#renamed-heading is accepted by default, and reported with the option on.
  • #renamed-heading, written on page.md itself, is reported either way, the page rendering no such id.
plugins:
  - htmlproofer:
      strict_anchors: True

skip_downloads

Optionally skip downloading of a remote URLs content via GET request. This can considerably reduce the time taken to validate URLs.

plugins:
  - htmlproofer:
      skip_downloads: True

retry_max_times

Sets the maximum number of HTTP request retries when checking an external URL. Defaults to 0 (no retries). Retries back off exponentially, starting at 2 seconds. Local links and anchors are not retried.

plugins:
  - htmlproofer:
      retry_max_times: 3

max_workers

Optionally set the maximum number of worker threads used to validate URLs concurrently. By default, this is not set and the default of Python's ThreadPoolExecutor is used.

plugins:
  - htmlproofer:
      max_workers: 16

user_agent

The User-Agent to send when requesting an external URL. A browser's by default, because sites and the CDNs in front of them increasingly answer anything else with a 403, which is reported as a broken link although the page opens in a browser.

Set it to identify your build instead:

plugins:
  - htmlproofer:
      user_agent: 'Bot (https://example.com/)'

Compatibility with attr_list extension

If you need to manually specify anchors make use of the attr_list extension in the markdown. This can be useful for multilingual documentation to keep anchors as language neutral permalinks in all languages.

  • A sample for a heading # Grüße {#greetings} (the slugified generated anchor grue is overwritten with greetings).
  • This also works for images this is a nice image ![](foo-bar.png){#nice-image}
  • And generally for paragraphs:
Listing: This is noteworthy.
{#paragraphanchor}

Improving

More information about plugins in the MkDocs documentation

Acknowledgement

This work is based on the mkdocs-markdownextradata-plugin project and the Finding and Fixing Website Link Rot with Python, BeautifulSoup and Requests article.

Metadata

Release files for mkdocs-htmlproofer-plugin 1.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mkdocs-htmlproofer-plugin 1.6.0
File Size Uploaded
mkdocs_htmlproofer_plugin-1.6.0.tar.gz 15.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mkdocs-htmlproofer-plugin 1.6.0
File Interpreter ABI Platform
mkdocs_htmlproofer_plugin-1.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 27.6 kB

Release files / mkdocs_htmlproofer_plugin-1.6.0.tar.gz

Download URL mkdocs_htmlproofer_plugin-1.6.0.tar.gz
Size 15.0 kB
Tags Source
SHA-256 checksum
How to use checksums
7991d835b2cf0798cbaf49d876164c67314051ceaaa5bd71a089bb6a33e4e0b5
BLAKE2b-256 checksum
How to use checksums
1f6a33126e357d06e082877f770a405cf44374f3a6501a359562004da866b739
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.5

Release files / mkdocs_htmlproofer_plugin-1.6.0-py3-none-any.whl

Download URL mkdocs_htmlproofer_plugin-1.6.0-py3-none-any.whl
Size 12.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
697eef6c44ff2d3dd61f2adb1c39bb3cb687ff38e8dd9f7d430cca638b5a6c26
BLAKE2b-256 checksum
How to use checksums
f3ceb97c19d190f2da00ca931121a4df81b1ed1057f6f48a22b67833a1627768
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.5
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page