Skip to main content

wagtail-machine-readable

Make your Wagtail site AI-ready: llms.txt, Markdown page variants and structured data in a single install.

PyPI CI Python License: MIT

LLMs, answer engines and AI crawlers are becoming the front door to your content, and they read llms.txt, not your carefully tuned templates. wagtail-machine-readable turns your existing Wagtail page tree into spec-compliant machine-readable outputs automatically. It serves live pages only, respects privacy restrictions, and gives every Site in a multi-site install its own hostname. Install it, add two lines, and /llms.txt works.

What you get

  • /llms.txt: a spec-compliant index of your site, with an H1 site name, a blockquote description, H2 sections derived from your top-level pages, and - [name](url): description link lists.
  • /llms-full.txt: the same page set with page content rendered to Markdown.
  • .md page variants: every visible page served as Markdown on its own URL (/about/team.md, /index.md for the site root), with content negotiation and rel="alternate" links so agents can find them.
  • StreamField to Markdown extraction: headings, links, images (alt text and URL), embeds and tables survive extraction instead of being stripped.
  • JSON-LD structured data: schema.org WebPage/Article, Organization and BreadcrumbList per page via a template tag, with a per-page-type builder mapping.
  • AI crawler robots.txt: opt-in allow and deny controls for known AI user agents (GPTBot, ClaudeBot, PerplexityBot and others) managed from settings.
  • AI crawler analytics: see which AI bots read your content in an admin report, recorded by an opt-in middleware.
  • Editor tools: a "Machine readable content" admin report and a page-editor preview mode that shows pages as machines see them.
  • Static export: a management command that writes every output to disk for static hosting and CDN workflows.
  • Wagtail-native visibility rules: only live pages, view-restricted (private) subtrees excluded, drafts excluded, multi-site aware, plus per-model and per-page opt-outs.

Quickstart

pip install wagtail-machine-readable
# settings.py
INSTALLED_APPS = [
    # ...
    "wagtail_machine_readable",
]
# urls.py, before Wagtail's catch-all page serving include
urlpatterns = [
    # ...
    path("", include("wagtail_machine_readable.urls")),
    path("", include(wagtail_urls)),
]

Zero configuration produces a valid llms.txt:

# Acme

> Acme makes modular widgets.

## Pages

- [Contact](https://example.com/contact/): Get in touch.

## About

- [About](https://example.com/about/): Who we are.
- [Team](https://example.com/about/team/): The people.

## Blog

- [Blog](https://example.com/blog/): News and articles.
- [First post](https://example.com/blog/first-post/): Our first post.
- [Second post](https://example.com/blog/second-post/)

Sections come from the children of each Site's root page, and descriptions come from search_description by default. Every Wagtail Site gets its own document on its own hostname.

Markdown page variants

Every visible page is also served as Markdown by appending .md to its slug path, so /about/team/ becomes /about/team.md, and the site root is /index.md. Responses use text/markdown and the same visibility rules as llms.txt, so drafts and private pages 404. Set MARKDOWN_ENABLED to False to turn the variants off.

Bodies are produced by MarkdownContentExtractor, which maps rich text and StreamField blocks to Markdown: headings (demoted below the page title), [text](url) links with rich-text page references expanded, ![alt](url) images, embeds and URL blocks as autolinks, and TableBlock values as Markdown tables.

Making the variants discoverable

Crawlers will not guess your URL convention, so advertise the variants. Add the middleware:

MIDDLEWARE = [
    # ...
    "wagtail_machine_readable.middleware.MachineReadableMiddleware",
]

This does two things for canonical page URLs:

  • Content negotiation: a request with Accept: text/markdown ranked above HTML gets the Markdown variant directly (with Vary: Accept and a Content-Location header). Browsers never send this, so regular visitors are unaffected.
  • Link headers: HTML page responses gain Link: <.../team.md>; rel="alternate"; type="text/markdown".

And in your page templates, advertise the variant in the head:

{% load machine_readable %}
{% markdown_alternate page %}

which renders <link rel="alternate" type="text/markdown" href="...">. The middleware resolves the page for each negotiated or HTML page response, costing one extra query. Skip it (and keep the template tag) if that matters at your scale.

Structured data (JSON-LD)

Add the template tag to your page template:

{% load machine_readable %}
{% structured_data page %}

This renders a <script type="application/ld+json"> element containing a schema.org @graph: the page node (WebPage by default), a BreadcrumbList from the page's ancestors and an Organization derived from the Wagtail Site. Map page types to other builders (Article ships ready to use) or to your own StructuredDataBuilder subclasses:

WAGTAIL_MACHINE_READABLE = {
    "STRUCTURED_DATA_BUILDERS": {
        "blog.BlogPage": "wagtail_machine_readable.structured_data.ArticleBuilder",
    },
}

Mappings match base classes too, so "wagtailcore.Page" changes the default for every page type.

AI crawler robots.txt

Opt in to serving /robots.txt with explicit rules for known AI user agents:

WAGTAIL_MACHINE_READABLE = {
    "ROBOTS_TXT_ENABLED": True,
    "ROBOTS_AI_DEFAULT": "allow",  # baseline for known AI agents
    "ROBOTS_AI_DENY": ["Bytespider"],  # per-agent overrides
    "ROBOTS_EXTRA": "Sitemap: https://example.com/sitemap.xml",
}

The built-in list covers GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot, Google-Extended, CCBot, Meta-ExternalAgent and others, and the allow and deny lists also accept agents not on the list. Other crawlers get User-agent: * / Allow: /. If your project already serves robots.txt, keep its URL pattern above the package include.

Editor tools

With wagtail.admin installed the package adds two reports (under Reports in the admin menu):

  • Machine readable content: every live page with its machine-readable status, covering whether a description is present, whether the extracted body is non-empty, and whether the page is excluded from outputs. Editors work through the "missing" and "empty" rows.
  • AI crawler activity: recent requests from known AI crawlers (see below).

For a "view as machine" preview inside the page editor, adopt the preview mixin (a plain class, no migration):

from wagtail_machine_readable.models import MachineReadablePreviewMixin


class ArticlePage(MachineReadablePreviewMixin, Page): ...

Editors get a "Machine readable" entry in the preview-mode dropdown showing the page exactly as its .md variant renders, so garbled extraction becomes visible while editing, not after publishing.

AI crawler analytics

Prove the machine-readable outputs are being read. Add the tracking middleware:

MIDDLEWARE = [
    # ...
    "wagtail_machine_readable.middleware.AICrawlerTrackingMiddleware",
]

Requests whose User-Agent matches a known AI crawler (the same list the robots.txt controls use, plus any agents in your allow and deny settings) are recorded with agent, host, path and status code, and surfaced in the AI crawler activity report with CSV and XLSX export. The hit model ships with its own migration, so run python manage.py migrate after installing or upgrading. Keep the table tidy from a cron job:

python manage.py prune_ai_crawler_hits --days 90

Static generation

python manage.py generate_machine_readable --output ./static-export
python manage.py generate_machine_readable --site example.com

Single-site projects write llms.txt, llms-full.txt, the .md page tree (index.md, about/team.md and so on) and, when enabled, robots.txt into the output directory. Multi-site projects get one subdirectory per hostname.

Caching

Responses carry Cache-Control: public, max-age=3600 by default (CACHE_MAX_AGE). For sites where generation itself is expensive, opt in to caching the generated documents server-side:

WAGTAIL_MACHINE_READABLE = {
    "GENERATION_CACHE_TIMEOUT": 86400,
}

Rendered documents are stored in Django's default cache and invalidated when a page is published or unpublished, so a long timeout is safe.

Settings

All configuration lives in a single dict. Every key has a working default, so configure only what you want to change:

WAGTAIL_MACHINE_READABLE = {
    "SITE_DESCRIPTION": "Acme makes modular widgets.",
    "MAX_PAGES_PER_SECTION": 25,
    "EXCLUDE_PAGE_MODELS": ["blog.BlogTagIndexPage"],
}
Key Default Purpose
SITE_DESCRIPTION None Blockquote description. A string, or a {hostname: description} dict for multi-site. Falls back to the root page's description chain.
SECTION_STRATEGY TopLevelSectionStrategy path Dotted path to the class deriving H2 sections from the page tree.
MAX_PAGES_PER_SECTION 50 Cap entries per section (None = unlimited).
RESPECT_SHOW_IN_MENUS False Only include pages with show_in_menus=True.
EXCLUDE_PAGE_MODELS [] Page models to exclude, as "app_label.ModelName" strings (exact type).
EXCLUDE_PAGE_IDS [] Specific page ids to exclude.
DESCRIPTION_FIELDS ["machine_readable_description", "search_description"] Per-page description fallback chain; first non-empty field wins.
FULL_TEXT_ENABLED True Serve/write llms-full.txt.
FULL_TEXT_MAX_PAGES None Cap the number of pages in llms-full.txt (None = unlimited).
MARKDOWN_ENABLED True Serve/write .md page variants.
EXTRACTOR MarkdownContentExtractor path Dotted path to the ContentExtractor used for page bodies.
CACHE_MAX_AGE 3600 Cache-Control: public, max-age=N on responses (0 = no header).
GENERATION_CACHE_TIMEOUT None Opt-in server-side caching of generated documents, in seconds.
STRUCTURED_DATA_BUILDERS {} Map "app_label.ModelName" to StructuredDataBuilder dotted paths.
ROBOTS_TXT_ENABLED False Serve/write robots.txt from this package.
ROBOTS_AI_DEFAULT "allow" Baseline policy ("allow"/"deny") for known AI user agents.
ROBOTS_AI_ALLOW [] Agents to explicitly allow, overriding the baseline.
ROBOTS_AI_DENY [] Agents to explicitly deny, overriding the baseline.
ROBOTS_EXTRA "" Raw text appended to robots.txt (for example a Sitemap: line).

Misconfigured keys are caught by Django system checks at startup.

Customisation

Exclude a page type with no mixin or migration needed:

class InternalToolPage(Page):
    exclude_from_machine_readable = True

Per-page editor controls: adopt the optional mixin for a dedicated AI-consumer description and a per-page exclusion flag, surfaced in the page editor by MachineReadablePanel:

from wagtail_machine_readable.models import MachineReadableMixin
from wagtail_machine_readable.panels import MachineReadablePanel


class ArticlePage(MachineReadableMixin, Page):
    promote_panels = Page.promote_panels + [MachineReadablePanel()]

The mixin adds fields to your page model, so run python manage.py makemigrations && python manage.py migrate after adopting it (and after upgrading from a version with fewer mixin fields). Until the migration is applied, any query against the page model, including the pages themselves, fails with a missing-column error.

Custom sections: subclass SectionStrategy and point SECTION_STRATEGY at it:

from wagtail_machine_readable.generators import SectionStrategy


class NavigationSections(SectionStrategy):
    def build_sections(self, site): ...

Custom content extraction: subclass ContentExtractor (or MarkdownContentExtractor, whose block handling and convert_html hook are overridable) and point EXTRACTOR at it. Set EXTRACTOR to "wagtail_machine_readable.extractors.DefaultContentExtractor" for the plain-text behaviour of v0.1.

Custom structured data: subclass StructuredDataBuilder (or WebPageBuilder/ArticleBuilder) and map page types to it via STRUCTURED_DATA_BUILDERS.

How it behaves

  • Visibility: a page appears only if it is live, has no view restriction on itself or an ancestor, and is not excluded by settings, the class attribute or the per-page flag. The private and draft rules match what anonymous visitors can already see, so nothing non-public leaks.
  • Multi-site: Site.find_for_request() scopes every request, so each hostname serves its own tree with absolute URLs.
  • Caching: responses carry Cache-Control: public, max-age=3600 by default. The views are cache_page-compatible (cache keys include the host), so wrapping them or enabling Django's cache middleware works per site out of the box. Server-side generation caching is opt-in via GENERATION_CACHE_TIMEOUT, invalidated on publish and unpublish.
  • Output safety: titles and descriptions are collapsed to single lines and escaped, so page content cannot inject sections or links into the document structure. JSON-LD payloads escape <, > and &, so page content cannot break out of the script element.

How is this different from django-llms-txt?

django-llms-txt is Django-generic: you describe your content to it. wagtail-machine-readable is Wagtail-native. It already understands the page tree (sections for free), live and privacy rules, multi-site scoping, search_description, and StreamField content. Point it at a Wagtail project and it produces the right documents with zero configuration.

Compatibility

Python 3.11 to 3.13. Django 4.2 and 5.2. Wagtail 6.3 to 7.x. The full matrix is tested in CI.

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

wagtail_machine_readable-0.5.0.tar.gz (46.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

wagtail_machine_readable-0.5.0-py3-none-any.whl (38.1 kB view details)

Uploaded Python 3

File details

Details for the file wagtail_machine_readable-0.5.0.tar.gz.

File metadata

  • Download URL: wagtail_machine_readable-0.5.0.tar.gz
  • Upload date:
  • Size: 46.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for wagtail_machine_readable-0.5.0.tar.gz
Algorithm Hash digest
SHA256 9cdefa5e7404130741af7876852c5b3c9691599839fb4b7a1727b4c1e5624bb6
MD5 91db91f5b5ea9cecf3fba05ebb2b7c9a
BLAKE2b-256 d2d124c511e8fa6b068eaa2f82aec47b22a9514cca9c4a0e3fbb43075ffc4591

See more details on using hashes here.

Provenance

The following attestation bundles were made for wagtail_machine_readable-0.5.0.tar.gz:

Publisher: publish.yml on brett-allard-amp/wagtail-machine-readable

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file wagtail_machine_readable-0.5.0-py3-none-any.whl.

File metadata

File hashes

Hashes for wagtail_machine_readable-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 03a2432ab603167d4859bd366823cecfaee8f3b3f704e66128bbb159e77fece2
MD5 7a408d7c0ebc3ab6433983d1f3f88ff8
BLAKE2b-256 48869f62208ca9a5ddba976a7663f3f69df0d30173c969179625e05a88b64b40

See more details on using hashes here.

Provenance

The following attestation bundles were made for wagtail_machine_readable-0.5.0-py3-none-any.whl:

Publisher: publish.yml on brett-allard-amp/wagtail-machine-readable

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 files

0.4.0

2 files

0.3.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page