A high-performance, pure-Python AsciiDoc parser based on Lark.
Project description
= AsciiDoctrine
Michael R. Bernstein <zopemaven@gmail.com>
v0.1.0
:toc: left
:icons: font
:idprefix:
:idseparator: -
:sectanchors:
:sectlinks:
:source-highlighter: highlight.js
image:https://img.shields.io/badge/License-Apache_2.0-blue.svg[License, link=https://opensource.org/licenses/Apache-2.0]
image:https://img.shields.io/badge/python-3.9+-blue.svg[Python Version]
image:https://img.shields.io/badge/status-alpha-orange.svg[Development Status]
image:https://img.shields.io/badge/docs-GitHub_Pages-blue[Documentation, link=https://webmaven.github.io/asciidoctrine/]
A high-performance, pure-Python AsciiDoc parser built with Lark, designed for compatibility with the official AsciiDoc specification and the TCK.
The complete, compiled user documentation is available online at link:https://webmaven.github.io/asciidoctrine/[https://webmaven.github.io/asciidoctrine/].
== Motivation
The Python ecosystem has long lacked a modern, maintainable, and specification-compliant AsciiDoc parser. Existing tools are often port-based or rely on regex-heavy implementations that struggle with the complex, context-sensitive nature of AsciiDoc.
AsciiDoctrine is built from the ground up to provide:
1. **Spec Alignment**: Strict adherence to the upcoming official AsciiDoc Language Specification.
2. **First-Class AST**: A structured, type-safe Abstract Syntax Tree that makes building renderers and static analysis tools a breeze.
3. **Performance**: Leveraging the Lark parsing engine for efficient processing of large documents.
== Architecture
The parser operates in a multi-pass pipeline to handle the inherent complexity of AsciiDoc:
[source,mermaid]
....
graph LR
A[Source] --> B(Lark Parser)
B --> C[Concrete Syntax Tree]
C --> D(Transformer)
D --> E[Structured AST]
E --> F(Semantic Passes)
F --> G[Resolved ASG]
....
* **AST (Abstract Syntax Tree)**: Represented in `nodes.py`, this is a structural tree of the document elements.
* **ASG (Abstract Semantic Graph)**: The final resolved state where attributes, cross-references, and includes are fully processed.
== Technical Choices
* **Lark Parsing Engine**: We use Lark because it supports multiple parsing algorithms (Earley, LALR) and has an experimental PEG mode. This allow us to handle the context-sensitive nature of AsciiDoc without the maintenance nightmare of large regex collections.
* **Two-Pass Pipeline**: Handling attributes and includes requires knowing the state of the whole document. Our multi-pass approach ensures that we can resolve semantic details (like cross-references) correctly.
* **Pure Python**: Zero C-extensions means easy installation on all platforms including Pyodide and WebAssembly.
== Installation & Setup
For developers and contributors, the recommended "happy path" is to set up a local virtual environment and install AsciiDoctrine in editable development mode:
[source,bash]
----
# 1. Clone the repository
git clone https://github.com/webmaven/asciidoctrine.git
cd asciidoctrine
# 2. Create and activate a Python virtual environment
python3 -m venv venv
source venv/bin/activate
# 3. Install in development mode with test and documentation dependencies
pip install -e ".[test,docs]"
----
> [!NOTE]
> This package is currently under active development in tandem with the official TCK integration.
== Quick Start
AsciiDoctrine implements a two-pass resolution pipeline. First, parse the raw source to a syntax-level AST. Then, resolve the AST to a semantic, queryable Abstract Semantic Graph (ASG):
[source,python]
----
from asciidoctrine import parse_to_ast
from asciidoctrine.resolver import ASGResolver
source = """
== Section Title
This is a *bold* word in a paragraph.
"""
# 1. Parse raw source to syntax-level AST
ast = parse_to_ast(source)
# 2. Resolve to semantic ASG (resolves attributes, includes, and filters comments)
resolver = ASGResolver(ast)
asg = resolver.resolve(ast)
# Iterate through sections in the ASG blocks list
for block in asg.blocks:
if block.name == "section":
title_text = "".join(node.value for node in block.title.inlines if node.name == "text")
print(f"Found section: {title_text}")
----
=== ASG Representation
Calling `asg.to_dict()` yields a structured, spec-compliant representation matching the official AsciiDoc Language ASG schema:
[source,json]
----
{
"name": "document",
"type": "block",
"blocks": [
{
"name": "section",
"type": "block",
"level": 1,
"title": {
"name": "title",
"type": "inline",
"inlines": [
{ "name": "text", "type": "string", "value": "Section Title" }
]
},
"blocks": [
{
"name": "paragraph",
"type": "block",
"inlines": [
{ "name": "text", "type": "string", "value": "This is a " },
{
"name": "span",
"type": "inline",
"variant": "strong",
"form": "constrained",
"inlines": [
{ "name": "text", "type": "string", "value": "bold" }
]
},
{ "name": "text", "type": "string", "value": " word in a paragraph." }
]
}
]
}
],
"attributes": {}
}
----
== Roadmap to Parity
The path to 1:1 parity with Asciidoctor is tracked through the following phases:
[cols="1,3,1"]
|===
| Phase | Focus | Status
| **0** | **Foundations**: PEG grammar, structured AST, and TCK harness. | ✅
| **1** | **Advanced Blocks**: Admonitions ✅, Sidebars ✅, Source blocks ✅, and Example blocks ✅. | ✅
| **2** | **Document Infra**: Headers ✅, attributes ✅, and includes ✅. | ✅
| **3** | **Tables & Description Lists**: Nested/mixed description lists and advanced table cell alignments/spans. | ✅
| **4** | **Testing & Developer Experience**: Code block attributes, callout stripping, and precise source location coordinate tracking. | ✅
| **5** | **Static Site Generator Support**: Auto-slugified section IDs, TOC outline extraction, and cross-references (Priority 3). | ⏳
| **6** | **Sphinx Extension Support**: Metadata alignment, complete node renderer visitor auditing, and Pygments styling (Priority 4). | ⏳
|===
== Project Structure
[source,text]
----
asciidoctrine/
├── src/
│ └── asciidoctrine/ # Core parser logic
│ ├── grammar.lark # EBNF Grammar
│ ├── lark_parser.py # Transformer and Parser entry point
│ ├── nodes.py # AST Node definitions
│ └── __init__.py # Public API
├── examples/ # Real-world usage samples
├── tests/ # Unit and integration tests
├── pyproject.toml # Build configuration
└── README.adoc # This file
----
== Testing & Compliance
We prioritize correctness by testing against three fronts:
1. **Unit & Integration Tests**: Granular tests for grammar and semantic components.
2. **External Corpus (DocTest)**: Real-world examples from the `asciidoctor-doctest` corpus.
3. **TCK (Technology Compatibility Kit)**: Direct compliance with the official AsciiDoc Language TCK suite.
Run the stable Pytest suite locally with:
[source,bash]
----
pytest -k "not functional"
----
Run the official TCK test suite with:
[source,bash]
----
./run-tck.sh
----
== Self-Hosted Documentation
This project's documentation is authored entirely in AsciiDoc, compiled using Sphinx along with our own self-hosted `asciidoctrine.sphinx_ext` plugin, and published online at link:https://webmaven.github.io/asciidoctrine/[https://webmaven.github.io/asciidoctrine/]!
Build and view the documentation locally:
[source,bash]
----
# Compile HTML documentation
sphinx-build -b html docs/ docs/_build/html
# Open in your browser (macOS example)
open docs/_build/html/index.html
----
== Contributing
We welcome contributions! Please see link:CONTRIBUTING.adoc[CONTRIBUTING.adoc] for detailed development workflow guidelines, TCK compliance processes, and code style.
== License
Distributed under the **Apache License 2.0**. See link:LICENSE[LICENSE] for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
asciidoctrine-0.1.0a7.tar.gz
(62.2 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file asciidoctrine-0.1.0a7.tar.gz.
File metadata
- Download URL: asciidoctrine-0.1.0a7.tar.gz
- Upload date:
- Size: 62.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b31cd06db78317eccec8a886147bca3b0888543894621dfd29f60783bd644e00
|
|
| MD5 |
5da0784a86f959b71c5a6271fa8a3b4b
|
|
| BLAKE2b-256 |
8635abacffdcff835520443a461d67370987667f3e82175a7746d13fb2c05ee7
|
File details
Details for the file asciidoctrine-0.1.0a7-py3-none-any.whl.
File metadata
- Download URL: asciidoctrine-0.1.0a7-py3-none-any.whl
- Upload date:
- Size: 46.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4e16260919f523361774f471f13c39eec697c8c77fecde1ca5c526b7b20d97d3
|
|
| MD5 |
d59e17b1b1e0b713c11a4325a548112c
|
|
| BLAKE2b-256 |
e4913dd9979a009f568b690ae10e01a1e7c0804660958e4c5a140460108e3805
|