Emboss — precision PDF generation with optimal typesetting and PDF/UA structure
Project description
Emboss
Precision PDF generation for Python. Describe a document with a declarative spec; the engine handles typography, layout, accessibility tagging, and the complete PDF structure. Output is deterministic and byte-identical across runs.
from emboss import Document
doc = Document(title="Q3 Report", style="finance")
doc.heading("Revenue Analysis", level=1)
doc.paragraph("Revenue increased 12% year over year.")
doc.table(headers=["Region", "Q3"], rows=[["North America", "$2.4M"]])
doc.save("report.pdf")
No coordinates. No manual page-break handling. No separate accessibility pass.
Table of Contents
- Why Emboss
- Installation
- Quick Start
- EmbossSpec Format
- Document Elements
- Style Presets
- Typography Engine
- Layout System
- Accessibility (PDF/UA)
- Templates
- Domain Features
- Cross-References and Numbering
- SVG Embedding
- Multi-Column Layout
- Math Notation
- Code Blocks
- Charts
- PDF/A Archival Output
- Digital Signatures
- Validation
- Adapters
- API Reference
- Performance
- Development
- Architecture
- Roadmap
- License
Why Emboss
Existing Python PDF libraries are page-description tools: you place ink at coordinates, and the resulting file has no idea that a given string was a heading. Accessibility tags, if available at all, are painted on after the fact and tend to disagree with the visible page.
Emboss inverts that. The semantic document is the source of truth, and both the visual output and the structure tree are derived from it. They cannot drift apart, because they come from the same description.
What it guarantees
Deterministic output. The same document produces byte-identical PDFs across runs and machines. No timestamps, no random identifiers; the file /ID is derived from content. This makes output hash-verifiable for filings and diffable in CI.
assert doc.render() == doc.render() # always true
Structural correctness. Cross-reference offsets are recorded as bytes are written, so the xref table cannot disagree with the file.
Content never overflows. Everything is measured before anything is placed, so text running off the page is not a failure mode.
Tagged by default. Every document ships with a complete PDF/UA structure tree.
Installation
pip install emboss-pdf
Requires Python 3.10+ and fonttools. Optional extras:
pip install emboss-pdf[all] # pydantic + pikepdf + cryptography
pip install emboss-pdf[llm] # pydantic schemas for LLM structured output
pip install emboss-pdf[verify] # pikepdf for PDF structural verification
pip install emboss-pdf[signing] # cryptography for digital signatures
pip install emboss-pdf[dev] # full dev environment (pytest, mypy, ruff)
Quick Start
Fluent API
from emboss import Document
doc = Document(title="Quarterly Report", author="Finance Team", style="finance")
doc.heading("Executive Summary", level=1)
doc.paragraph("Revenue grew 12% year over year, driven by enterprise expansion.")
doc.table(
headers=["Metric", "Q3 2025", "Q3 2024", "Change"],
rows=[
["Revenue", "$24.1M", "$21.5M", "+12%"],
["EBITDA", "$8.2M", "$6.9M", "+19%"],
],
)
doc.bullets(["North America: +15%", "EMEA: +8%", "APAC: +11%"])
doc.save("quarterly_report.pdf")
Declarative Spec
from emboss.spec import Document, Heading, Paragraph, Table
doc = Document(
title="Quarterly Report",
style="finance",
content=[
Heading(text="Executive Summary", level=1),
Paragraph(content="Revenue grew 12% year over year."),
Table(
headers=["Metric", "Value"],
rows=[["Revenue", "$24.1M"], ["EBITDA", "$8.2M"]],
),
],
)
pdf_bytes = doc.render()
LLM Integration (3 tiers)
Tier 1: Spec prompt (full quality, recommended) -- pass the EmbossSpec format with your prompt:
from emboss import spec_prompt, Document
system = spec_prompt(style="finance") # compact EmbossSpec description for LLMs
response = client.messages.create(
model="claude-sonnet-5",
system=system,
messages=[{"role": "user", "content": "Create a Q3 financial report"}],
)
doc = Document.from_json(response.text)
doc.save("report.pdf")
Tier 2: One-liner (zero friction):
from emboss import generate
generate("Create a quarterly financial report",
style="finance", output="report.pdf",
provider="anthropic") # or "openai"
Tier 3: Markdown (works with any LLM output):
from emboss import Document
md = llm.generate("Write a quarterly report...") # any LLM, any provider
doc = Document.from_markdown(md, style="finance")
doc.save("report.pdf")
All three tiers produce identical PDF quality -- the typography engine (Knuth-Plass, optical margins, kerning) runs the same regardless of input format.
EmbossSpec Format
Documents are defined using EmbossSpec, a declarative JSON-serializable specification. Every document element is a typed Python dataclass that can be constructed programmatically, deserialized from JSON, or generated by an LLM via structured output.
The spec is the single source of truth. From it, Emboss derives:
- The visual PDF (typography, layout, pagination)
- The accessibility structure tree (PDF/UA tags)
- HTML, Markdown, and DOCX exports (via adapters)
{
"title": "Annual Report",
"style": "corporate",
"content": [
{"type": "heading", "text": "Overview", "level": 1},
{"type": "paragraph", "content": "Company performance exceeded targets."},
{"type": "table", "headers": ["Q1", "Q2"], "rows": [["$1M", "$1.2M"]]}
]
}
Document Elements
| Element | Description | Key Parameters |
|---|---|---|
Heading |
Section heading (H1-H6) | text, level (1-6) |
Paragraph |
Body text with inline formatting | content (str, TextRun, or list) |
BulletList |
Unordered list with nesting | items (strings or sublists) |
NumberedList |
Ordered list with nesting | items, start |
Table |
Data table with header row | headers, rows, caption |
Image |
Embedded image | source (path or bytes), width, caption |
CodeBlock |
Syntax-highlighted code | code, language, line_numbers |
MathBlock |
LaTeX-style math notation | source, display |
Chart |
Data visualization | chart_type, labels, values |
Callout |
Admonition box (note/warning/tip) | content, variant, title |
Footnote |
Numbered footnote | content |
SvgBlock |
Inline SVG vector graphic | source, width, height |
BibliographyBlock |
Formatted citation list | citations |
HorizontalRule |
Visual separator | thickness, color |
PageBreak |
Force a page break | -- |
Inline Formatting
from emboss import Paragraph, TextRun
Paragraph(content=[
TextRun(text="Regular text, "),
TextRun(text="bold text", bold=True),
TextRun(text=", "),
TextRun(text="italic text", italic=True),
TextRun(text=", and "),
TextRun(text="colored text", color="2563eb"),
])
Style Presets
Five professionally designed stylesheets are built in. Each encodes type scale, leading, table rules, and spacing so output looks authored without choosing a single measurement.
| Preset | Body Font | Heading Font | Body Size | Alignment | Use Case |
|---|---|---|---|---|---|
legal |
Times | Times | 11.5pt | Justified | Contracts, briefs, pleadings |
finance |
Helvetica | Helvetica | 10pt | Left | Reports, filings, data sheets |
academic |
Times | Helvetica | 11.5pt | Justified | Papers, dissertations, theses |
corporate |
Helvetica | Helvetica | 10.5pt | Left | Memos, policies, manuals |
minimal |
Helvetica | Helvetica | 9.5pt | Left | Data exports, compact reports |
# Use a preset by name
doc = Document(style="finance")
# Or create a custom stylesheet
from emboss import Style, StyleSheet
custom = StyleSheet(
name="custom",
body=Style(font_family="Helvetica", font_size=11.0, align="justify"),
h1=Style(font_size=18.0, bold=True),
# ... other elements
)
doc = Document(style=custom)
Typography Engine
Knuth-Plass Line Breaking
Emboss uses the Knuth-Plass algorithm from Knuth and Plass (1981) for optimal paragraph line breaking. Unlike greedy algorithms that fill each line independently, Knuth-Plass treats the paragraph as a shortest-path problem and minimizes total demerits across all lines:
greedy widths: [216, 172, 168, 162, 176] variance 367
optimal widths: [216, 224, 202, 198] variance 110
The algorithm:
- Models text as Box (word), Glue (elastic space), and Penalty (breakpoint) items
- Considers all legal breakpoints simultaneously
- Penalizes consecutive hyphenation, abrupt spacing changes, and tight/loose lines
- Falls back to greedy breaking only for pathological input (long URLs, extreme narrow columns)
Knuth-Liang Hyphenation
Pattern-based hyphenation with a no-break list for domain terms (Inc., LLC, EBITDA, plaintiff) that should never be split.
Real Font Metrics
Glyph advances and kerning come from the font file (via fonttools). Text is emitted with the PDF TJ operator so kerning pairs are actually applied between glyphs. Base-14 font metrics are built in, so valid PDFs are produced with no font files present.
Smart Typography
Automatic curly quotes, em/en dashes, fractions, ellipses, and non-breaking spaces before units. Applied during text processing, not as a post-process.
Ligature Support
Automatic fi, fl, ffi, and ffl ligature substitution for fonts that support them.
Layout System
Page Setup
from emboss import PageSpec
# Standard sizes
page = PageSpec.letter() # 8.5 x 11 in
page = PageSpec.a4() # 210 x 297 mm
page = PageSpec.legal() # 8.5 x 14 in
# Custom dimensions and margins (in points, 72pt = 1 inch)
page = PageSpec(
width=612, height=792,
margin_top=72, margin_bottom=72,
margin_left=90, margin_right=72,
)
# Multi-column layout
page = PageSpec.a4(columns=2, column_gap=18)
Pagination Features
- Widow and orphan control: prevents single lines stranded at page tops/bottoms
- Keep-with-next: headings stay with their following paragraph
- Keep-together: atomic blocks (callouts, small tables) never split across pages
- Table splitting: large tables break across pages with header row repetition
- Column balancing: multi-column layouts balance content across columns
Accessibility (PDF/UA)
Every document is tagged with a complete PDF/UA structure tree:
- Headings become
/H1through/H6 - Paragraphs become
/P - Tables become
/Tablewith/THand/TDcells; header cells carry/Scope - Lists become
/Lwith/LI,/Lbl, and/LBody - Images get
/Figurewith alt text - Running heads, page numbers, Bates stamps, and watermarks are marked
/Artifact
from emboss.pdf.verify import verify_pdf
result = verify_pdf(doc.render())
print(result) # VerifyResult(valid=True, structure_tree=True, ...)
Templates
Pre-configured document factories for common formats:
from emboss.templates import memo, report, letter, invoice
# Create a memo with pre-set styling
doc = memo(title="Project Update", author="Engineering Team")
doc.heading("Status", level=2)
doc.paragraph("All milestones are on track.")
doc.save("memo.pdf")
# Create an invoice
doc = invoice(title="Invoice #2024-0147", author="Acme Corp")
doc.table(
headers=["Item", "Qty", "Price"],
rows=[["Widget", "100", "$5.00"], ["Gadget", "50", "$12.00"]],
)
doc.save("invoice.pdf")
Available templates: memo, report, letter, invoice, academic_paper, legal_brief, slide_deck, data_sheet.
Domain Features
Legal and Financial Documents
from emboss import Document, LegalFeatures, PageSpec
doc = Document(
title="Memorandum of Understanding",
style="legal",
page=PageSpec.letter(margin_left=108),
legal=LegalFeatures(
watermark="CONFIDENTIAL",
watermark_opacity=0.12,
line_numbering=True,
bates_prefix="ACME-",
bates_start=1,
bates_digits=6,
bates_position="bottom-right",
),
)
- Bates numbering: sequential identifiers for legal discovery
- Line numbering: continuous line numbers in the left margin (court filings)
- Watermarks: diagonal text overlay with configurable opacity
Custom Headers and Footers
from emboss import HeaderFooter
doc = Document(
header=HeaderFooter(
left="Acme Corp",
center="Confidential",
right="{page} of {pages}",
separator=True,
),
footer=HeaderFooter(
center="Generated by Emboss",
),
)
Placeholders {page} and {pages} are resolved at render time.
Cross-References and Numbering
Automatic figure, table, and equation numbering with label-based cross-references:
from emboss import Document, Table, Image
from emboss.crossref import CrossReferenceIndex
doc = Document(title="Research Paper", style="academic")
doc.table(
headers=["Variable", "Value"],
rows=[["x", "42"]],
caption="Experimental results",
label="results-table",
)
doc.paragraph("As shown in @results-table, the value of x is 42.")
The @label syntax resolves to "Table 1", "Figure 2", etc., with automatic counter tracking via NumberingContext.
SVG Embedding
Embed SVG vector graphics as native PDF paths (no rasterization):
doc.svg("""
<svg width="200" height="100">
<rect x="10" y="10" width="80" height="80" fill="#2563eb"/>
<circle cx="150" cy="50" r="40" fill="#dc2626"/>
</svg>
""", width=200, height=100)
Supported elements: rect, circle, ellipse, line, polygon, polyline, path (M/L/H/V/C/Z commands).
Multi-Column Layout
from emboss import Document, PageSpec
doc = Document(
title="Newsletter",
page=PageSpec.a4(columns=2, column_gap=18),
)
doc.heading("Top Stories", level=1)
doc.paragraph("Content flows automatically across columns...")
Content flows between columns automatically. Column-spanning elements (headings, wide tables) are supported via the column_span style property.
Math Notation
LaTeX-style math rendering with support for common commands:
doc.math(r"E = mc^2")
doc.math(r"\int_{0}^{\infty} e^{-x^2} dx = \frac{\sqrt{\pi}}{2}")
doc.math(r"\sum_{i=1}^{n} x_i = \frac{n(n+1)}{2}")
Supported: superscripts, subscripts, fractions (\frac), square roots (\sqrt), integrals, summations, Greek letters, \text{}, \mathcal, \mathbb, \mathbf, \mathrm, \operatorname, and more.
Code Blocks
Syntax-highlighted code with optional line numbers:
doc.code_block(
code='def hello():\n print("Hello, world!")',
language="python",
line_numbers=True,
)
Built-in highlighters for Python, JavaScript, TypeScript, SQL, JSON, HTML, CSS, Rust, Go, Java, C, and more.
Charts
Data visualization rendered as native PDF vector graphics:
doc.chart(
chart_type="bar",
labels=["Q1", "Q2", "Q3", "Q4"],
values=[120, 150, 180, 210],
caption="Quarterly Revenue ($M)",
)
Chart types: bar, line, pie, scatter.
PDF/A Archival Output
Generate PDF/A-2b compliant output for long-term archival:
doc = Document(
title="Annual Filing",
pdfa=True, # enables PDF/A-2b conformance
tagged=True, # required for PDF/A
)
Includes XMP metadata, sRGB ICC output intent, and all required PDF/A catalog entries.
Digital Signatures
Sign PDFs with X.509 certificates:
from emboss.signing import sign_pdf
signed_bytes = sign_pdf(
pdf_bytes,
cert_path="signer.pem",
key_path="signer-key.pem",
reason="Approved",
location="San Francisco",
)
Requires pip install emboss-pdf[signing].
Validation
Constraints are checked before rendering. Repairable issues are fixed automatically; genuine errors are reported:
from emboss import ConstraintValidator
result = ConstraintValidator().validate(doc)
for issue in result.issues:
print(issue)
# fixed/mathematical: column widths rescaled to fit page
# fixed/structural: orphan heading moved to next page
Adapters
Export to multiple formats from a single document spec:
from emboss.adapters.html_export import to_html
from emboss.adapters.markdown_export import to_markdown
from emboss.adapters.docx_export import to_docx
html_str = to_html(doc)
md_str = to_markdown(doc)
docx_bytes = to_docx(doc)
Pydantic Schema (LLM Integration)
from emboss.adapters.pydantic_schema import build_json_schema, parse_document
# Get the JSON Schema for LLM structured output
schema = build_json_schema()
# Parse an LLM-generated JSON document
doc = parse_document(json_string)
pdf_bytes = doc.render()
API Reference
Document
Document(
title="", # Document title (appears in PDF metadata and title block)
author="", # Author name
subject="", # Subject line
keywords="", # Comma-separated keywords
language="en-US", # BCP 47 language tag
style="corporate", # Preset name or StyleSheet instance
page=PageSpec(), # Page dimensions and margins
content=[], # List of block elements
header=None, # HeaderFooter for page headers
footer=None, # HeaderFooter for page footers
page_numbers=True, # Show page numbers
tagged=True, # Generate PDF/UA structure tree
legal=None, # LegalFeatures (Bates, line numbers, watermark)
pdfa=False, # Generate PDF/A-2b compliant output
toc=False, # Generate table of contents
signatures=None, # Digital signature fields
)
Convenience Methods
doc.heading(text, level=1)
doc.paragraph(content)
doc.bullets(items)
doc.numbered(items)
doc.table(headers, rows)
doc.image(source, width=None, caption=None)
doc.chart(chart_type, labels, values)
doc.code_block(code, language="text")
doc.math(source)
doc.footnote(content)
doc.callout(content, variant="note")
doc.svg(source)
doc.rule()
doc.page_break()
doc.render() -> bytes
doc.save(path)
Performance
- Layout is pure Python with Knuth-Plass optimal line breaking
- Font metrics are cached with LRU caching
- A typical 10-page document renders in ~50-100ms
- 200-page documents are supported with ~20-50MB peak memory
- O(n) linear scaling with document size
Development
Setup
git clone https://github.com/GGChamp85/Emboss.git
cd Emboss
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
Running Tests
pytest # run all tests
pytest tests/test_emboss.py # run core tests
pytest -q # quiet mode
Code Quality
ruff check src/ tests/ # linting
mypy src/emboss/ # type checking
Generating Example PDFs
python examples/financial_report.py
python examples/legal_pleading.py
python examples/showcase.py
Architecture
Document (EmbossSpec)
|
v
ConstraintValidator --> validate + auto-fix
|
v
LayoutEngine --> measure (font metrics) --> paginate (Knuth-Plass)
|
v
Renderer --> ContentStream (PDF operators) + StructureTree (PDF/UA tags)
|
v
PDFAssembler --> byte-exact PDF with content-derived /ID
Package Structure
src/emboss/
spec.py # Document data model (EmbossSpec)
writer.py # Render pipeline: measure -> paginate -> render
styles.py # Cascading style system + presets
constraints.py # Validation + auto-fix
layout/
engine.py # Layout engine: measurement + pagination
typography/
font_metrics.py # Glyph metrics, kerning, subsetting
line_breaking.py # Knuth-Plass line breaker
hyphenation.py # Knuth-Liang hyphenation
ligatures.py # fi/fl/ffi/ffl substitution
pdf/
assembler.py # Low-level PDF object writer
fonts.py # Font resource building + subsetting
objects.py # PDF object types (Dict, Array, Stream)
streams.py # Content stream operators (text, shapes)
tags.py # PDF/UA structure tree builder
verify.py # Structural verification
adapters/
html_export.py # HTML export
markdown_export.py # Markdown export
docx_export.py # DOCX export
pydantic_schema.py # JSON Schema / Pydantic models for LLM
math_render.py # LaTeX math parser + renderer
code_highlight.py # Syntax highlighting
charts.py # Chart rendering
images.py # Image handling
svg.py # SVG parser + renderer
crossref.py # Cross-reference resolution
numbering.py # Figure/table auto-numbering
templates.py # Document templates
bibliography.py # Citation formatting
colors.py # Color palettes + theming
intelligence.py # Content analysis + smart typography
toc.py # Table of contents generation
pdfa.py # PDF/A-2b conformance
signing.py # Digital signatures
redaction.py # Content redaction
slides.py # Slide/presentation layout
Roadmap
| Phase | Scope | Status |
|---|---|---|
| 1 | Core engine, typography, layout, tagging | Done |
| 2 | Unicode CIDFont, code highlighting, math, bibliography, charts | Done |
| 3 | Images, SVG, TOC, multi-column, templates | Done |
| 4 | PDF/A, redaction, digital signatures | Done |
| 5 | Cross-references, custom headers/footers, numbered lists | Done |
| 6 | Micro-typography, GPOS kerning, performance, CMYK | In progress |
License
Apache-2.0. See LICENSE for the full text.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file emboss_pdf-0.1.0.tar.gz.
File metadata
- Download URL: emboss_pdf-0.1.0.tar.gz
- Upload date:
- Size: 195.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
18086bb8f62efa53bef2bc3b9cfec0520ef54fcb542a3ad2aa440191b212b6b0
|
|
| MD5 |
2f660384bd9d0c59ee21a239064f1dce
|
|
| BLAKE2b-256 |
b1cb90d83da13038a68ae55855bfd95dfeb408d38feeb14a083fec57c07516e0
|
Provenance
The following attestation bundles were made for emboss_pdf-0.1.0.tar.gz:
Publisher:
publish.yml on GGChamp85/Emboss
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
emboss_pdf-0.1.0.tar.gz -
Subject digest:
18086bb8f62efa53bef2bc3b9cfec0520ef54fcb542a3ad2aa440191b212b6b0 - Sigstore transparency entry: 2256756422
- Sigstore integration time:
-
Permalink:
GGChamp85/Emboss@94d71d062c68165aec0492cf3b24aca8748e8e3b -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/GGChamp85
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@94d71d062c68165aec0492cf3b24aca8748e8e3b -
Trigger Event:
release
-
Statement type:
File details
Details for the file emboss_pdf-0.1.0-py3-none-any.whl.
File metadata
- Download URL: emboss_pdf-0.1.0-py3-none-any.whl
- Upload date:
- Size: 156.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
de8817b6ab1a161b993998af2280b9aa9ffd31f693db6d6fddf664a15afd22a9
|
|
| MD5 |
1587e9b34d1e3f027883323d0a98e393
|
|
| BLAKE2b-256 |
997d208ed152e06281f104c67ff3349f82b654cd6335eaf6d94cbb4623cc15e4
|
Provenance
The following attestation bundles were made for emboss_pdf-0.1.0-py3-none-any.whl:
Publisher:
publish.yml on GGChamp85/Emboss
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
emboss_pdf-0.1.0-py3-none-any.whl -
Subject digest:
de8817b6ab1a161b993998af2280b9aa9ffd31f693db6d6fddf664a15afd22a9 - Sigstore transparency entry: 2256756428
- Sigstore integration time:
-
Permalink:
GGChamp85/Emboss@94d71d062c68165aec0492cf3b24aca8748e8e3b -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/GGChamp85
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@94d71d062c68165aec0492cf3b24aca8748e8e3b -
Trigger Event:
release
-
Statement type: