Skip to main content

to-CEI

Project description

to_cei — Monasterium CEI builder

A small Python library for assembling CEI (Charter Encoding Initiative) charters and charter groups as objects, then serialising them to XML files ready for import into Monasterium.net.

Requires Python 3.13+.

Install

pip install to_cei
# or
uv add to_cei

Quick start

Build a minimal valid charter and write it to disk:

from to_cei import Charter

charter = Charter(
    id_text="1457_03_15",
    id_norm="urn:example:1457_03_15",
    date_value="14570315",
    abstract="Konrad von Lintz beurkundet einen Vertrag.",
    issuers="Konrad von Lintz",
    issued_place="Wien",
)

charter.to_file("./out")  # writes ./out/1457_03_15.cei.xml

charter.to_xml() returns an lxml._Element if you would rather embed the result in a larger pipeline.

Dates

Real archival data routinely mixes complete dates with partial dates ("only the year is known", "sometime in the 15th century") and "no date" markers. to_cei.dates.parse handles all three: a Time for exact dates, a (Time, Time) range for partial ones, None for explicit no-date markers.

from to_cei import dates

# Numeric / standard
dates.parse("1457-03-15")               # ISO 8601         -> Time
dates.parse("14570315")                 # MOM exact        -> Time
dates.parse("14570399")                 # MOM unknown day  -> month range
dates.parse("14579999")                 # MOM unknown month-> year range
dates.parse("15.03.1457")               # German DMY       -> Time
dates.parse("1457.03.15")               # YMD dotted       -> Time
dates.parse("15/03/1457")               # slashed DMY      -> Time

# Year-only and German archival hedging
dates.parse("1457")                     # bare year        -> full year
dates.parse("um 1457")                  # ca./circa/etwa   -> full year
dates.parse("1457 oder 1458")           # either-or        -> 2-year span
dates.parse("zwischen 1457 und 1460")   # between          -> 4-year span
dates.parse("15. Jahrhundert")          # century          -> 1401-1500
dates.parse("Anfang des 15. Jahrhunderts")  # century third

# Explicit range
dates.parse(("1457-03-15", "1457-03-20"))

Recognised string formats (in priority order)

# Format Examples
1 ISO 8601 "1457-03-15"
2 MOM numeric (YYYYMMDD) "14570315", "14570399", "14579999"
3 Dotted/slashed numeric "15.03.1457" (German DMY), "1457.03.15" (YMD), "15/03/1457", "15-03-1457"
4 Bare year "1457" → full-year range
5 German fuzzy "um/ca./circa/gegen/etwa 1457", "1457 oder 1458", "zwischen X und Y", "15. Jahrhundert", "15. Jh.", "Anfang/Mitte/Ende des 15. Jahrhunderts"

If parsing fails, the exception lists every format that was tried with its rejection reason — so you can see at a glance whether your input just needs a small reformat or a new parser.

Plug in your own format

If your data has a dialect the built-ins don't cover, pass a parser via extra_parsers=. It receives a stripped, non-empty, non-no-date string and must either return a Time / (Time, Time) or raise ValueError to mean "not my format" — exactly the contract the built-ins follow. Extra parsers run after the built-ins, so they only see strings nothing else could handle:

def julian_day(text: str) -> Time:
    if not text.startswith("JD"):
        raise ValueError("not a julian-day string")
    return Time(float(text[2:]), format="jd", scale="ut1")

dates.parse("JD2451545.0", extra_parsers=[julian_day])

Charter(date_value=...) accepts the same shapes (it delegates to parse), so this also works at construction time.

"No date" markers

Real catalogues use a small zoo of "no date" labels. parse recognises them and returns None — they're a meaningful value, not an error:

dates.parse("sine dato")    # Latin    -> None
dates.parse("s.d.")
dates.parse("ohne Datum")   # German   -> None
dates.parse("o.D.")
dates.parse("undatiert")
dates.parse("n.d.")         # English  -> None
dates.parse("undated")
dates.parse("99999999")     # MOM sentinel -> None
dates.parse("")             # empty -> None

Charter(date_value=...) delegates to parse, so any of these strings produce date_value=None (serialized as @value="99999999"). Use to_cei.dates.is_no_date(text) for an explicit predicate; the full list lives in to_cei.dates.NO_DATE_MARKERS.

Multiple issuers

Charter(
    id_text="example",
    issuers=["Konrad von Lintz", "Heinrich der Praitenvelder"],
)

Abstract as XML

When you need structured markup inside the abstract (e.g. nested persName elements), pass an lxml element instead of a string. Note that doing so currently means you must build the issuer markup into the abstract yourself — issuers= and an XML abstract are mutually exclusive at serialisation time.

from lxml import etree
from to_cei.config import CEI

abstract = CEI.abstract(
    "Konrad von Lintz beurkundet einen Vertrag mit ",
    CEI.persName("Heinrich"),
    ".",
)

Charter(id_text="example", abstract=abstract)

Image URLs (graphic_urls)

graphic_urls accepts anything from a full URL down to a bare filename — the CEI schema doesn't care, and Monasterium resolves short forms against a per-fond / per-collection base_url configured on the server. The most common new-user trap is assuming you must always supply a fully-qualified URL.

# All three are valid CEI:
Charter("1", graphic_urls="https://example.org/charters/1.jpg")  # full URL
Charter("1", graphic_urls="2024/scans/1.jpg")                    # path fragment
Charter("1", graphic_urls="1.jpg")                               # bare filename

When to use which:

  • Full URL — required if the image lives somewhere outside the fond/collection's configured base_url, or if you're producing CEI for use outside Monasterium.
  • Path fragment / bare filename — preferred when all images for a fond live under a common base. Monasterium joins them with the fond's base_url at display time, so you avoid hard-coding the host in every charter and the fond can be moved later without rewriting the data. Coordinate the base path with whoever administers the fond on Monasterium.

Pass a list for multiple images:

Charter("1", graphic_urls=["1_recto.jpg", "1_verso.jpg"])

Charter groups

from to_cei import Charter, CharterGroup

group = CharterGroup(
    name="Schotten OSB",
    charters=[Charter(id_text="a"), Charter(id_text="b")],
)
group.to_file("./out")

Validation

from to_cei import Charter, Validator

v = Validator()                     # downloads + caches the CEI XSD once
v.validate_cei(charter.to_xml())    # raises XMLSchemaValidationError on failure
v.is_valid_cei(charter.to_xml())    # bool variant

The XSD is cached on disk under ~/.cache/to-cei/ after the first fetch.

Field reference

See docs/fields.md for the complete table of every Charter, Seal, and CharterGroup keyword argument with the CEI element each one produces.

Develop

uv sync           # create the venv with deps
uv run pytest     # run tests
just test         # same, via the justfile

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

to_cei-0.4.0.tar.gz (46.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

to_cei-0.4.0-py3-none-any.whl (34.9 kB view details)

Uploaded Python 3

File details

Details for the file to_cei-0.4.0.tar.gz.

File metadata

  • Download URL: to_cei-0.4.0.tar.gz
  • Upload date:
  • Size: 46.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for to_cei-0.4.0.tar.gz
Algorithm Hash digest
SHA256 3f6b1ffecbc5ac4811f8466f90a982217ae2889701175e898e4828a6a90f1cd9
MD5 2379af7cb8b113e30d48c2d0cc789344
BLAKE2b-256 77cb938b9d0100d5b8e30b4fb74582ad6ef6a45e6f42979189fe87fe0ce4c85d

See more details on using hashes here.

File details

Details for the file to_cei-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: to_cei-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 34.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for to_cei-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 49f7123fdf1749982b48acb514c3b1899506c49a23fc8c9d36b5a71075e4e8ba
MD5 013d28dd12884cacd707a781fe9b810c
BLAKE2b-256 6544cca10318c35009042088b5a8600cc596aa57352dc4dd04e25c3a8958a391

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page