Skip to main content

Spans and Trees

A small Python library for converting between XML trees and a span-based structure. This can be useful for extracting sections of text from XML documents and doing special things with some of the tags.

The two main functions are tree_to_spans and spans_to_tree for converting between an ElementTree element and text with a list of spans. Examples are shown below.

tree_to_spans

First create a little example XML tree to convert.

import xml.etree.ElementTree as ET

xmlstring = "<doc><title>Important document</title><contents>Empty</contents></doc>"
root = ET.ElementTree(ET.fromstring(xmlstring)).getroot()

Then use the tree_to_spans function to convert the XML document into the text content with spans.

from spans_and_trees import tree_to_spans

text, spans = tree_to_spans(root)

print(text)  # Important documentEmpty
print(spans) # [(0, 18, 'title', {}), (18, 5, 'contents', {})]

The format of the spans are a tuple of length 4. The element contents are:

  1. The start location of the span
  2. The length of the span
  3. The tag of the span
  4. A dictionary of the attributes of the span.

spans_to_tree

Now we create a dummy document with a block of text and a span at particular offset.

from spans_and_trees import spans_to_tree

text = 'The quick brown fox jumped over the lazy dog'
spans = [ (10,5,'colour',{'dummy_attrib':'5'}) ] # The span starts at 10, has length of 5, is a 'colour' tag and has a dummy attribute.

root = spans_to_tree(text, spans)

print(type(root)) # <class 'xml.etree.ElementTree.Element'>

We can check the XML tree that has been created:

xmlstr = ET.tostring(root)

print(xmlstr) # b'<tree>The quick <colour dummy_attrib="5">brown</colour> fox jumped over the lazy dog</tree>'

The root element's tag defaults to "tree", since spans don't record a tag for the whole document (tree_to_spans doesn't include one either). Pass root_tag to use something else, e.g. spans_to_tree(text, spans, root_tag="doc").

spans_to_passages

spans_to_passages takes the text/spans output of tree_to_spans and splits it into a list of text passages, one per split_tags element (e.g. p, sec), dropping the content of any ignore_tags element (e.g. table), and attaching any keep_tags spans (e.g. bold) that fall within each passage.

from spans_and_trees import tree_to_spans, spans_to_passages

xmlstring = """
<article>
	<sec>
		<title>Introduction</title>
		<p>This is <bold>important</bold> background text.</p>
		<table-wrap>Some table content we want to ignore.</table-wrap>
		<p>A second paragraph follows.</p>
	</sec>
</article>
"""

root = ET.ElementTree(ET.fromstring(xmlstring)).getroot()
text, spans = tree_to_spans(root)

passages = spans_to_passages(text, spans, ignore_tags={'table-wrap'}, split_tags={'title','p'}, keep_tags={'bold'})
for p in passages:
	print(p)

# {'start': 5, 'end': 17, 'text': 'Introduction', 'spans': []}
# {'start': 20, 'end': 54, 'text': 'This is important background text.', 'spans': [(8, 9, 'bold', {})]}
# {'start': 97, 'end': 124, 'text': 'A second paragraph follows.', 'spans': []}

Each passage's start/end are offsets into the original text; each attached span's offsets are relative to the passage's own text.

Example: a real PMC article via Entrez

Here the spans_and_trees.pmc tag sets are applied to a real open-access article fetched from NCBI's Entrez E-utilities. For PMC articles, there is a helper function cleanup_pmc_text which does some cleaning of common Unicode problems. There are also pre-prepared tag lists for PMC to ignore, split on and keep.

from spans_and_trees.pmc import cleanup_pmc_text, PMC_IGNORE_TAGS, PMC_SPLIT_TAGS, PMC_KEEP_TAGS

import urllib.request

pmcid = "13488400"  # https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13488400/
url = f"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pmc&id={pmcid}&rettype=full&retmode=xml"

with urllib.request.urlopen(url) as response:
	root = ET.fromstring(response.read())

body = root.find(".//body")
text, spans = tree_to_spans(body)
text = cleanup_pmc_text(text)
passages = spans_to_passages(text, spans, ignore_tags=PMC_IGNORE_TAGS, split_tags=PMC_SPLIT_TAGS, keep_tags=PMC_KEEP_TAGS)

print(passages[0])
# {'start': 0, 'end': 12, 'text': 'Introduction', 'spans': []}

Since keep_tags includes sup, a passage containing e.g. an isotope like "¹H" keeps that as a span rather than losing it. Passing a passage's text/spans back into spans_to_tree rebuilds that markup:

passage = passages[22]
print(passage)
# {'start': 14361, 'end': 14380, 'text': '1H NMR Spectroscopy', 'spans': [(0, 1, 'sup', {})]}

elem = spans_to_tree(passage["text"], passage["spans"])
print(ET.tostring(elem))
# b'<tree><sup>1</sup>H NMR Spectroscopy</tree>'

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

spans_and_trees-0.2.0.tar.gz (8.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

spans_and_trees-0.2.0-py3-none-any.whl (7.3 kB view details)

Uploaded Python 3

File details

Details for the file spans_and_trees-0.2.0.tar.gz.

File metadata

  • Download URL: spans_and_trees-0.2.0.tar.gz
  • Upload date:
  • Size: 8.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.15

File hashes

Hashes for spans_and_trees-0.2.0.tar.gz
Algorithm Hash digest
SHA256 fc861f394fe4dd59abfe96e3c30ede1c616b7f048f5f9adb9bed2090e9cfc35e
MD5 4034d188fcb14ebdc2d1c147c7676222
BLAKE2b-256 7f6b26b7d1b65ad0f4de5ee4d541dcc4baa0759ec2ec0a271149eb7db66680b9

See more details on using hashes here.

File details

Details for the file spans_and_trees-0.2.0-py3-none-any.whl.

File metadata

File hashes

Hashes for spans_and_trees-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 69f730b394e9d800db71f9e9d503707f10cebf6679514450e82df4ba7e6e8ad1
MD5 525c8b5c9f0252170360a00de3419cda
BLAKE2b-256 d0b60edcefa30e3a4209283590113400d976ed83c960b8d87c8d8e09ce7962f1

See more details on using hashes here.

Release history Release notifications | RSS feed

0.3.1

2 files

0.3.0

2 files

This release

0.2.0 This release

2 files

0.1.4

1 file

0.1.3

1 file

0.1.2

1 file

0.1.1

1 file

0.1.0

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page