Skip to main content

The ietfdata library - Access the IETF Datatracker and related resources

This project contains Python 3 libraries to retrieve and work with data from the IETF Datatracker, IETF Mail Archive, RFC index, and related resources.

Installation

The ietfdata library is distributed as a Python package. You should be able to install via pip in the usual manner:

pip install ietfdata

Accessing the IETF Datatracker

The DataTracker class provides an interface for programmatic access to the IETF Datatracker, providing metadata about the development of IETF standards.

Instantiation

There are two ways to instantiate this class, depending on how it is to be used. The normal way, when writing code to perform analysis of a snapshot of the IETF data, for example if writing a research paper, a dissertation, or as part of a student project, is to use an archive file:

dt = DataTracker(DTBackendArchive("archive/ietf-dt.sqlite"))

When instantiated in this manner, the DataTracker class will read from the specified sqlite database.

If the specified sqlite database does not exist, then the DataTracker class will fetch a complete copy of the data from the IETF Datatracker. This will take around 24 hours, and will produce database that is about 2GB in size (if interrupted, it is safe to rerun the above operation and the download will resume where it left-off). Once the sqlite database is downloaded, future instantiations of the DataTracker will read from it directly and will not access the online IETF Datatracker, making them much faster and avoiding overloading the IETF's servers.

The following can be run from the command line to fetch a copy of the database:

  python3 -m ietfdata.tools.download_dt archive/ietf-dt.sqlite

If you are working on a paper, project, or dissertation with a group of people, one person should create the sqlite database and share a copy with the others. This avoids overloading the IETF's servers, and ensures that everyone working in the group generates the same results.

Alternatively, when writing code to perform live queries of the IETF Datatracker, for example as part of a tool that provides an interactive dashboard or status report, the DataTracker should be instantiated as follows:

dt = DataTracker(DTBackendLive())

In this case, the DataTracker class will directly query the online IETF Datatracker for every request you make. This is appropriate when making small numbers of queries, for exploratory programming or when performing a live status check, but must not be used for tasks that need to make large numbers of queries. The IETF will block your access if you make many queries using DTBackendLive().

Usage

The DataTracker provides an extensive API that is best explored by reading the source code for datatracker.py and datatracker_types.py. The examples/ directory contains a number of examples of how to use the library.

Start by importing and instantiating the library:

from ietfdata.datatracker import *

dt = DataTracker(DTBackendArchive("archive/ietf-dt.sqlite"))

Then follow the suggestions below, and read the relevant sections of the datatracker.py source code, for examples of how to access the data.

People

To find information about a person:

p = dt.person_from_email("csp@csperkins.org")
print(p.name)
print(p.biography)

Documents

To find information about a document:

d = dt.document_from_rfc("RFC9000")
print(d.title)
print(d.group)

See also the discussion of Datatracker Extensions below.

Groups

To find information about a group:

d = dt.document_from_rfc("RFC9000")
g = dt.group(d.group)
print(g.acronym)

for e in dt.group_events(group = g):
  print(e.time)
  print(e.desc)

Meetings

(tbd)

Intellectual Property Rights Disclosures

(tbd)

Accessing the IETF Datatracker Extensions

The DataTrackerExt class is a subclass of DataTracker that provides additional features on top of those provided by the IETF Datatracker.

Instantiation

The DataTrackerExt class is instantiated in an analogous manner to the DataTracker class:

from ietfdata.datatracker_ext import *

dte = DataTrackerExt(DTBackendArchive("archive/ietf-dt.sqlite"))

Usage

Since it's a subclass of the DataTracker, any of the methods that can be used on the DataTracker can also be used with DataTrackerExt.

The DataTrackerExt offers a number of other useful features including the ability to find the history of an RFC:

from ietfdata.datatracker_ext import *
from ietfdata.rfcindex        import *

dte = DataTrackerExt(DTBackendArchive("archive/ietfdata-dt.sqlite"))
ri  = RFCIndex(rfc_index="archive/rfc-index.xml")
rfc = ri.rfc("RFC9000")
for d in dte.draft_history_for_rfc(rfc):
    print("    {0: <50} | {1} | {2}".format(d.draft.name, d.rev, d.date.strftime("%Y-%m-%d")))

or the history of an Internet-draft:

dte = DataTrackerExt(DTBackendArchive("archive/ietfdata-dt.sqlite"))
doc = dt.document_from_draft("draft-ietf-avtcore-ecn-for-rtp")
for d in dte.draft_history(doc):
    print("    {0: <50} | {1} | {2}".format(d.draft.name, d.rev, d.date.strftime("%Y-%m-%d")))

It also contains methods to find the people who currently hold various leadership roles in the IETF, IRTF, and IAB, and the set of currently active working groups and research groups, for example:

c = dte.ietf_chair()
print(c.name)

for p in dte.working_group_chairs():
    print(p.name)

Finally, the DataTrackerExt class contains a method that given a name and an email address, for example as might be extracted from an email "From:" header, tries to find a person in the DataTracker. This uses a number of heuristics to find the right person even if there is no exact match:

p1 = dte.person_from_name_email("Colin Perkins", "csp@csperkins.org")
print(p1.id)
p2 = dte.person_from_name_email("Colin Perkins via Datatracker", "noreply@ietf.org")
print(p2.id)

Accessing the IETF Mail Archive

The MailArchive3 class provides an interface to accessing the IETF email archive.

Instantiation

The MailArchive3 class is instantiated as follows, giving a path to an sqlite database containing a copy of the archive:

from ietfdata.mailarchive3 import *
ma = MailArchive("archive/ietf-ma.sqlite")

If the specified sqlite database does not exist, the ma.update() method can be called to download a complete copy of the mail archive and store it in the database. The mail archive is approximately 40 gigabytes in size and will take around 24 hours to download. If the sqlite database file already exists, calling ma.update() will only fetch new messages, and so will be much faster.

The following can be run from the command line to fetch a copy of the mail archive and create the sqlite database:

  python3 -m ietfdata.tools.download_ma_ietf archive/ietf-ma.sqlite

If you are working on a paper, project, or dissertation with a group of people, one person should create the sqlite database and share a copy with the others. This avoids overloading the IETF's servers, and ensures that everyone working in the group generates the same results.

Usage

Once you have a copy of the sqlite database containing the mail archive, start by importing and instantiating the library:

from ietfdata.mailarchive3 import *
ma = MailArchive("archive/ietf-ma.sqlite")

Once this is done, you can find the mailing list names:

for ml_name in ma.mailing_list_names()
  print(ml_name)

You can find information about a particular mailing list:

ml = ma.mailing_list("quic")
print(ml.num_messages())

Each mailing list is represented by a MailingList object. That has a messages() method to retrieve the messages, and a threads() method to retrieve all discussion threads.

You can find information about the messages sent to a mailing list:

ml = ma.mailing_list("quic")
for envelope in ml.messages():
  print(f"From:    {envelope.from_()}")
  print(f"To:      {envelope.to()}")
  print(f"Subject: {envelope.subject()}")
  print(f"Date:    {envelope.date()}")
  print(f"Message-Id:  {envelope.message_id()}")
  print("")

Each email message is represented by an Envelope object. The envelope has methods (from_(), to(), subject(), etc.) to access the header fields, a contents() method to retrieve the message contents, and replies() and in_reply_to() methods to follow the thread of discussion.

Each email message on the server is uniquely identified by the combination of the name of the mailing list it was sent to, and the uidvalidity() and uid() fields of the message. Each message also has a message_id() that identifies the message.

If a message is sent copied to several different mailing lists, then it will appear in the mail archive several times, one copy in each mailing list. Each copy will have a different mailing list, uidvalidity() and uid(), but all will have the same message_id().

Read the source code for mailarchive3.py for details.

Accessing the RFC Index

(tbd)

See rfcindex.py

Entity Resolution

One of the challenges in working with the IETF data is determining whether different names or identifiers represent the same person or organisation (this is known as "entity resolution"). For example, the email addresses csp@csperkins.org, colin.perkins@glasgow.ac.uk, csp@isi.edu, and c.perkins@cs.ucl.ac.uk all represent the same person, but working in different jobs at different stages of their career. Similarly, "Technische Universität München", "TU Munich", and "TU Muenchen" all represent the same university.

The ietfdata library contains code that (attempts to) perform entity resolution. This can be run from the command lines as follows:

python3 -m ietfdata.tools.participants archive/ietf-dt.sqlite archive/ietf-ma.sqlite participants.json

python3 -m ietfdata.tools.organisations archive/ietf-dt.sqlite archive/rfc-index.xml organisations.json

python3 -m ietfdata.tools.affiliations archive/ietf-dt.sqlite archive/rfc-index.xml participants.json organisations.json affiliations.json

Running these commands will generate three files:

  • The file participants.json contains information about the people, giving each participant in IETF a unique identifier (e.g., PID:063009) that is associated with their names, email addresses, DataTracker identifier, GitHub username, any other identifying information that can be extracted.

  • The file organisations.json contains information about organisations, giving each a unique identifier (e.g., ORG:001156) that's associated with the different names the organisation has been given and the domain names it uses.

  • The file affiliations.json, matches participants to organisations at different stages of their career.

As of September 2026, the entity resolution code runs but has known problems and limitations that mean the results are not always accurate.

GitHub Access

IETF working groups increasing make use of GitHub to prepare documents. The ietfdata library contains minimal, extremely limited, code to fetch relevant data from GitHub:

from ietfdata.github import GitHub

gh = GitHub()
for issue in gh.issues("quicwg", "base-drafts"):
    print(issue)

for comment in gh.comments_for_issue("quicwg", "base-drafts", "5010"):
    print(comment)

user = gh.user("csperkins")
print(user)

for repo in gh.repos_for_user("csperkins"):
    print(repo)

NOTE: GitHub aggressively rate limits access for unauthenticated users to 60 requests per hour. Set the environment variable GITHUB_API_TOKEN to your GitHub access token before using this code to receive the higher rate limit (5000 requests per hour) available to logged-in GitHub users. If you don't have a GitHub access token, see https://github.com/settings/tokens when logged in to GitHub and select "Generate new token".

Development

To modify the ietfdata library, clone from GitHub then follow the instructions below to install dependencies and test the results. If you just intend to use the library to support writing a paper, as part of a student project, or to perform some other analysis, you can skip the remainder of this document.

Create a virtual environment and install dependencies in the usual manner:

python3 -m venv venv/
source venv/bin/activate
python3 -m pip install -e .

Once the virtual environment is started, running:

python3 tests/test_datatracker.py 

will run the test suite for the datatracker module. Running:

python3 tests/test_rfcindex.py

Will test the rfcindex module.

Release Process

  • Edit CHANGELOG.md and ensure up-to-date
  • Edit pyproject.toml to ensure the correct version number is present
  • Edit ietfdata/dt_backend.py to ensure the correct version number
  • Edit ietfdata/github.py to ensure the correct version number
  • Run make test to run the test suite. If any tests fail, fix then restart the release process
  • Commit changes and push to GitHub
  • Check that the GitHub Continuous Integration run succeeds, and fix any problems (this runs with a fresh cache, so can sometimes catch problems that aren't found by local tests).
  • Run python3 -m build --wheel to prepare the package
  • Run python3 -m twine upload dist/* to upload the package
  • Commit the packages files in dist/* push to GitHub
  • Tag the release in GitHub

Release files for ietfdata 0.9.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ietfdata 0.9.1
File Size Uploaded
ietfdata-0.9.1.tar.gz 109.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ietfdata 0.9.1
File Interpreter ABI Platform
ietfdata-0.9.1-py3-none-any.whl Python 3 none any Details

Total release size: 181.3 kB

Release files / ietfdata-0.9.1.tar.gz

Download URL ietfdata-0.9.1.tar.gz
Size 109.4 kB
Tags Source
SHA-256 checksum
How to use checksums
29ac8311cfd695ebed0dd6444d9e243e945fcd51b9b97976ba43ea0b44eba09c
BLAKE2b-256 checksum
How to use checksums
9373a20fabacd0fd3b8e2835c6ad6e4aab919879577a80859b4284fa61fc7a25
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release files / ietfdata-0.9.1-py3-none-any.whl

Download URL ietfdata-0.9.1-py3-none-any.whl
Size 71.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cdce203e6edd3579cfee55c31127cad1512a035417eb6aa14451d25bbd27e695
BLAKE2b-256 checksum
How to use checksums
5fa5f51bfad6f177cee411d4c4f0c5bf65da3d4eb3cf2c9c097e43b0a90605f9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release history Release notifications | RSS feed

0.9.2

2 release files

This release

0.9.1 This release

2 release files

0.9.0

2 release files

0.8.3

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.8

2 release files

0.6.7

2 release files

0.6.6

2 release files

0.6.5

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page