Skip to main content

dofjson

Prototype client for the JSON open data service exposed by sidof.segob.gob.mx, the Secretaría de Gobernación's system for Mexico's official gazette (DOF, Diario Oficial de la Federación).

The service's public docs only show sample responses, but its real, unauthenticated endpoints were found under https://sidof.segob.gob.mx/dof/sidof/:

Endpoint Description
GET /diarios/porFecha/DD-MM-YYYY Edition metadata for a date (Matutina/Vespertina/Extraordinaria)
GET /diarios/{YYYY} A whole year's FechasSinPublicacion — the dates it claims had no gazette
GET /notas/DD-MM-YYYY Notes/documents published on a date
GET /notas/nota/{codNota} Full detail of a single note, including its HTML content
GET /indicadores/DD-MM-YYYY Economic indicators (exchange rate, TIIE, UDIS)

Note that this service reports a missing day as 200 OK with empty note lists, not as an error, and that some of the dates it lists as unpublished were in fact published — see the days SIDOF loses.

This is an experimental package for evaluating whether this service is a viable alternative (or complement) to dof2md's PDF download + Markdown conversion pipeline — notes already come with structured HTML content, which may be easier to work with than OCR'd PDFs.

On top of the raw endpoints, the client offers note-scoped downloads that resolve a note's page span (infer_paginas) and fetch it in whichever form you want:

  • download_nota_imagenes(codNota) — the note's scanned page image(s).
  • download_nota_pdf(codNota) — the note as its own PDF: the whole edition PDF (there is no per-note PDF endpoint) sliced to just the note's pages, using pypdf.

Usage

pip install -e "packages/dofjson[test]"
dofjson 2026-07-16 --endpoint notas --outdir output

Building a local archive of daily indexes (--archivo)

dofjson --archivo downloads the daily notes index incrementally, day by day, over a whole date range (by default from January 2, 1917 to today). For each date it does exactly what dofjson YYYY-MM-DD --endpoint notas does — get_notas(date) filtered with quita_notas_sin_titulo — and saves one JSON per day. It does not download each note's content or scanned images: only the index.

dofjson --archivo                                      # 1917-01-02 -> today
dofjson --archivo --desde 1980-01-01 --hasta 1980-12-31
dofjson --archivo --pausa 1.0                          # slower (kinder to the server)

Output goes to notas-archivo/ (configurable with --outdir), a local, never-committed directory (it is in the repo's .gitignore), with the same per-day filenames the plain command produces:

notas-archivo/
  .completados                 # registry of finished days (for resuming)
  2026/
    15072026-notas.json        # index for 2026-07-15 (get_notas, filtered)
    16072026-notas.json
  1980/
    02011980-notas.json

The mode is resumable and idempotent: the .completados registry records the finished days, so each run only fetches what is missing. Days that fail with network errors are not marked and get retried on the next run; days with no edition (holidays, weekends) are marked so they are not retried forever. "Today" is never marked, so late additions are picked up by a later run. You can interrupt with Ctrl-C and resume at any time.

The full range is ~40,000 days: a long download, meant to be run in parts. Start with a bounded range via --desde/--hasta if you only need an era.

The days SIDOF loses, and where they are recovered from (--respaldo)

SIDOF does not report a day it is missing as an error. It answers 200 OK with every note list empty — which is also how it reports a Sunday — and lists the date under FechasSinPublicacion in GET /diarios/{year}. Most of those dates are genuine: weekends and holidays. Some are not.

On 8 March 1999 the DOF published the decree amending articles 16, 19, 22 and 123 of the Constitution. SIDOF has no trace of it: the day is empty, the note's codNota returns {"Nota": []}, and its codDiario 404s. Sampling four years (1999, 2006, 2010, 2020) turned up eight such days — dates SIDOF calls unpublished that were published.

www.dof.gob.mx, the DOF's own website, is a separate system on a separate database, and it has them. So an empty answer is no longer taken at face value: on a weekday, the day is put to the website before being written off.

dofjson --archivo                        # habiles (default): re-check Mon-Fri
dofjson --archivo --respaldo todos       # also weekends (~10,000 more requests)
dofjson --archivo --respaldo nunca       # trust SIDOF alone
[1999-03-08] SIDOF no la tiene; recuperada de dof.gob.mx

The same applies to a single date, so a lost day is reachable directly:

dofjson 1999-03-08 --endpoint notas      # -> "fuente": "dof.gob.mx", 22 notas

The text of those notes is recoverable too, not just their titles. dofweb.get_nota(codNota) reads a note's page on the website and returns it in the shape client.get_nota() uses, with the note's HTML in cadenaContenido — the same string SIDOF would have served, so nota2md converts it to Markdown by the ordinary HTML path:

from dofjson import dofweb

dofweb.get_nota(4997808)["Nota"]["cadenaContenido"]   # DOF 03-03-1999

On a note both sources have, the recovered HTML differs from SIDOF's only in escaping its accents as entities, and the Markdown built from either is identical. The page carries no codDiario, codEdicion or pagina, so the image and PDF paths stay SIDOF-only; a note the site has no HTML for comes back with existeHtml "N", and an unknown codNota with "Nota": [], as SIDOF answers.

Which source a day came from is recorded, never inferred. Every saved day carries a fuente key ("sidof" or "dof.gob.mx"), and the registry stores it next to the date, so provenance can be audited after the fact:

1999-03-06	sin-edicion
1999-03-08	dof.gob.mx
1999-03-09	sidof

In the compact --titulos dataset the marker rides along on the notes it applies to: a note carries fuente only when its day did not come from SIDOF, since repeating "sidof" on all ~1.2 million rows would cost more than it says.

What the fallback carries, and what it does not

The website's daily index lists the substantive gazette — PE, PJ, PL, OA, OD — and leaves out the three bulk-announcement groups, which are reachable on the site only through its POST search form:

CV convocatorias for public-sector procurement
VG convocatorias for civil-service vacancies
AV avisos judiciales y generales

On days both sources have, the recovered set of codNota matches SIDOF's exactly once those three are excluded (checked on days sampled from 1999 through 2026). A recovered day is therefore complete with respect to what the gazette enacted and short of what it announced, and says so in notasIncompletas rather than passing for a whole day.

The website's per-note index starts in January 1999; before that it holds only scanned images, so an older day returns an edition with no index. Those come back in edicionesSinIndice as {"codEdicion", "codDiario"}: no titles to list, but proof the gazette was published, which is what keeps the day off the empty pile. Every day confirmed lost from SIDOF so far is 1999 or later, inside the range where titles can actually be recovered.

On a page served for the wrong date. The index prints the date it is actually serving, and it has been seen answering with a different day's page. Since the parser stamps each note with the date that was asked for, taking such a page at face value would file real notes under the wrong day. Every page carrying content is therefore checked against what it claims to be, and a mismatch raises dofweb.PaginaDeOtroDia. --archivo treats that like a network error and leaves the day to retry — believing it would corrupt the day, and calling the day empty would bury it for good. Editions the gazette never ran carry no date and no content, which is not a mix-up and is not treated as one.

On TLS. www.dof.gob.mx serves its leaf certificate without the intermediate that signs it, so verification fails with "unable to get local issuer certificate" on any client that does not chase the issuer itself. The missing GoDaddy intermediate — and its root — ship in dofjson/certs/dof-gob-mx-chain.pem. The system trust store is tried first and the bundled chain only on failure. Certificate verification is never disabled.

Building a compact titulo dataset from the release (--titulos)

dofjson --titulos builds a small codNota + titulo + fecha dataset out of every note ever published, sourced from the notas-archivo GitHub release (one notas-YYYY.tgz per year, 1917 to last year, plus one notas-YYYY-MM.tgz per month of the current year). Each asset is downloaded straight into memory, its daily JSON indexes are read without ever writing them to disk, and only codNota/titulo/fecha are kept from every note (titulo is Spanish for "title", fecha for "date") — codNota to fetch that note's full content later, titulo for exploratory analysis of the titles themselves, fecha to place each title in time. The result is a single gzip-compressed JSONL file (~1.2 million notes fit in a few tens of MB): small enough to move to a Colab GPU runtime for experiments.

dofjson --titulos                    # -> titulos/titulos.jsonl.gz
dofjson --titulos --outdir /content  # e.g. from a Colab notebook
import gzip, json
with gzip.open("titulos/titulos.jsonl.gz", "rt", encoding="utf-8") as f:
    notas = [json.loads(line) for line in f]
# notas[0] == {"codNota": 4434476, "titulo": "CIRCULAR nº. 164, ...", "fecha": "23-03-1917"}

Or use the function directly:

from pathlib import Path
from dofjson.titulos import download_titulos

download_titulos(Path("titulos.jsonl.gz"))

Development

pytest packages/dofjson

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dofjson-0.2.0.tar.gz (42.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dofjson-0.2.0-py3-none-any.whl (28.4 kB view details)

Uploaded Python 3

File details

Details for the file dofjson-0.2.0.tar.gz.

File metadata

  • Download URL: dofjson-0.2.0.tar.gz
  • Upload date:
  • Size: 42.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.13

File hashes

Hashes for dofjson-0.2.0.tar.gz
Algorithm Hash digest
SHA256 1889e55a6960c3e6d60abe84dab6730a8f3f689d2aa8c9a08b31a6ba26092bc0
MD5 bcbffd4367d11666041f312e63ceeb1c
BLAKE2b-256 93092d0c58f1d95f8b25905d1e1c0e9959b21e7a44cdce798fc21d1d0a987487

See more details on using hashes here.

File details

Details for the file dofjson-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: dofjson-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 28.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.13

File hashes

Hashes for dofjson-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 74c47f4a91ffe4a5b0f1385c0a5b3b7aedc6a5b5ca4f1f3f51dfef54a4111b18
MD5 519fb8f46e1d5be6dd3529f649b2073a
BLAKE2b-256 831f7a60ca65d78fc1da19eace6d09268b26d9ba95da68fda3e4f383fabd3275

See more details on using hashes here.

Release history Release notifications | RSS feed

0.4.1

2 files

0.4.0

2 files

0.3.0

2 files

This release

0.2.0 This release

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page