Skip to main content

lintle

Validate and clean Two-Line Element (TLE) satellite corpora — correctness-first.

lintle audits TLE files exported from space-track.org against the standardized TLE spec, repairs the systematic export defects, and emits a uniform, de-defected corpus that any SGP4 / orbital-mechanics library can ingest directly. Records it cannot safely repair are quarantined — never silently mangled — into a per-file sidecar detailed enough to file a defect report with space-track.

  • Correctness over recovery — every emitted record is re-validated and valid by construction; on any doubt a record is quarantined, never guessed.
  • Constant memory — streams a 3 GB file line-by-line; the whole ~30 GB corpus never loads into RAM.
  • Byte-deterministic output — same input → identical bytes every run (diff-able, CI-friendly).

On the bundled 29-file corpus (~232 M records), run with --reconstruct-checksum: 99.96 % cleaned, 0.044 % quarantined — every quarantined record fell into an anticipated defect category. (Missing-checksum reconstruction is opt-in as of 0.6.0; without the flag the 71.3 M checksumless records are quarantined instead of cleaned — see Results.)


What problem it solves

A TLE record is two fixed-width lines, each exactly 69 ASCII columns, with a mod-10 checksum in column 69. Bulk historical exports from space-track carry two systematic, era-specific defects:

  • Trailing \ artifact — almost every Line 1 has an extra \ byte appended before the newline.
  • Missing checksum digit — many records were exported without their column-69 checksum, leaving 68-column lines.

These appear independently and in combination, and a small fraction of records are genuinely corrupt (garbled columns, orphaned lines, wrong lengths). lintle distinguishes the safely-repairable from the genuinely-corrupt and treats each correctly.

Installation

Requires Python 3.14+ and uv. The only runtime dependency is rich (>=15,<16, terminal rendering for the clean progress UI); everything else is standard library. (sgp4 is a dev-only test oracle.)

uv sync

No build step is needed to run the tool.

Usage

The console script is lintle (python -m lintle … is equivalent):

# Produce cleaned output + quarantine sidecars
uv run lintle clean [path]

# Audit a clean run's output for corruption and contradictions
uv run lintle verify [out-dir] --source [src-dir]

# Emit a de-duplicated "latest re-issue only" import list
uv run lintle dedup [out-dir]

# Extract one satellite's complete deduped TLE history + stats sidecar
uv run lintle extract <NORAD-ID>... [--out-dir DIR] [--dest DIR]

# Re-render a prior clean run's aggregate summary from its report.json
uv run lintle report [out-dir]

# Explain a rule ID or fix tag — definition, examples, source citation
uv run lintle explain <TAG>

# Compare two clean runs' findings (per-rule deltas)
uv run lintle diff <run-a> <run-b>

Arguments and options:

Option Default Meaning
path data/source A single file or directory. A directory is globbed for tle*.txt (tool output *.cleaned.txt / *.broken.txt is excluded).
--out-dir DIR data/output Where clean writes its output. Created if absent.
--jobs N CPU count − 1 Files processed in parallel. Lower it if a slow disk causes I/O contention.
--report text|json text Summary format.
--max-quarantined N[%] 0 Exit non-zero only if MORE than N records were quarantined; or, with a trailing %, more than N% of routed records (clean + quarantined). Default 0 ≡ "any quarantine fails".
--resume / --no-resume (clean only) Resume an interrupted run without prompting / ignore any checkpoint and start fresh. See Cancelling and resuming.
--reconstruct-checksum off (clean only) Opt in to tier-2 missing-checksum reconstruction: recompute and append a dropped column-69 checksum instead of quarantining the 68-char line. Off by default because a dropped trailing data character is indistinguishable from a dropped checksum. Part of the resume run-identity, so toggling it forces a fresh run.

Examples:

# Clean the whole corpus
uv run lintle clean data/source

# Clean one file to a custom location
uv run lintle clean data/source/tle2022.txt --out-dir data/output

# Clean the corpus, capture a machine-readable summary
uv run lintle clean data/source --report json > run-summary.json

# CI gate: fail only if more than 100 records (or 1% of routed records) are quarantined
uv run lintle clean data/source --max-quarantined 100 --report json > run-summary.json
uv run lintle clean data/source --max-quarantined 1%  --report json > run-summary.json

# Look up what a rule ID or fix tag means, with a verified example
uv run lintle explain TLE-CHK-001
uv run lintle explain reconstructed-checksum

Exit codes:

Code Meaning
0 Quarantine count (or rate) is at or below --max-quarantined (default 0).
1 Quarantine count (or rate) exceeded --max-quarantined.
2 Operational error — no input files, disk shortfall, lock held, stale/corrupt/declined resume, or a file that failed to process.
129 / 130 / 143 Killed by SIGHUP / Ctrl-C (SIGINT) / SIGTERM.

Repairable defects (including the near-universal trailing \) do not raise the exit code above 0 — almost every raw file contains them. --max-quarantined preserves the meaningful 2 (operational error) and 130 (Ctrl-C) signals that a lintle … || true pipe would swallow.

Auditing a clean run — lintle verify

clean promises byte-faithful, always-valid output; verify is the independent second opinion that checks the promise held. It reads a finished run's <out-dir>/01-cleaned/ (never mutating it) and runs three sgp4-free checks:

  • re-validate — every cleaned record must still pass the one validator (tle.py);
  • contradiction — no two records may share a (catalog, epoch) and an element-set number yet carry a different orbital state (benign same-epoch re-issues, which carry a new element-set number, are counted in a census, not flagged);
  • source byte-diff — with --source, every cleaned line must be a sanctioned edge edit of a real source line; any interior-column change is corruption.

It can also run an opt-in --orbit physics pass (goal 2): for a sample of satellites it propagates each cleaned TLE forward to its neighbour's epoch with sgp4 and flags position-residual outliers across the track. An outlier is inconclusive (soft) — a real manoeuvre and a corruption look alike from one pair — so it never fails the run; only sgp4 rejecting an element set as physically impossible is hard (and ~never happens on cleaned data).

# Audit the default run against its source
uv run lintle verify data/output --source data/source

# Structural checks only (re-validate + contradiction), no source needed
uv run lintle verify data/output --no-source-diff

# Add the sampled sgp4 orbit-consistency pass (3000 satellites by default)
uv run lintle verify data/output --orbit --all
Option Default Meaning
out-dir stored config, else data/output The finished clean run to audit (reads <out-dir>/01-cleaned/).
--source DIR stored config, else data/source Original source tree for the byte-diff (goal 1).
--no-source-diff off Skip the byte-diff; re-validate and check contradictions only.
--orbit off Run the sampled sgp4 orbit-consistency pass (goal 2).
--sample N 3000 (--orbit) satellites to sample; ignored with --all.
--all off (--orbit) check every satellite, not a sample.

The suspects report is written under <out-dir>/04-verify/ (suspects.jsonl + summary.{json,md}). Exit codes: 0 clean, 1 at least one hard suspect (a re-validation failure, an element-set contradiction, an interior mutation, or an sgp4-unphysical element set), 2 operational error (no cleaned output to audit). Orbit-residual outliers are soft (inconclusive) and never raise the exit code. summary.json/summary.md also carry an informational epoch_distribution — a {"YYYY-MM": count} record-density histogram over valid records — that never affects the exit code or the suspect list.

A "latest only" import list — lintle dedup

The 01-cleaned/ archive faithfully keeps every published record — including the near-identical re-issues space-track emits for the same satellite at the same epoch (same orbit, just a bumped element-set or revolution number). A consumer that wants "the orbit for satellite X at time T" then has to pick one. lintle dedup produces that pick as a separate artifact — 01-cleaned/ is never touched:

  • one card per (catalog, epoch), keeping the latest re-issue (highest element-set number);
  • benign re-issues (identical orbital state) collapse silently;
  • a genuine contradiction (same satellite + epoch, a different orbit) is kept-latest and flagged — never resolved in silence;
  • if a verify run's suspects.jsonl exists, hard suspects are excluded from the list first.
# De-duplicate the default run -> data/output/05-dedup/import.txt
uv run lintle dedup data/output

Writes <out-dir>/05-dedup/: import.NNNNN.txt (the chunked ingest list, sorted by (catalog, epoch)), notes.NNNNN.jsonl (one entry per collapsed group — kept vs dropped), manifest.jsonl (one row per satellite — record count, epoch span, median spacing, largest gap — a single plain file, never chunked; the substrate for jq | shuf | xargs lintle extract coverage queries), and summary.json (counts). This realises the output tiering source/01-cleaned/ (immutable archive) → 05-dedup/import.txt (verified + deduped, ready to ingest). Exit codes: 0 clean, 1 a genuine contradiction was arbitrated (review notes.jsonl), 2 no cleaned output.

One satellite's history — lintle extract

lintle extract <NORAD-ID>... binary-searches a prior dedup run's sorted import set and writes each satellite's complete deduped history as <id>.txt (pure TLE lines, epoch-ascending) plus a <id>.json stats sidecar. With no --dest, output goes to <out-dir>/06-extract/ (which also gets a README); an explicit --dest DIR writes there instead and is never decorated with a README — it's the user's own directory. extract also warns (never fails) if 01-cleaned/ has changed since the dedup run it's reading from, by comparing a cheap stat-only fingerprint stored in 05-dedup/summary.json.

Interactive wizard & remembered paths — .lintle.json

Run lintle with no subcommand on a terminal and it opens a small menu — configure, clean, verify, report, quit — instead of printing help. The source and output directories you pick are remembered in a project-local ./.lintle.json so the everyday commands run without repeating paths. Precedence is always explicit CLI arg → stored config → built-in default, and a stored path that has since vanished is re-prompted. Off a TTY (scripts, CI, pipes) a bare lintle keeps the old behaviour — prints help and exits 2 — so nothing that pipes lintle ever blocks on a prompt.

Correctness guarantees & limits

This is the heart of the tool. The cleaner never applies a fix and hopes: it applies a candidate fix, re-runs the full validator, and commits only if the result passes — so the output cannot contain a wrong-but-valid-looking record. One validator (tle.py) defines what "perfect" means; clean checks every candidate repair against it before committing — so correctness is structural, not assumed.

lintle never invents data. The single sanctioned reconstruction is the column-69 checksum — safe only because it is a deterministic mod-10 function of columns 1–68, so recomputing a missing one asserts nothing the record didn't already say (the redundancy paradox: the only field safe to rebuild is the one that was redundant to begin with). A mod-10 checksum accepts a wrong line 1-in-10 times by luck, so guessing an orbital-data character risks a record that looks valid but is silently wrong — the one outcome worse than dropping it. So anything requiring such a guess (bad checksum, wrong length, orphan line, garbled columns) is quarantined, not repaired.

Even the checksum recompute is opt-in as of 0.6.0 (--reconstruct-checksum): by default a checksumless 68-char line is quarantined, because a dropped trailing data character is indistinguishable from a dropped checksum, so reconstructing one by default could silently emit wrong-but-valid data.

Fixes fall into five classes in decreasing order of safety — content-preserving (trailing \, CRLF, trailing whitespace), reconstructed-checksum, content-shifting (leading trim), structural (drop blanks), and corrupt (quarantine).

→ Full fix-class table, repair tiers, and the stable rule registry: ARCHITECTURE.md §1 and §4.

Output

clean (and, optionally, verify/dedup/extract after it) lays --out-dir out as one flat level of directories, numbered in pipeline order:

<out-dir>/
├── README.md          — overview: the six dirs, order of operations, regen note
├── 01-cleaned/        — tleYYYY.00001.cleaned.txt, .00002… chunk set per input file  + README.md
├── 02-broken/         — tleYYYY.00001.broken.txt, .00002…  chunk set per input file  + README.md
├── 03-report/         — report.md · report.json · report.00001.jsonl ·
│                        broken-noradids.ndjson                                       + README.md
├── 04-verify/         — `lintle verify` output (suspects.NNNNN.jsonl, summary.{json,md}) + README.md
├── 05-dedup/          — `lintle dedup` output (import.NNNNN.txt, notes.NNNNN.jsonl,
│                        manifest.jsonl, summary.json)                                + README.md
└── 06-extract/        — `lintle extract` output (<id>.txt, <id>.json)                + README.md

The out-dir root has one numbered directory per pipeline step — 01-cleaned, 02-broken, and 03-report (written by clean), 04-verify (written by verify), 05-dedup (written by dedup), and 06-extract (written by extract, its default destination) — plus a self-describing root README.md. Every populated step directory also carries its own static README.md. This is a breaking layout change as of 0.11.0: the 0.10.1 data/ grouping is retired, and outputs from ≤ 0.10.3 must be regenerated (a fresh clean run scrubs both legacy layouts). Every record/line output stream is split into a <stem>.NNNNN.<suffix> chunk set of --chunk-records records each (default 1,000,000 ≈ 140 MB, so no single output file is ever huge), tunable via --chunk-records N on clean/dedup/verify (0 = a single .00001 chunk). cat <stem>.*.<suffix> reproduces the pre-chunking single file byte-for-byte. Aggregate summaries (report.md, report.json, broken-noradids.ndjson, */summary.*) stay single files.

  • 01-cleaned/tleYYYY.NNNNN.cleaned.txt — standard 2-line TLE text, every record verified valid and ready for downstream ingestion.
  • 02-broken/tleYYYY.NNNNN.broken.txt — the quarantine sidecar: source line number(s), a human-readable reason, and the offending line(s) copied byte-faithfully, with a header formatted to paste into a space-track defect report.
  • broken-noradids.ndjson — one {"noradId":N} per line, the deduplicated, sorted set of NORAD catalog numbers quarantined anywhere in the run (for programmatic consumers).
  • report.md — human-readable run report: corpus totals, % cleaned/quarantined, fix counts, the per-rule defect breakdown, a per-file table, a per-NORAD breakdown, and — when any input file failed to process — a ## Failures table naming each failed file and its error.
  • report.json — the machine-readable run envelope, byte-identical to the --report json stdout output. Persisted on every clean run so lintle report can re-render the summary later without re-processing the corpus.

At the end of a clean run an aggregate summary panel is rendered to stderr — corpus totals, % cleaned/quarantined, and the top fix / quarantine rules — sized to the terminal width (with an ASCII-bar fallback off a TTY). Text-mode stdout stays empty; the full machine summary is report.json (or --report json on stdout). records counts paired 2-line entries; clean are those that passed and were written; quarantined is everything routed to 02-broken/ (failed records and every orphan line). The invariant is records + orphan == clean + quarantined. Defects key by the stable RuleID registry (TLE-CHK-001, TLE-PAIR-001, …) so one identifier names a defect across every artifact.

lintle report [out-dir] re-renders that panel to stdout from a prior run's report.json (or echoes the JSON verbatim with --report json); a missing or unreadable report.json exits 2.

Live progress during a long run is also written to stderr (so it never pollutes the stdout --report json pipe): a size roster up front, per-file byte/record progress with throughput and ETA, and an [k/N] line as each file finishes.

→ Machine-readable contracts (--report json envelope, report.jsonl, the .broken.txt format, the checkpoint): ARCHITECTURE.md §6.

Results on the bundled corpus

A full run over the 29-file corpus (tle2004tle2025, ~232 million records), with --reconstruct-checksum:

  • 99.96 % cleaned — 187.9 M trailing-\ artifacts stripped, 71.3 M missing checksums reconstructed.
  • 0.044 % quarantined (103,228 records) as genuinely corrupt — every quarantined record fell into an anticipated category; no unknown defect type surfaced.

Since 0.6.0 the missing-checksum recompute is opt-in (default off). A default run over the same corpus quarantines those 71.3 M checksumless records instead of reconstructing them, so the default-mode cleaned rate is correspondingly lower — pass --reconstruct-checksum to reproduce the figures above.

Operational notes

Cancelling and resuming

A long clean can be interrupted (Ctrl-C, a closed laptop, SIGTERM/SIGHUP). Re-run the same command (same --out-dir, unchanged inputs) to resume; on a TTY it prompts, in CI it auto-resumes. Resume granularity is a whole file: completed files are skipped and the file in flight at the interruption restarts from the beginning — so a multi-file corpus run benefits, but a single-file run gains nothing. --no-resume discards the checkpoint and starts fresh (clearing prior outputs).

→ Checkpoint shape and the resume-decision matrix: ARCHITECTURE.md §5.

Disk space

Every record is routed to exactly one of 01-cleaned/ or 02-broken/ — never duplicated — so the output is roughly the input's size plus tiny metadata. As a guard, lintle requires ~2× the total input size free on the --out-dir volume before starting, aborting with exit 2 if short (and warning on stderr in the 2×–2.5× borderline band). Rule of thumb for the ~30 GB corpus: keep ~60 GB free to clear the abort floor, ~75 GB to clear the warning. (The 12 GB TLEs.zip is not an input and is never read.)

Development

uv sync                          # install dev dependencies
uv run pytest                    # run the test suite
uv run pytest --cov=lintle       # with a coverage report
uv run ruff check                # lint
uv run ruff format               # auto-format

The suite includes per-module unit tests, an asymmetric cross-check against the trusted sgp4 parser (a known-good TLE must be accepted by both), and end-to-end integration tests (golden output, idempotence, re-validation). See CONTRIBUTING.md for setup, testing, and the git workflow.

Further reading

ARCHITECTURE.md is the living design reference — the validator definition, the module map and data flow, the repair model, streaming/durability/resume, the machine-readable output-format contracts, and the runtime-dependency policy. Dated design specs, plans, and corpus-run summaries are kept for historical rationale under docs/superpowers/archive/.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

lintle-0.12.0.tar.gz (727.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

lintle-0.12.0-py3-none-any.whl (149.3 kB view details)

Uploaded Python 3

File details

Details for the file lintle-0.12.0.tar.gz.

File metadata

  • Download URL: lintle-0.12.0.tar.gz
  • Upload date:
  • Size: 727.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for lintle-0.12.0.tar.gz
Algorithm Hash digest
SHA256 aed51e1986a9017949a1c2d8bbf621c2938267ea6ee476a2748bb97221e5e034
MD5 9eb77e15f0e41b1f9265d34846d6ea9d
BLAKE2b-256 f17bc6f212d495791ca1041a6dd5c9d7e864284f11f77dcc8ae33cb97f28fb4a

See more details on using hashes here.

File details

Details for the file lintle-0.12.0-py3-none-any.whl.

File metadata

  • Download URL: lintle-0.12.0-py3-none-any.whl
  • Upload date:
  • Size: 149.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for lintle-0.12.0-py3-none-any.whl
Algorithm Hash digest
SHA256 091bb2803b643c5e6d4d69fd3371de713bfa73dfea1036be5cd5f5a0dc7b312e
MD5 48b33abca4a678fa245ce53e765f63ac
BLAKE2b-256 03063e8673e84cbb98e88a104de94ca587eedfb46a06ac97d19685955e351a1e

See more details on using hashes here.

Release history Release notifications | RSS feed

0.14.1

2 files

0.14.0

2 files

0.13.6

2 files

This release

0.12.0 This release

2 files

0.11.1

2 files

0.11.0

2 files

0.10.3

2 files

0.10.2

2 files

0.10.1

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.1

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.1

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page