jsonl-diff
jsonl-diff performs strict, identity-based reconciliation of JSONL/NDJSON
datasets. It matches records by one or more top-level fields, then classifies
each identity as equal, added, deleted, or modified.
Why jsonl-diff?
A text diff compares lines. That is usually the wrong model for datasets: reordering unchanged records creates noise, and inserting one record can make every later line appear different.
jsonl-diff fills the gap between line-oriented tools and in-memory dataframe comparisons:
- match records by identity, not physical position;
- compare datasets larger than available RAM using temporary SQLite storage;
- preserve original OLD and NEW line numbers for auditability;
- expose the same behavior through a CLI and a small Python API.
Use it for snapshots, ETL validation, migrations, exports, and CI checks. Use
diff, jq, or a visual JSON diff instead when you need line-level,
field-level, or JSON Patch output.
The comparison is backed by a temporary SQLite index. Complete inputs and complete result sets do not need to fit in memory. Process memory is bounded by the largest record plus parser and SQLite buffers, while indexed data lives on temporary disk.
Features
- Single or composite top-level identities, independent of input order.
- Configurable duplicate handling with strict failure by default.
- Exact RFC 6901 object-member ignores.
- Semantic number comparison without binary floating-point rounding.
- Deterministic summaries and changed-identity iteration.
- Original OLD and NEW physical line numbers for every change.
- Local, HTTP/HTTPS, file-like, and supported compressed sources through
py-jsonl. - The same comparison engine through the CLI and Python API.
- Disk-backed comparison with configurable
jsonl-difftemporary storage.
Requirements and installation
jsonl-diff supports Python 3.8 through 3.14.
Install it in a project managed by uv:
uv add jsonl-diff
Install the CLI as a standalone tool:
uv tool install jsonl-diff
pip is also supported:
python -m pip install jsonl-diff
CLI quick start
Given:
{"id":1,"name":"Ada"}
{"id":2,"name":"Grace"}
and a reordered, changed dataset:
{"id":2,"name":"Grace Hopper"}
{"name":"Ada","id":1}
{"id":3,"name":"Linus"}
compare them by id:
jsonl-diff old.jsonl new.jsonl --key id
Standard output contains only the text summary:
Records:
equal: 1
added: 1
deleted: 0
modified: 1
The result means one record is unchanged, one was added, and one existing record changed. Physical order does not affect these counts.
Diagnostics are written to standard error. Use --quiet when only the exit
status or details file is needed.
Common pipelines
Filter JSONL with jq before comparing it. -c keeps one compact JSON object
per line:
jq -c 'select(has("deleted_at") and .deleted_at == null)' old.jsonl |
jsonl-diff - new.jsonl --key id
Read compressed input directly when no preprocessing is required:
jsonl-diff old.jsonl.xz new.jsonl.gz --key id
For two independent pipelines on Bash-compatible systems, use process substitution:
jsonl-diff \
<(head -n 1000 old.jsonl | jq -c 'select(.active)') \
<(head -n 1000 new.jsonl | jq -c 'select(.active)') \
--key id
- represents stdin. Both inputs cannot use the same stdin stream.
CLI syntax and options
jsonl-diff [-h] --key KEY [--ignore IGNORE]
[--duplicates {error,first,last}] [--details FILE] [--quiet]
[--max-temp MAX_TEMP]
old new
| Argument | Meaning |
|---|---|
old |
OLD local path, HTTP/HTTPS source, or - for stdin. |
new |
NEW local path, HTTP/HTTPS source, or - for stdin. |
--key KEY |
Required top-level identity field. Repeat it or use a comma-separated value for a composite identity. |
--ignore JSON_POINTER |
Exact RFC 6901 pointer to an object member to remove before content comparison. Repeatable. |
--duplicates POLICY |
Handle repeated identities with error (default), first, or last. |
--details FILE |
Write machine-readable JSONL details to FILE; details are never written to stdout. |
--quiet |
Suppress the normal stdout summary. Errors still go to stderr. |
--max-temp MAX_TEMP |
Limit storage owned by jsonl-diff to a positive integer number of bytes. |
-h, --help |
Show command help and exit. |
These composite-key forms are equivalent:
jsonl-diff old.jsonl new.jsonl --key country,customerId,type
jsonl-diff old.jsonl new.jsonl \
--key country \
--key customerId \
--key type
Ignore volatile fields by exact path:
jsonl-diff old.jsonl new.jsonl \
--key id \
--ignore /updated_at \
--ignore /metadata/request_id
Exit codes
| Code | Meaning |
|---|---|
0 |
The inputs are equal under the configured rules. |
1 |
At least one added, deleted, or modified identity was found. |
2 |
A handled input, source, resource, output, or interruption error occurred. |
3 |
CLI usage or comparison configuration is invalid. |
Invalid argparse usage exits directly with code 3. The Python API uses
exceptions instead of these process exit codes.
Details JSONL
--details FILE streams a deterministic JSONL report through py-jsonl after
both inputs have been completely indexed and validated:
jsonl-diff old.jsonl new.jsonl --key id --details changes.jsonl
For an OLD source containing identities 1 and 2, and a NEW source
containing changed identity 2 and added identity 3, the file is:
{"duplicates":"error","ignore":[],"key":["id"],"type":"meta"}
{"key":[1],"old_line":1,"op":"deleted","type":"change"}
{"key":[2],"new_line":1,"old_line":2,"op":"modified","type":"change"}
{"key":[3],"new_line":2,"op":"added","type":"change"}
{"added":1,"deleted":1,"equal":0,"modified":1,"type":"summary"}
Record types are:
meta: always first.keycontains normalized identity-field names;ignorecontains the configured ignore pointers.change: one per changed identity.opisadded,deleted, ormodified.old_lineis present for deletions and modifications;new_lineis present for additions and modifications. Equal records are not emitted.summary: always last after a successful write, with all four totals.
The normal stdout summary is unchanged when --details is used, unless
--quiet suppresses it. A write failure can leave a partial details file and
returns exit code 2.
Python API
Comparing sources
from jsonl_diff import ChangeOperation, diff
with diff(
"old.jsonl.gz",
"new.jsonl.gz",
key=("country", "customerId"),
ignore=("/updated_at",),
duplicates="error",
max_temp=2_000_000_000,
) as result:
print(result.summary)
for change in result.changes(ChangeOperation.MODIFIED):
print(change.key, change.old_line, change.new_line)
The callable signature is:
from typing import Any, Optional, Sequence, Union
from jsonl_diff import DiffResult, DuplicatePolicy
def diff(
old: Any,
new: Any,
*,
key: Union[str, Sequence[str]],
ignore: Sequence[str] = (),
duplicates: Union[str, DuplicatePolicy] = DuplicatePolicy.ERROR,
max_temp: Optional[int] = None,
) -> DiffResult:
...
diff() creates a comparison session without opening either source. Entering
its context fully reads, indexes, and validates both sources before exposing
the result. DiffResult.changes() is then a lazy iterator over the disk-backed
result rather than a list held in memory.
Using DiffResult as a context manager is required. It owns the temporary
resources and removes its private workspace on exit. A result cannot be entered
more than once. Its summary becomes available after entering and remains
available after closing; iterating changes requires the context to remain open.
Result models
result.summary is an immutable Summary:
Summary(equal=10, added=2, deleted=1, modified=3)
It exposes integer fields equal, added, deleted, and modified, plus the
boolean property different. The same values are available directly as
read-only DiffResult properties.
Each item from result.changes() is an immutable Change with:
operation: aChangeOperationenum value (ADDED,DELETED, orMODIFIED);key: a typed tuple in normalized identity-field order;oldandnew: optionalSourceLocationvalues;old_lineandnew_line: convenience properties returning a one-based physical line orNone.
For additions, old/old_line are None; for deletions,
new/new_line are None; modifications have both locations. Location
sources are labeled OLD and NEW.
Pass a ChangeOperation to filter without materializing all changes:
with diff("old.jsonl", "new.jsonl", key="id") as result:
added = result.changes(ChangeOperation.ADDED)
for change in added:
print(change.key, change.new_line)
With numeric keys, API key components are decimal.Decimal values. For
example, JSON identity 7 is returned as Decimal("7"); JSON strings and
booleans retain their types. Details JSONL writes numeric keys as JSON numbers,
including arbitrary-precision integers.
API errors
The public error hierarchy starts with JsonlDiffError:
ConfigurationError: invalid keys, ignore pointers, ormax_temp;InputError: an invalid source or record; exposessourceand optionalline;DuplicateKeyError: anInputErrorwithkeyand the first/repeated physical lines inlines;ResourceError: the temporary index cannot be created, written, or kept within its configured limit.
Failures during diff() clean up the workspace before the exception is
raised.
Sources and compression
Source opening and decompression are delegated to py-jsonl:
| Source | CLI | Python API |
|---|---|---|
| Local path | Yes | Yes, including path-like values |
| HTTP/HTTPS URL | Yes | Yes |
| File-like object | No | Yes |
gzip (.gz) |
Yes | Yes |
bzip2 (.bz2) |
Yes | Yes |
xz (.xz) |
Yes | Yes |
Zstandard (.zst) |
Python 3.14 only | Python 3.14 only |
| ZIP archive | No | No |
Zstandard availability follows py-jsonl and its use of Python 3.14's
standard-library zstd support; it is not supported by this project on earlier
Python versions. Compression and HTTP transfer boundaries do not affect the
reported decoded JSONL line numbers.
Comparison semantics
Strict JSONL records
- Every physical line must be non-blank, valid JSON containing exactly one top-level object.
- Local paths and binary streams are decoded as UTF-8. Text streams are
already decoded, while URL response charset handling follows
py-jsonl(defaulting to UTF-8 when no charset is declared). - Blank lines, malformed JSON, duplicate object property names,
NaN,Infinity, and-Infinityare errors. - Top-level arrays, strings, numbers, booleans, and
nullare rejected. - Input errors identify
OLDorNEWand include the physical line when available.
No normal summary or details output begins until both inputs pass complete input and duplicate validation.
Identity and duplicates
Identity fields are top-level object-member names. Each component must exist and be a non-null string, number, or boolean; objects and arrays are invalid. Nested identity paths and automatic key detection are not supported.
Composite identity field names are sorted lexically before their values are
extracted. Consequently, key=("b", "a") and key=("a", "b") both produce
keys in (a, b) order and match identically. Repeated or empty identity-field
names are configuration errors.
JSON types remain significant: "1" is different from 1, and true is
different from 1.
Every normalized identity must be unique within OLD and within NEW by default.
The first duplicate aborts the comparison; DuplicateKeyError reports the
first occurrence and repeated physical line. Set duplicates="first" to keep
the first occurrence or duplicates="last" to keep the last occurrence and
its physical line. These tolerant policies depend on physical input order.
Ignored object members
Ignore expressions are exact
RFC 6901 JSON Pointers. Escapes such
as ~1 for / and ~0 for ~ are supported. An absent path is harmless,
and /updated_at does not match /metadata/updated_at.
Ignores are applied symmetrically before content canonicalization. A pointer may remove an entire array-valued object member, but it may not traverse an array or address an array element. Ignoring a top-level identity field is also invalid. Wildcards, JSONPath, and recursive name matching are not supported.
Canonical content and numbers
Records with the same identity are compared after ignored members are removed:
- JSON object property order is insignificant at every depth.
- Array order is significant.
- Unicode strings are compared by code-point sequence; no Unicode normalization is applied.
- Numbers are parsed as
Decimaland compared by mathematical value.1,1.0, and1e0are equal, and negative zero equals zero. - Arbitrary-precision integers and decimal values are not rounded through binary floating point.
Canonical content is stored with its length and SHA-256 digest. Matching canonical bytes are still compared, so a digest match alone is not treated as proof of equality. Numbers are written using compact scientific notation; a large exponent does not expand into a large string of zeroes.
Determinism
Summary counts do not depend on either source's physical record order. Changes, including details records, use the stable lexical byte order of each canonical typed identity. This is a reproducibility guarantee, not numeric, locale-aware, or human-oriented sorting.
Disk-backed architecture and temporary files
Each record is parsed and validated incrementally, normalized, and inserted into a private SQLite database under an operating-system temporary directory. The database stores the typed canonical identity, original line, canonical content, content length, and digest. A uniqueness constraint detects duplicate identities. SQL joins calculate the summary, and ordered SQLite cursors drive lazy change iteration.
This architecture bounds memory by the records currently being processed and
database buffers; it does not keep the complete decoded inputs or all changes
in RAM. It does require temporary disk space. Canonical records and SQLite
indexes remain until the DiffResult is closed, so temporary usage can exceed
twice the combined decompressed input size.
max_temp / --max-temp accepts a positive byte count. The implementation
limits and checks files in the workspace owned by jsonl-diff, raising
ResourceError (CLI exit 2) when the index cannot stay within that budget.
Choose a limit with room for SQLite pages and index overhead.
py-jsonl may create its own temporary staging files for remote or compressed
sources. Those files follow py-jsonl's resource policy and are not counted by
jsonl-diff's max_temp limit. Total system temporary usage can therefore
exceed the configured value.
The private workspace is removed on context-manager exit, explicit close(),
or a handled failure during construction. Cleanup failures are emitted as
warnings rather than replacing the primary error.
Limitations and non-goals
jsonl-diff intentionally does not provide:
- line- or physical-order comparison;
- equal-record events or complete record payloads in details output;
- field-level diffs, JSON Patch, move detection, or rename detection;
- nested identities, fuzzy matching, or key autodetection;
- numeric tolerances, unordered-array comparison, or Unicode normalization;
- wildcard/JSONPath ignores or ignores that traverse arrays;
- schema validation or repair of malformed input;
- ZIP input, database-table input, cloud-provider SDK integrations, GUI, or HTML reports.
It is a reconciliation tool for strict JSONL datasets, not a general-purpose visual JSON diff.
Development
Clone the repository and sync the test and lint dependency groups with uv:
uv sync --group test --group lint
Run the test suite and Ruff:
uv run pytest
uv run ruff check --quiet --output-format=concise .
The tests cover CLI behavior, identity and canonicalization rules, strict validation, details output, source types, compression, temporary limits, and resource cleanup.
License
See LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file jsonl_diff-0.1.0.tar.gz.
File metadata
- Download URL: jsonl_diff-0.1.0.tar.gz
- Upload date:
- Size: 24.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
248d60c408dbfb00dab5e179ee95bdfe293e9d3cd7cf6530a55a102722f27d7d
|
|
| MD5 |
9eee06acdbbdc489a397d98a6fcc8a5d
|
|
| BLAKE2b-256 |
f260df78dec780b5c5f24265baad36b7e511cb320285ec32fd668400633654b4
|
File details
Details for the file jsonl_diff-0.1.0-py3-none-any.whl.
File metadata
- Download URL: jsonl_diff-0.1.0-py3-none-any.whl
- Upload date:
- Size: 16.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
795745887d316068290d5455ac613b869b418a0ca22090ccf23d71756ae5b0b7
|
|
| MD5 |
1808e48bf86846f6fbbeb75214cdcbdb
|
|
| BLAKE2b-256 |
14028d7a86af21a1cac1b0856f9cce05bd946aafc370f257dcce36f0f9704e53
|