Skip to main content

goldrush

Build status

goldrush is a Python implementation of the Gold Rush match key algorithm for identifying "duplicate" MARC records, or records that appear to be about the same bibliographic item. It provides a function that you pass a MARC record and get back the Gold Rush key as a string, plus a goldrush command line tool that keys the records in a MARC or MARCXML file. Both pymarc and mrrc records are supported — goldrush is duck-typed and imports neither backend itself, so you use whichever one you already parse records with.

goldrush is a class-for-class port of the canonical, production-verified Java reference implementation, coalliance-matchkey, maintained by the Colorado Alliance of Research Libraries (CoAlliance). It targets algorithm version _v08142026, whose version string is embedded in every generated key (just before the trailing format character) so you can tell which algorithm produced a given key. See docs/CoAlliance_Match_Key.md for the field-by-field specification.

The initial code was found in pymarc_dedupe created by Max Kadel and Jane Sandberg at Princeton University Library. The goldrush module was created because pymarc_dedupe included other functionality that was unrelated to Gold Rush, and pymarc_dedupe was not installable as a module via PyPI.

Usage

First install goldrush together with a MARC backend. goldrush does not depend on either pymarc or mrrc directly, so pick one via an extra:

pip install goldrush[pymarc]     # or: pip install goldrush[mrrc]

Then load a record with your backend and generate a key:

>>> from goldrush import goldrush
>>> from pymarc import MARCReader          # or: from mrrc import MARCReader
>>> record = next(MARCReader(open('marc.dat', 'rb')))
>>> goldrush(record)
'pragmaticprogrammerfromjourneymantomaster___________________________________________________________2000____1__addisa________________________________________hunt_________________v08142026p'

The same call works with an mrrc record.

The CoAlliance indexer optionally uses the MARC filename as a hint when deciding whether a record is electronic or print. To match that behaviour, pass the filename:

>>> goldrush(record, marc_filename_hint='cu-electronic-2026-04.marc')

Pass nothing (the default) for pure-MARC format detection.

Command line

Installing goldrush also installs a goldrush command for keying records straight from a file, which is handy for a quick look or for shell pipelines. It reads MARC21 (binary) or MARCXML, from files or standard input, and writes one key per line:

$ goldrush docs/examples/duplicates.mrc
pragmaticprogrammerformachinelearning[...]scuta________________v08142026p
pragmaticprogrammerformachinelearning[...]scuta________________v08142026p
pragmaticprogrammerformachinelearning[...]scuta________________v08142026e

(The keys are 188 characters wide; [...] stands in for the middle of each one here.)

So finding the duplicates within a file is a pipeline:

$ goldrush records.mrc | sort | uniq -d

Nothing but the key is printed by default, because the record's own control number is not part of the key and should not end up inside one. When you need to know which record a key came from, --id adds the 001 as a leading, tab-separated column:

$ goldrush --id docs/examples/duplicates.mrc
22866507	pragmaticprogrammerformachinelearning[...]scuta________________v08142026p
1841930199	pragmaticprogrammerformachinelearning[...]scuta________________v08142026p
22926520	pragmaticprogrammerformachinelearning[...]scuta________________v08142026e

That is the output shape of the reference MatchKeyCli, so the two can be diffed directly.

The input format is detected from the file extension and then the file's first bytes; pass --format marc or --format marcxml to say so explicitly. Records are parsed with pymarc if it is installed and mrrc otherwise (--backend picks one). Use --filename-hint TEXT to supply a format-detection hint for every record, or --hint-filenames to use each input's own filename as its hint (what the CoAlliance indexer does).

Records that cannot be parsed are reported on stderr and skipped; the exit status is 1 if anything was skipped or any input yielded no records.

Matching records

Two records match when their keys are equal. Comparing keys is all the matching there is. docs/examples/duplicates.mrc holds three real records for one 2023 CRC Press title: the print edition as cataloged by the Library of Congress, the same printed book as cataloged independently in K10plus, and the e-book as cataloged by the Library of Congress.

>>> from goldrush import goldrush
>>> from pymarc import MARCReader
>>> with open('docs/examples/duplicates.mrc', 'rb') as fh:
...     lc_print, k10_print, lc_electronic = [goldrush(r) for r in MARCReader(fh)]
>>> lc_print == k10_print
True
>>> lc_print == lc_electronic
False

The two print records match even though the catalogers disagreed about the leading article ($aA pragmatic programmer vs $aThe pragmatic programmer), whether to record an edition statement, whether to bracket the date ([2023] vs 2023), and how to punctuate the publisher. Gold Rush normalizes these details away.

The e-book does not match, and the reason is visible in the keys: they differ at exactly one character, the trailing format byte.

>>> lc_print[-10:]
'v08142026p'
>>> lc_electronic[-10:]
'v08142026e'
>>> lc_print[:-1] == lc_electronic[:-1]
True

See docs/examples/README.md for record provenance; tests/test_examples.py asserts these relationships under both backends.

Masking parts of the key

Sometimes a component of the key encodes a distinction you do not care about for the task in front of you. Because the key is a fixed-width positional string, you can ignore a component by slicing it out or blanking it before you compare. Nothing in goldrush does this for you; it is your string, and masking is ordinary Python.

The print/electronic distinction is a good worked example. Keeping print and electronic apart is usually what you want. But if you are clustering at the work level, "how many editions of this do we hold, in any carrier?", you can mask it out. The format character is the last character of the key, so dropping it is the whole recipe:

>>> def format_agnostic(key):
...     """Gold Rush key with the trailing print/electronic byte removed."""
...     return key[:-1]
...
>>> format_agnostic(lc_print) == format_agnostic(lc_electronic)
True
>>> len({format_agnostic(k) for k in (lc_print, k10_print, lc_electronic)})
1

All three example records collapse to a single masked key.

The same move works on any other component, depending on what you are doing. These are the ranges, as generated by generator.generate:

Component Range Width
title 0:95 95
GMD (disabled since 2022, always underscores) 95:100 5
publication year 100:104 4
pagination 104:108 4
edition 108:111 3
publisher 111:116 5
leader type 116:117 1
title part 117:147 30
title number 147:157 10
author 157:162 5
title dates 162:177 15
algorithm version 177:187 10
format (p/e) 187:188 1

So, blank 104:108 if you want records to match across differing pagination — common between a print record and an $a1 online resource record, and in fact the format byte alone is often not enough to collapse print and electronic for that reason. Blank 100:104 to ignore a one-year publication-date disagreement. Blank 157:162 if you are matching on title and imprint alone. Keep the width when you blank, so later components stay aligned:

>>> def work_level(key):
...     """Key with both pagination and the format byte masked out."""
...     return key[:104] + '____' + key[108:-1]
...
>>> len({work_level(k) for k in (lc_print, k10_print, lc_electronic)})
1

Masking widens matches, so it also admits false positives. The more you blank, the more genuinely different manifestations collapse together. Note also that a masked key is no longer a Gold Rush key, so don't store one where something else expects the real thing. See docs/CoAlliance_Match_Key.md for what each component is derived from.

Verifying against the reference

The reference implementation is the behavioural contract. The tests port its JUnit suite to pytest and additionally diff goldrush's output for every record in tests/marc.dat against a golden file (tests/marc.dat.keys) generated by the reference coa_matchkey_v08142026.jar. To regenerate the golden file (requires Java and marc4j 2.9.6 on the classpath):

java -Dorg.coalliance.indexing.fileName= \
     -cp coa_matchkey_v08142026.jar:marc4j-2.9.6.jar \
     org.coalliance.matchkey.cli.MatchKeyCli tests/marc.dat > tests/marc.dat.keys

Credits and licensing

goldrush is licensed under the Apache License, Version 2.0.

The Gold Rush matchKey algorithm is the work of the Colorado Alliance of Research Libraries, and goldrush's field-extraction and normalization logic (src/goldrush/fields/, src/goldrush/util/, generator.py, and version.py) is a port of their Apache-2.0 reference implementation, coalliance-matchkey (Copyright 2026 Colorado Alliance of Research Libraries). The ISBN-validation logic additionally derives from the solrmarc-marc4j project. These attributions are recorded in NOTICE.

Metadata

Release files for goldrush 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for goldrush 0.6.0
File Size Uploaded
goldrush-0.6.0.tar.gz 19.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for goldrush 0.6.0
File Interpreter ABI Platform
goldrush-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 49.9 kB

Release files / goldrush-0.6.0.tar.gz

Download URL goldrush-0.6.0.tar.gz
Size 19.2 kB
Tags Source
SHA-256 checksum
How to use checksums
6d236458c389cf99382b654c8659b943f48e4c851781817502c23782ebc2f3a7
BLAKE2b-256 checksum
How to use checksums
b8b71f47a0bb31b808aacc254220cadfa906aac8ab1efee28ea67aad22fb3ee0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / goldrush-0.6.0-py3-none-any.whl

Download URL goldrush-0.6.0-py3-none-any.whl
Size 30.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
64cfc6bafc1336b09f751db71654d0bd29cd4a9334bdb607014dc5f37d52ec82
BLAKE2b-256 checksum
How to use checksums
7bd4c5b84f95f831db75bc5fe04b47047fac0b0c6f3fb3a7cc26cdd2d1e3e04b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.2.0

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page