Fast, explainable book metadata matching with pluggable sources
Project description
book-match
Fast, explainable book metadata matching with pluggable sources.
Features
- Fast: RapidFuzz-powered string matching (8x faster than alternatives)
- Explainable: Human-readable explanations for every match decision
- Pluggable Sources: Query Google Books, OpenLibrary, or add your own
- Batch Processing: Deduplicate thousands of books with blocking strategies
- Domain-Aware: Understands book titles, authors, ISBNs, editions
- Type-Safe: Full type hints with runtime validation
Installation
# Core library (matching only)
pip install book-match
# With metadata source integrations
pip install book-match[sources]
# Everything
pip install book-match[all]
Quick Start
Basic Matching
from book_match import Book, BookMatcher
# Create books to compare
local_book = Book(
title="The Great Gatsby",
authors=("F. Scott Fitzgerald",),
year=1925,
)
remote_book = Book(
title="Great Gatsby, The: A Novel",
authors=("Fitzgerald, F. Scott",),
isbn_13="9780743273565",
year=1925,
)
# Match them
matcher = BookMatcher()
result = matcher.match(local_book, remote_book)
print(f"Confidence: {result.confidence:.0%}") # 94%
print(f"Verdict: {result.verdict}") # MatchVerdict.AUTO_ACCEPT
print(f"Explanation: {result.explanation}")
# "Strong match (94% confidence). Title excellent match after subtitle
# handling: "The Great Gatsby" ↔ "Great Gatsby, The: A Novel". Author
# excellent match: "F. Scott Fitzgerald" ↔ "Fitzgerald, F. Scott"..."
With Metadata Sources
import asyncio
from book_match import Book, BookResolver, OpenLibrarySource, GoogleBooksSource
async def find_book():
# Create resolver with multiple sources
async with BookResolver(
sources=[OpenLibrarySource(), GoogleBooksSource()]
) as resolver:
# Your book with partial metadata
my_book = Book(
title="Dune",
authors=("Frank Herbert",),
)
# Find matches across all sources
matches = await resolver.resolve(my_book, min_confidence=0.8)
for match in matches[:3]:
print(f"{match.confidence:.0%}: {match.remote_book.title}")
print(f" ISBN: {match.remote_book.isbn_13}")
print(f" Source: {match.remote_book.source}")
asyncio.run(find_book())
Batch Deduplication
from book_match import Book, BatchMatcher
# Your library of books
books = [
Book(title="The Hobbit", authors=("J.R.R. Tolkien",)),
Book(title="Hobbit, The", authors=("Tolkien, J. R. R.",)),
Book(title="Lord of the Rings", authors=("Tolkien",)),
# ... thousands more
]
# Find duplicates
batch = BatchMatcher()
for duplicate in batch.deduplicate(books):
print(f"Duplicate found ({duplicate.confidence:.0%}):")
print(f" {duplicate.local_book.title}")
print(f" {duplicate.remote_book.title}")
Core Concepts
Books
The Book dataclass represents book metadata:
from book_match import Book
book = Book(
title="The Lord of the Rings: The Fellowship of the Ring",
authors=("J.R.R. Tolkien",),
isbn_10="0618346252",
isbn_13="9780618346257",
language="en",
year=1954,
publisher="Houghton Mifflin",
source="google_books", # Where this data came from
source_id="abc123", # ID in that source
)
Match Results
Every match returns detailed, explainable results:
result = matcher.match(book1, book2)
# Overall confidence (0.0 to 1.0)
result.confidence # 0.87
# Verdict: AUTO_ACCEPT, REVIEW, or REJECT
result.verdict # MatchVerdict.REVIEW
# Human-readable explanation
result.explanation # "Possible match (87% confidence)..."
# Individual factors
for factor in result.factors:
print(f"{factor.name}: {factor.similarity:.0%}")
print(f" Weight: {factor.weight}")
print(f" Details: {factor.details}")
Configuration
Customize matching behavior:
from book_match import BookMatcher, MatchConfig
# Strict matching (fewer false positives)
strict_matcher = BookMatcher(MatchConfig.strict())
# Lenient matching (fewer false negatives)
lenient_matcher = BookMatcher(MatchConfig.lenient())
# Custom configuration
custom_config = MatchConfig(
title_weight=0.6,
author_weight=0.3,
auto_accept_threshold=0.95,
strip_subtitles=True,
)
ISBN Handling
Proper ISBN validation with checksums:
from book_match import (
is_valid_isbn,
validate_isbn,
isbn10_to_isbn13,
normalize_isbn,
extract_isbns,
)
# Validation (checks checksum!)
is_valid_isbn("9780306406157") # True
is_valid_isbn("1234567890") # False (invalid checksum)
# Conversion
isbn10_to_isbn13("0306406152") # "9780306406157"
# Extract from text
text = "ISBN: 978-0-306-40615-7 and also 0-306-40615-2"
extract_isbns(text) # ["9780306406157", "0306406152"]
Custom Metadata Sources
Add your own sources:
from book_match import BaseSource, Book, SearchQuery
class MyLibrarySource(BaseSource):
@property
def name(self) -> str:
return "my_library"
async def search(self, query: SearchQuery, limit: int = 10) -> list[Book]:
# Query your database/API
results = await my_api.search(
title=query.title,
author=query.authors[0] if query.authors else None,
)
return [
Book(
title=r["title"],
authors=tuple(r["authors"]),
isbn_13=r.get("isbn"),
source=self.name,
source_id=r["id"],
)
for r in results
]
async def fetch_by_isbn(self, isbn: str) -> Book | None:
result = await my_api.get_by_isbn(isbn)
if result:
return Book(...)
return None
# Use it
resolver = BookResolver(sources=[MyLibrarySource()])
Batch Processing
For large datasets, use blocking to avoid O(n²) comparisons:
from book_match import (
BatchMatcher,
TitlePrefix,
FirstAuthorSurname,
BatchConfig,
)
# Custom blocking rules
batch = BatchMatcher(
blocking_rules=[
TitlePrefix(4), # Group by first 4 chars of title
FirstAuthorSurname(), # Group by author surname
],
batch_config=BatchConfig(
min_confidence=0.7,
max_results_per_book=5,
),
)
# Deduplicate with progress
def on_progress(progress):
print(f"{progress.percent_complete:.1f}% - {progress.matches_found} matches")
duplicates = list(batch.deduplicate(books, on_progress=on_progress))
Performance
Benchmarks on a MacBook Pro M2:
| Operation | book-match | Alternative |
|---|---|---|
| Single match | ~50μs | ~400μs (jellyfish) |
| 10k deduplication | ~2s | ~15s (naive) |
| ISBN validation | ~1μs | ~5μs (isbnlib) |
API Reference
Core Types
Book- Immutable book metadataMatchResult- Complete match result with explanationMatchFactor- Individual scoring factorMatchVerdict- AUTO_ACCEPT, REVIEW, REJECT
Matching
BookMatcher- Core matching engineMatchConfig- Matching configuration
Sources
BookResolver- Multi-source orchestratorGoogleBooksSource- Google Books APIOpenLibrarySource- OpenLibrary.org APIMetadataSource- Protocol for custom sources
Batch Processing
BatchMatcher- Batch deduplication and linkingBlockingRule- Base class for blocking strategies
ISBN
is_valid_isbn(),validate_isbn()- Validationisbn10_to_isbn13(),isbn13_to_isbn10()- Conversionnormalize_isbn(),extract_isbns()- Normalization
Contributing
Contributions welcome! Please read our contributing guidelines first.
# Setup development environment
git clone https://github.com/bibliostack/book-match
cd book-match
pip install -e ".[dev]"
# Run tests
pytest
# Run linter
ruff check .
# Type checking
mypy src/
License
MIT License - see LICENSE for details.
Acknowledgments
- RapidFuzz for fast string matching
- OpenLibrary for free book metadata
- Splink for record linkage inspiration
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file book_match-1.0.0.tar.gz.
File metadata
- Download URL: book_match-1.0.0.tar.gz
- Upload date:
- Size: 57.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
454beb618b53195481cb50daf01c146359578c8d00da76d88da1928364309f41
|
|
| MD5 |
6470bf948c33f2b4df4335b9bc77686e
|
|
| BLAKE2b-256 |
df95466f52757ddbc4476e3c68a3419822bfb754cff8f5a0c48851c7d1ea3a59
|
Provenance
The following attestation bundles were made for book_match-1.0.0.tar.gz:
Publisher:
publish.yml on bibliostack/book-match
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
book_match-1.0.0.tar.gz -
Subject digest:
454beb618b53195481cb50daf01c146359578c8d00da76d88da1928364309f41 - Sigstore transparency entry: 1059754089
- Sigstore integration time:
-
Permalink:
bibliostack/book-match@471ad94b04e0f8f62c546033cfe36c738d6f3a2e -
Branch / Tag:
refs/tags/v1.0.0 - Owner: https://github.com/bibliostack
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@471ad94b04e0f8f62c546033cfe36c738d6f3a2e -
Trigger Event:
release
-
Statement type:
File details
Details for the file book_match-1.0.0-py3-none-any.whl.
File metadata
- Download URL: book_match-1.0.0-py3-none-any.whl
- Upload date:
- Size: 51.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
896d176cdf9d730eafda168d0957331f94aa346850e0abe0059e2aafe873ff96
|
|
| MD5 |
fb1e685dcdf63f0bf17fbf184fc0590e
|
|
| BLAKE2b-256 |
55de9c7f9f437a9ae66845b8d97d471b6d8169a5a18ac33cab6b5e21c7140053
|
Provenance
The following attestation bundles were made for book_match-1.0.0-py3-none-any.whl:
Publisher:
publish.yml on bibliostack/book-match
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
book_match-1.0.0-py3-none-any.whl -
Subject digest:
896d176cdf9d730eafda168d0957331f94aa346850e0abe0059e2aafe873ff96 - Sigstore transparency entry: 1059754090
- Sigstore integration time:
-
Permalink:
bibliostack/book-match@471ad94b04e0f8f62c546033cfe36c738d6f3a2e -
Branch / Tag:
refs/tags/v1.0.0 - Owner: https://github.com/bibliostack
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@471ad94b04e0f8f62c546033cfe36c738d6f3a2e -
Trigger Event:
release
-
Statement type: