docculus
Overview
docculus is a lightweight Python library that makes it easy to analyze, transform, hash, and
validate
langchain_core.documents.Document
objects.
Quick Links:
Why docculus?
Working with collections of LangChain documents often means writing the same boilerplate over and
over: counting duplicates, filtering by metadata, formatting documents into an LLM-friendly
prompt, or hashing content to detect changes. docculus packages these operations into a small,
well-tested API:
Deduplicate and format documents for a prompt:
>>> from langchain_core.documents import Document
>>> from docculus.transform import deduplicate_documents, format_documents
>>> docs = [
... Document(id="1", page_content="The cat sat on the mat."),
... Document(id="2", page_content="The cat sat on the mat."),
... Document(id="3", page_content="The dog chased the ball."),
... ]
>>> unique_docs = deduplicate_documents(docs)
>>> print(format_documents(unique_docs, output_format="markdown"))
## Document 1
<BLANKLINE>
The cat sat on the mat.
<BLANKLINE>
## Document 2
<BLANKLINE>
The cat sat on the mat.
<BLANKLINE>
## Document 3
<BLANKLINE>
The dog chased the ball.
Inspect a corpus at a glance:
>>> from docculus.analysis import compute_content_stats_exact
>>> stats = compute_content_stats_exact(docs)
>>> stats["count"], stats["duplicate_count"]
(3, 1)
Check consistency across documents sharing an id:
>>> from docculus.validation import validate_document_consistency
>>> validate_document_consistency(docs)
True
Features
docculus provides a comprehensive set of utilities for working with Document objects:
📊 Analysis
Corpus-wide, read-only inspection utilities:
- Content statistics (exact and approximate) with
compute_content_stats_exact()andcompute_content_stats_approx() - Metadata statistics with
compute_metadata_stats() - Duplicate and empty document detection with
find_duplicate_document_ids(),find_empty_documents(), andfind_empty_document_ids() - Human-readable report printing with
print_content_stats_report()andprint_metadata_stats_report()
🔄 Transform
Utilities that produce a new document list:
- Deduplicate documents with
deduplicate_documents() - Filter by metadata with
filter_by_metadata(),filter_by_metadata_range(), andfilter_by_metadata_values() - Sort with
sort_by_metadata()and truncate withtruncate_documents() - Assign or copy ids with
assign_ids()andcopy_ids_to_metadata() - Format documents into LLM-friendly strings (XML, Markdown, JSON) with
format_documents()
📄 Document
Per-document utilities:
- Id generation with
generate_id(),generate_random_id(), andgenerate_deterministic_id() - Emptiness checks with
is_empty()andis_whitespace_only() - Length helpers with
get_length(),get_lengths(),get_longest_document(), andget_shortest_document()
#️⃣ Hashing
Deterministic hashing of documents:
- Hash a document or a sequence of documents with
hash_document()andhash_documents() - Generate a stable UUID from a document with
hash_document_to_uuid() - Register custom hashing strategies with
register_document_hasher()
✅ Validation
Consistency checks across documents sharing an id, via validate_document_consistency().
🖨️ Display
Pretty-print documents and their metadata to the terminal with print_document(),
print_documents(), and print_documents_metadata().
🗄️ Store
Persist and retrieve documents with a common BaseDocumentStore interface, with
InMemoryDocumentStore, SQLiteDocumentStore, and DuckDBDocumentStore implementations
(plus typed variants).
Installation
We highly recommend installing
docculus in
a virtual environment
to avoid dependency conflicts.
Using uv (recommended)
uv is a fast Python package installer and resolver:
uv pip install docculus
Install with all optional dependencies:
uv pip install docculus[all]
Install with specific optional dependencies:
uv pip install docculus[duckdb,rich] # with DuckDB and rich
Using pip
Alternatively, you can use pip:
pip install docculus
Install with all optional dependencies:
pip install docculus[all]
Requirements
- Python: 3.10 or higher
- Core dependencies:
coola,langchain-core - Optional dependencies (install with
docculus[all]): DuckDB • Faker • persista • rich
Compatibility Matrix
coola |
coola |
langchain-core |
duckdb* |
faker* |
persista* |
rich* |
python |
|---|---|---|---|---|---|---|---|
main |
>=1.1.10,<1.0 |
>=1.4,<2.0 |
>=1.3,<2.0 |
>=40.0,<41.0 |
>=0.0.3,<1.0 |
>=14.0.0,<16.0 |
>=3.10 |
0.0.1 |
>=1.1.10,<1.0 |
>=1.4,<2.0 |
>=1.3,<2.0 |
>=40.0,<41.0 |
>=0.0.3,<1.0 |
>=14.0.0,<16.0 |
>=3.10 |
* indicates an optional dependency
Contributing
Contributions are welcome! We appreciate bug fixes, feature additions, documentation improvements, and more. Please check the contributing guidelines for details on:
- Setting up the development environment
- Code style and testing requirements
- Submitting pull requests
Whether you're fixing a bug or proposing a new feature, please open an issue first to discuss your changes.
API Stability
:warning: Important: As docculus is under active development, its API is not yet stable and
may
change between releases. We recommend pinning a specific version in your project’s dependencies to
ensure consistent behavior.
License
docculus is licensed under BSD 3-Clause "New" or "Revised" license available in LICENSE
file.
Metadata
Release files for docculus 0.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| docculus-0.0.1.tar.gz | 45.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| docculus-0.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 114.5 kB
Release files / docculus-0.0.1.tar.gz
| Download URL | docculus-0.0.1.tar.gz |
|---|---|
| Size | 45.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
408668d1af9c2dfb057d002b8d1bf71d8ec0c29d120d05b7f77a2a6a5da79057
|
|
BLAKE2b-256 checksum How to use checksums |
3c4bc8a5b6223091d7a016662e5ab8ee865d1f2cc5908615f42ceb7689202f2a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.1 {"installer":{"name":"uv","version":"0.12.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / docculus-0.0.1-py3-none-any.whl
| Download URL | docculus-0.0.1-py3-none-any.whl |
|---|---|
| Size | 68.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2b9bbfa6d1c43637bd8c2d891138d600c394f0c7fa226d6722132dbc12a8968f
|
|
BLAKE2b-256 checksum How to use checksums |
ac9547e1905aa412f68b411785f8d71e6efe035a4f8eb272f8604941e1d17103
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.1 {"installer":{"name":"uv","version":"0.12.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|