UniProt reader for LlamaIndex

These details have not been verified by PyPI

Project description

UniProt Reader for LlamaIndex

This package provides a reader for UniProt Swiss-Prot format files, allowing you to load protein data into LlamaIndex for further processing and analysis.

Features

Efficient parsing of large UniProt files with optional lazy loading.
Structured output with both text containing entire UniProt record and metadata containing protein ID.
Configurable field selection

Installation

pip install llama-index-readers-uniprot

Usage

from llama_index.readers.uniprot import UniProtReader

# Initialize the reader
reader = UniProtReader()

# Load data from a UniProt file
documents = reader.load_data("path/to/uniprot_sprot.dat")

# Access the documents
for doc in documents:
    print(f"Protein ID: {doc.metadata['id']}")

Lazy Loading for Large Files

Since UniProt files are large (several GB) it's recommended to use lazy loading to process records one at a time, without loading the entire database into memory:

# Initialize the reader
reader = UniProtReader()

# Load data lazily from a UniProt file
for doc in reader.lazy_load_data("path/to/uniprot_sprot.dat"):
    print(f"Protein ID: {doc.metadata['id']}")
    print("---")

Example of building an index from a lazy loaded UniProt file

from llama_index.readers.uniprot import UniProtReader
from llama_index.core import VectorStoreIndex
from llama_index.core.node_parser import SentenceSplitter

reader = UniProtReader(max_records=10000)

# Load existing protein IDs from the index
existing_protein_ids = {
    node.metadata.get('id')
    for node in index.storage_context.docstore.docs.values()
    if node.metadata.get('id')
}

text_splitter = SentenceSplitter(chunk_size=2048)
index = VectorStoreIndex([], transformations=[text_splitter], show_progress=True)
documents_gen = reader.lazy_load_data("path/to/uniprot_sprot.dat")

# Process documents in batches
batch_size = 10
current_batch = []

for doc in documents_gen:
  protein_id = doc.metadata.get('id')

  if protein_id in existing_protein_ids:
    print(f"Skipping document {protein_id} - already indexed")
    continue


  current_batch.append(doc)

  if len(current_batch) >= batch_size:
      index.refresh_ref_docs(documents=current_batch)
      current_batch = []

# Process any remaining documents
if current_batch:
    index.refresh_ref_docs(documents=current_batch)

# Define persist directory
persist_dir = "path/to/persist/directory"
index.storage_context.persist(persist_dir=persist_dir)

Customizing Field Selection

You can specify which fields to include in the output:

# Only include specific fields
reader = UniProtReader(include_fields={"id", "description", "sequence"})
documents = reader.load_data("path/to/uniprot_sprot.dat")

Available fields:

id: Protein identifier
accession: Accession numbers
description: Protein description
gene_names: Gene names
organism: Organism name
comments: Comments and annotations
keywords: Keywords
sequence_length: Length of the protein sequence
sequence_mw: Molecular weight of the protein
taxonomy: Taxonomic classification
taxonomy_id: Taxonomic database identifiers
citations: Literature citations
cross_references: Cross-references to other databases
features: Protein features

By default, all fields are included.

Limiting Number of Records

You can limit the number of records to parse using the max_records parameter:

# Parse only first 1000 records
reader = UniProtReader(max_records=1000)
documents = reader.load_data("path/to/uniprot_sprot.dat")

# Works with lazy loading too
for doc in reader.lazy_load_data(
    "path/to/uniprot_sprot.dat", max_records=1000
):
    print(f"Protein ID: {doc.metadata['id']}")

Contributing

We welcome contributions! Please see our contributing guidelines for details.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Project details

These details have not been verified by PyPI

Release history Release notifications | RSS feed

0.3.0

Mar 12, 2026

0.2.1

Sep 8, 2025

0.2.0

Jul 30, 2025

This version

0.1.0

Apr 4, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llama_index_readers_uniprot-0.1.0.tar.gz (6.0 kB view details)

Uploaded Apr 4, 2025 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

llama_index_readers_uniprot-0.1.0-py3-none-any.whl (6.8 kB view details)

Uploaded Apr 4, 2025 Python 3

File details

Details for the file llama_index_readers_uniprot-0.1.0.tar.gz.

File metadata

Download URL: llama_index_readers_uniprot-0.1.0.tar.gz
Upload date: Apr 4, 2025
Size: 6.0 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: poetry/1.8.3 CPython/3.12.3 Linux/6.8.0-1021-azure

File hashes

Hashes for llama_index_readers_uniprot-0.1.0.tar.gz
Algorithm	Hash digest
SHA256	`fdd5eed37471a70b67b10773dd19d9b2e149961802ff73344dd55d81fb73c68f`
MD5	`e583c5ed6972c98e8ac6f3283fb95d87`
BLAKE2b-256	`629cb03bf33a8cf6186b0ba40a1177c97ae6416af9bc707a6f572d627468448b`

See more details on using hashes here.

File details

Details for the file llama_index_readers_uniprot-0.1.0-py3-none-any.whl.

File metadata

Download URL: llama_index_readers_uniprot-0.1.0-py3-none-any.whl
Upload date: Apr 4, 2025
Size: 6.8 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: poetry/1.8.3 CPython/3.12.3 Linux/6.8.0-1021-azure

File hashes

Hashes for llama_index_readers_uniprot-0.1.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`4907097df9132d81e850b35a94240a5a9890abcae431e419f41bfd67d3c16de5`
MD5	`e1fd17595eb3636d01fc79bf8a87a5b8`
BLAKE2b-256	`239ed9f921b09bdcecb67392ee8872689b52603526769aa83d92417e349764fb`

See more details on using hashes here.

llama-index-readers-uniprot 0.1.0

Navigation

Verified details

Maintainers

Unverified details

Meta

Classifiers

Project description

UniProt Reader for LlamaIndex

Features

Installation

Usage

Lazy Loading for Large Files

Example of building an index from a lazy loaded UniProt file

Customizing Field Selection

Limiting Number of Records

Contributing

License

Project details

Verified details

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes