Skip to main content

Stars PyPI Total downloads Monthly downloads

cytome

A single-file format for single-cell multi-omics data.

cytome replaces ad-hoc file stacks with one SQLite-backed .cytome file that keeps matrices, metadata, embeddings, fragments, and provenance together.

cytome stores expression matrices, cell metadata, genomic fragments, embeddings, graphs, and computational provenance in a single SQLite file. It opens from manifest metadata, supports SQL-style metadata filtering, and provides a tested CLI for conversion, merge, subset, export, and validation.

cytome is part of the PIASO toolkit for single-cell analysis. PIASO natively supports cytome datasets for streaming single-cell ATAC-seq and RNA-seq workflows.

Contents

Why cytome?

Capability AnnData (.h5ad) BPCells TileDB-SOMA cytome (.cytome)
Single portable file Yes No (directory layout) No (array store) Yes
Python stdlib core storage dependency No (h5py) No No Yes (sqlite3)
SQL metadata queries No No Limited via APIs Yes
Fragment genomic range queries No native index Limited Yes Yes (SQLite R-tree)
Native provenance table No No Partial workflow metadata Yes
Streaming merge API Partial Strong compute streaming Distributed workflows Yes, chunk-aware
Full matrix load speed Fast Fast Varies Fast
Cloud-native multi-writer scale-out No No Yes No (single-writer SQLite)

Quick start

No optional dependencies — this runs with pip install cytome alone.

import cytome
import numpy as np
import pandas as pd
import scipy.sparse as sp

rng = np.random.default_rng(0)
counts = sp.random(1000, 500, density=0.05, format="csr", dtype=np.float32,
                   random_state=rng)

ds = cytome.create("example.cytome")
ds.set_entity("cells", pd.DataFrame({
    "barcode": [f"cell_{i}" for i in range(1000)],
    "cell_type": rng.choice(["A", "B"], 1000),
}))
ds.set_entity("genes", pd.DataFrame({"gene_id": [f"G{i}" for i in range(500)]}))
ds.add_matrix("RNA_counts", counts)
ds.flush()

print(ds)                                    # metadata-first summary
b_cells = ds.cells.query("cell_type == 'B'")  # SQL-style metadata query
block = ds.RNA.counts[:100, :50]              # random access
for start, end, chunk in ds.RNA.counts.iter_rows():
    ...                                       # bounded-memory streaming
ds.close()

Marker genes, straight off the file

COSG reads a .cytome in chunks, so peak memory does not scale with the number of cells and nothing is converted to AnnData first.

# pip install cosg
import cosg

markers = cosg.cosg("example.cytome", groupby="cell_type",
                    modality="RNA", layer="log1p", n_genes_user=10)
markers["names"]   # per-cell-type marker genes

Coming from AnnData

# pip install "cytome[anndata]"
ds = cytome.from_anndata(adata, modality="RNA", output="rna.cytome")

Key features

  • Instant metadata-first open (CytomeDataset initialization reads manifest and table metadata)
  • SQL-queryable entity tables (cells, genes, peaks, samples)
  • Merge, subset, and downsample APIs with chunk-aware implementations
  • Fragment storage (chunked, compressed) with genomic range queries and export
  • Analysis results live on the file: embeddings (ds.embeddings), neighbor graphs (ds.graphs), per-cell / per-feature columns, and arbitrary analysis artifacts (ds.metadata) — PIASO writes results back onto the cytome
  • Provenance logging with parameters, dependency versions, and methods-text export
  • ACID write transactions via SQLite
  • Multi-modal naming convention (RNA_counts, ATAC_counts, etc.)
  • Optional CSC feature index (build_feature_index) for faster column iteration
  • JSON metadata store (ds.metadata) for arbitrary nested analysis metadata
  • CLI (cytome) for convert, info, merge, subset, downsample, export, validate, provenance, and copy

Installation

pip install cytome

Optional extras:

pip install "cytome[full]"
pip install "cytome[dev]"

Usage examples

Convert from AnnData

import cytome
ds = cytome.from_anndata(adata, modality="RNA", output="rna.cytome")
ds.close()

Open and query

import cytome
ds = cytome.open("rna.cytome")
cd8 = ds.cells.query("cell_type == 'CD8 T'")
x = ds.RNA.counts[:200, :200]
ds.close()

Read without writing

ds = cytome.open("downloaded.cytome", mode="r")   # the file is never changed
ds.read_only                                       # True
ds.close()

A store that cannot be written (a read-only file, or a read-only folder such as a shared data directory) opens for reading on its own, with a warning; a write to it raises cytome.ReadOnlyError saying why.

Merge datasets

import cytome
merged = cytome.merge(["s1.cytome", "s2.cytome"], output="merged.cytome")
merged.close()

Command line

cytome info merged.cytome
cytome merge s1.cytome s2.cytome -o merged.cytome
cytome subset merged.cytome -o cd8.cytome --query "cell_type == 'CD8 T'"

ATAC fragments

ds = cytome.open("multiome.cytome")
fr = ds.ATAC.fragments.query_region("chr1", 1_000_000, 2_000_000)
ds.ATAC.fragments.export("fragments.tsv.gz")
ds.close()

With PIASO

cytome integrates with PIASO (pip install piaso-tools) for single-cell analysis. PIASO supports cytome datasets directly — streaming normalization, dimensionality reduction, clustering, and visualization without loading full matrices into memory.

PIASO functions are self-contained on a cytome: they stream from the file and write their results back onto it (graphs, embeddings, cluster labels, per-cell/per-feature metrics, markers) — they don't return large objects, so you read results from the dataset afterwards.

import cytome
import piaso

# Open a cytome dataset and analyse it with PIASO (results persist on ds)
ds = cytome.open("atac.cytome")

piaso.pp.calculateCellMetrics(ds)         # per-cell QC -> ds.cells
piaso.tl.runSVDLazy(ds)                   # X_svd embedding -> ds.embeddings
piaso.tl.neighbors(ds)                    # connectivities/distances -> ds.graphs
piaso.tl.leiden(ds, key_added="leiden")   # labels -> ds.cells["leiden"]
piaso.tl.umap(ds)                         # X_umap -> ds.embeddings
piaso.pl.plotUMAP(ds, color="leiden")

# Marker genes can persist to the cytome too
import cosg
cosg.run_cosg_cytome(ds, groupby="leiden", modality="ATAC",
                     write_to_cytome=True)        # -> ds.metadata["cosg"]

ds.close()

See piaso.org for documentation and tutorials.

File format

A .cytome file is a SQLite database with:

  • Entity tables (cells, genes, peaks, ...)
  • Chunked compressed sparse matrices (matrix_chunks + matrix_meta)
  • Optional CSC chunks (matrix_csc_chunks)
  • Dense chunked embeddings (dense_chunks + embedding_meta)
  • Fragment storage (chunked compressed BLOBs) + genomic R-tree indices
  • Provenance (_provenance) and metadata (_metadata)

Comparison with existing tools

Trade-offs in the current implementation:

  • Full matrix materialization can be slower than direct HDF5 reads for some workloads.
  • Merge/subset APIs are chunk-aware, but current merge implementation may still materialize intermediate matrices for gene remapping.
  • SQLite provides strong local reliability and portability but is not a distributed/cloud-native storage engine.

Citation

If you use cytome, please cite this repository (a citation entry will be added here when available).

License

BSD 3-Clause License. See LICENSE for details.

Metadata

Release files for cytome 0.3.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cytome 0.3.6
File Size Uploaded
cytome-0.3.6.tar.gz 212.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cytome 0.3.6
File Interpreter ABI Platform
cytome-0.3.6-py3-none-any.whl Python 3 none any Details

Total release size: 392.9 kB

Release files / cytome-0.3.6.tar.gz

Download URL cytome-0.3.6.tar.gz
Size 212.8 kB
Tags Source
SHA-256 checksum
How to use checksums
f99f18666bb540bd5782d76b35e26cc6b55cb18b4bdbf4f24cc54f043b9cb345
BLAKE2b-256 checksum
How to use checksums
5290d55fe3ab8744cd48dcc55db127ffdfe9c65f69619d28b7b5b2ff5eb8dfdf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / cytome-0.3.6-py3-none-any.whl

Download URL cytome-0.3.6-py3-none-any.whl
Size 180.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b022c7a928bcd6d4d585d7db7e373dacf2bdd9f71bd9b20ce3761a620d270577
BLAKE2b-256 checksum
How to use checksums
2dd55fe1a0b98f40fd04e00dfc41545afebaec5a9fe3365709886df8dd4d040d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.6 This release

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page