Skip to main content

lgopy-catalog

Reusable block catalog for LgoPy pipelines.

Installation

pip install lgopy-catalog

For local workspace development, run from the repository root:

uv sync

Usage

lgopy-catalog works with block packages generated by Block.build():

blocks/
  normalize/
    1.0.0/
      block.py
      __init__.py
      requirements.txt
      manifest.json
      schema.json

Publish an existing package directory into a catalog:

from lgopy_catalog import BlockCatalog, FSSpecBlockStore

catalog = BlockCatalog(block_store=FSSpecBlockStore("gs://my-lgopy-catalog"))
catalog.publish_package(".lgopy/blocks/normalize/1.0.0")

Search and load packages:

matches = catalog.search("normalization")
Normalize = catalog.load("normalize", "1.0.0")
block = catalog.create("normalize", version="1.0.0", scale=10.0)

Build a pipeline from published blocks

Install lgopy in the execution environment and install the selected packages' requirements.txt files; the catalog does not install them automatically. For a published normalize version 1.0.0 whose constructor accepts scale:

from lgopy.core import LgoPipeline

steps = [
    {"block": "normalize", "version": "1.0.0", "args": {"scale": 10.0}},
]
report = LgoPipeline.validate(steps, catalog=catalog)
if not report["valid"]:
    raise ValueError(report["issues"])
pipeline = LgoPipeline.from_list(steps, catalog=catalog)
# result = pipeline(your_inputs)
pipeline.save("pipeline.json")
restored = LgoPipeline.from_file("pipeline.json", catalog=catalog)

No import of the publisher's block module is needed. See the complete publishing and pipeline tutorial for runnable examples and dependency preparation.

Semantic search

Install the optional semantic-search dependencies:

pip install "lgopy-catalog[rag]"

Configure PostgreSQL and the embedding model with environment variables:

export LGOPY_CATALOG_DB_HOST=localhost
export LGOPY_CATALOG_DB_PORT=5432
export LGOPY_CATALOG_DB_NAME=lgopy_catalog
export LGOPY_CATALOG_DB_USER=postgres
export LGOPY_CATALOG_DB_PASSWORD=postgres
export LGOPY_CATALOG_EMBEDDING_DIM=768

Pass an embedding adapter to enable semantic indexing and search:

from lgopy_catalog import BlockCatalog, FSSpecBlockStore, GeminiEmbedding

catalog = BlockCatalog(
    block_store=FSSpecBlockStore("gs://my-lgopy-catalog"),
    embeddings=GeminiEmbedding(model_id="gemini-embedding-001"),
)

catalog.publish_package(".lgopy/blocks/normalize/1.0.0")
matches = catalog.semantic_search("normalize multispectral imagery before NDVI", k=5)

Published block embeddings are generated from manifest metadata, schema details, and the block class call method signature, docstring, and source extracted from block.py.

The catalog package uses a compact layout:

lgopy_catalog.catalog      # BlockCatalog API
lgopy_catalog.store        # block package stores
lgopy_catalog.models       # small dataclasses and protocols
lgopy_catalog.schemas      # database table schemas
lgopy_catalog.utils        # manifest, call-method, runtime helpers
lgopy_catalog.rag          # embeddings and vector search

Built-in embedding adapters:

GeminiEmbedding
OllamaEmbedding

Use Ollama instead of Gemini:

from lgopy_catalog import BlockCatalog, FSSpecBlockStore, OllamaEmbedding

catalog = BlockCatalog(
    block_store=FSSpecBlockStore("gs://my-lgopy-catalog"),
    embeddings=OllamaEmbedding(model_id="embeddinggemma"),
)

matches = catalog.semantic_search(
    "block for vegetation index calculation",
)

Gemini configuration:

export GOOGLE_API_KEY=...
export LGOPY_CATALOG_GEMINI_EMBEDDING_MODEL_ID=gemini-embedding-001
export LGOPY_CATALOG_GEMINI_OUTPUT_DIM=768

Ollama configuration:

export LGOPY_CATALOG_OLLAMA_BASE_URL=http://localhost:11434
export LGOPY_CATALOG_OLLAMA_EMBEDDING_MODEL_ID=embeddinggemma

You can also provide your own embedding adapter by implementing the EmbeddingModel protocol:

import numpy as np

class CustomEmbedding:
    model_name = "custom"

    async def embed(self, text: str) -> np.ndarray:
        return np.array(my_embedding_function(text), dtype=np.float32)

catalog = BlockCatalog(
    block_store=FSSpecBlockStore("gs://my-lgopy-catalog"),
    embeddings=CustomEmbedding(),
)

Remove a package version:

catalog.remove_package("normalize", "1.0.0")

catalog.search() performs text matching with metadata filters; semantic search ranks indexed candidates by meaning. Each semantic result includes a name, version, manifest, schema, model name, and cosine distance (lower is closer). Use distance_threshold for an optional maximum distance, not a confidence score. Inspect a candidate's schema and validate its arguments before execution.

The built-in index needs a reachable PostgreSQL server with pgvector and suitable initialization permissions. Existing packages are not automatically indexed when embeddings are enabled: republish them through the configured catalog. Query and index embeddings must use the same model and vector dimension. Changing models requires indexing packages for that model; give distinct configurations distinct model_name values. See the semantic-search guide for complete setup, indexing, and result-to-pipeline examples.

Metadata

Release files for lgopy-catalog 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lgopy-catalog 2.0.0
File Size Uploaded
lgopy_catalog-2.0.0.tar.gz 21.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lgopy-catalog 2.0.0
File Interpreter ABI Platform
lgopy_catalog-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 47.9 kB

Release files / lgopy_catalog-2.0.0.tar.gz

Download URL lgopy_catalog-2.0.0.tar.gz
Size 21.3 kB
Tags Source
SHA-256 checksum
How to use checksums
384d40ef7b3c88104610001462871815e03532543bba7f02f5abdb710092d4dc
BLAKE2b-256 checksum
How to use checksums
c943ee44584d203a9b2466c3f2a5a23212e1918a21133b93fffdb4ba84c4b84c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release files / lgopy_catalog-2.0.0-py3-none-any.whl

Download URL lgopy_catalog-2.0.0-py3-none-any.whl
Size 26.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d0c9ebcbfec7e68e2c541a39574713648fbf416e3c88c7ccae454f7f0f30ef83
BLAKE2b-256 checksum
How to use checksums
5cbfd8a2e0d6b8cad106523daaec62b183c738157abe096eb8d09f2a22b19cde
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page