Skip to main content

DisSCube

CI Docs License: MIT

Status: Alpha — stable APIs for the core pipeline; declarative models still evolving.

DisSCube is the spatial data cube engine of the DisSModel ecosystem. It converts raw geospatial sources (rasters, vectors) into derived variables aligned to LUCC (Land Use and Cover Change) modeling grids, ready for Cellular Automata models and spatio-temporal analysis.

📖 Documentation: https://dissmodel.github.io/disscube/

Core concept

SpatialSource  →  Derivation  →  Variable  →  DerivedVariable (Zarr)

A source (SpatialSource) goes through a derivation (SpatialDerivation or Derivation) that applies an operator on a grid (GridSpec), producing a derived variable registered in the SQLite catalog and stored in Zarr.

Installation

git clone https://github.com/DisSModel/disscube.git
cd disscube
python -m venv .venv && source .venv/bin/activate
pip install -e .

Basic usage

1. Initialize the catalog and register a grid

from disscube.client import CubeClient
from disscube.utils.grids import register_local_grid

cube = CubeClient(catalog="catalog.db", store="./data/")

grid = register_local_grid(
    cube,
    name="AC",
    bbox_geo=(-73.99, -11.15, -66.62, -7.11),
    resolution=5_000.0,
)

2. Register a source

from disscube.models import SpatialSource

cube.register_spatial_source(SpatialSource(
    id="mapbiomas_2020",
    name="MapBiomas Acre 2020",
    format="raster",
    asset_url="data/raw/mapbiomas_2020.tif",
    crs="EPSG:4326",
    time=2020,
))
from disscube.derivation import Derivation

d = Derivation(
    target="forest_pct",
    source_id="mapbiomas_2020",
    operator="percentage",
    class_code=3,
    role="driver",
    valid_from="2020",
    valid_until="2020",
)

cube.derive_declarative(d, grid_id="AC/5km")

4. Derive — direct mode

from disscube.models import SpatialDerivation, Variable

cube.derive(SpatialDerivation(
    source_id="mapbiomas_2020",
    grid_id="AC/5km",
    role="driver",
    variables=[Variable(name="forest_pct", operator="percentage", class_code=3)],
    valid_from="2020",
    valid_until="2020",
))

5. Load the result

da = cube.load("forest_pct", grid_id="AC/5km")
print(da.shape)   # (rows, cols)

6. Get the cube out

ds = cube.to_dataset(["forest_pct", "dist_roads"], grid_id="AC/5km", period=("2015", "2020"))
# xarray.Dataset: (y, x) static and (time, y, x) temporal variables, CRS and transform via ds.rio

cube.export_geotiff(["forest_pct"], "forest.tif", grid_id="AC/5km")   # one band per variable and year
cube.export_netcdf(["forest_pct"], "cube.nc", grid_id="AC/5km")       # CF-1.8; pip install "disscube[netcdf]"

Exports carry their provenance: each GeoTIFF band and netCDF variable records the spec_hash, the content_hash of the stored data and the source_checksum of the input it came from (cube.provenance("forest_pct") lists them per year).

DisSCube does not need DisSModel. To hand a cube to a DisSModel model, install disscube[dissmodel] and use cube.to_raster_backend(...), which returns a RasterBackend.

Pipeline files (TOML)

A whole data preparation — grid, sources, derived variables — can be declared in one TOML file, the counterpart of a TerraME fill script, and run from the command line:

disscube validate examples/pipelines/quickstart.toml   # no downloads
disscube run      examples/pipelines/quickstart.toml --workspace outputs/quickstart
schema = 1
name = "Quickstart"

[grid]
name = "demo/300m"
crs = "EPSG:31983"
bbox = [570000.0, 9708000.0, 582000.0, 9720000.0]
resolution = 300

[[source]]
id = "landuse"
type = "file"
path = "../data/quickstart/landuse.tif"
crs = "EPSG:31983"

[[derive]]
target = "forest_pct"
source = "landuse"
operator = "percentage"
class_code = 3

See docs/guides/pipeline_files.md and examples/pipelines/.

Examples

examples/ has runnable, self-contained examples that run offline in seconds:

  • 01–03, synthetic data — a quickstart with raster operators, vector drivers, and time series handed off to DisSModel.
  • quickstart.toml — declarative pipeline equivalent to example 01.
python examples/01_quickstart.py
disscube run examples/pipelines/quickstart.toml

Real data workflows, TerraME parity benchmarks, and large-scale case studies are maintained in LambdaGeo/disscube-recipes (cases/terrame_fill, cases/ilha_maranhao, cases/prodes_br163, cases/luccme_br).

See examples/README.md for the full list, and docs/terrame_fill_correspondence.md for how DisSCube's operators relate to TerraME's Fill.

Available operators

Operator Type Resampling Requires class_code
mean zonal average no
sum zonal sum no
std zonal nearest no
min zonal min no
max zonal max no
majority zonal nearest¹ no
minority zonal nearest¹ no
percentage zonal nearest¹ yes
attribute zonal nearest no
presence zonal nearest no
distance proximity (exact, cell centre → nearest feature, source not clipped) — no
min_distance proximity (raster approximation, features inside the grid) nearest no
count proximity nearest no
area polygons (exact share of the cell covered) — no

distance takes params = {crs = …} to measure in another CRS (metres on a geographic grid); the ¹ operators take params = {subcells = n} to cap the fine pixels per cell. Any derivation takes fill = "nearest".

¹ These use needs_fine_alignment=True: GridAligner resamples with nearest at high resolution; the actual reduction (per-window counting) is done by the operator.

Pipeline

SpatialSource
    │
    ▼
Normalizer        — validates / loads a GeoDataFrame (vector) or opens the raster
    │
    ▼
GridAligner       — reprojects per variable with the operator's Resampling
    │
    ▼
Aggregator        — delegates to operator.compute() → one xr.DataArray per variable
    │
    ▼
VariableWriter    — writes Zarr + registers the DerivedVariable in the catalog

Storage layout

data/derived/{grid_id}/{partition}/{spec_hash}/{variable_name}.zarr
  • partition = tile_id, or global for untiled derivations.
  • spec_hash = SHA-256 of the derivation (source + grid + variables + time window, plus the source's checksum when it has one — replacing a source file and registering its new checksum recomputes instead of returning a stale product).

Project structure

disscube/
├── client.py         CubeClient — public entry point
├── models/           GridSpec, SpatialSource, SpatialDerivation, Variable, Derivation…
├── operators/        Operators as classes (self-registered via __init_subclass__)
│   ├── base.py       Operator ABC + OPERATOR_REGISTRY
│   ├── zonal.py      mean, sum, majority, percentage, attribute, presence…
│   └── proximity.py  distance, min_distance, count
├── pipeline/         Pipeline execution & planning (schema, runner) + internal stages
├── catalog/          CatalogStore (Protocol) + SQLite and JSON implementations
├── storage.py        AssetStore (fsspec — local and S3)
├── cli.py            `disscube validate` / `disscube run` / `disscube export`
├── sources/          Adapters that bring external data in as SpatialSources,
│   │                 each with a checksum and a provenance.json sidecar
│   ├── _raster.py    Window2D, windowed reads, composites, mosaics, register_raster
│   ├── bdc.py        Brazil Data Cube cubes via STAC (`bdc` extra)
│   ├── _categorical.py  legends (.qml/.json/.csv) and reclassification
│   ├── mapbiomas.py  MapBiomas annual land-cover maps (Collection 11, 10 m series)
│   ├── prodes.py     PRODES deforestation (download + cache, legend from the .qml)
│   └── classified.py any classified map with its legend, e.g. from SITS
└── utils.py          Checksums (sha256_file) and BDC tile importer (import_bdc_grids)

Adding a new operator

Create a subclass of Operator in any module imported at startup:

from rasterio.warp import Resampling
from disscube.operators.base import Operator

class WeightedMeanOperator(Operator):
    name = "weighted_mean"
    _resampling = Resampling.average

    def compute(self, data, var, grid):
        # data is an xr.DataArray (raster) or a GeoDataFrame (vector)
        ...

The operator is registered automatically and accepted by Derivation / SpatialDerivation with no other change.

Known limitations

The limitations below are scope decisions for the current version, not bugs. They are documented so that users and reviewers understand what is implemented versus what is planned.

In-memory, single-tile processing Each call to derive() loads a tile's full data into memory. There is no lazy (Dask) or distributed processing. For continental-scale grids (e.g. BR/1km), use the tile loop — each tile is processed and saved independently.

Vector aggregation by rasterization (not area-weighted) Operators over vector sources (majority, percentage, attribute, presence, minority) convert geometries to raster before aggregating pixels. For the share of each cell covered by polygons, use area, which intersects the polygons with the cells exactly.

Tile disambiguation in load() CubeClient.load(name) without tile_id raises ValueError when multiple tiles of the same variable exist on the same grid. Automatic mosaicking is not implemented. Always pass tile_id in multi-tile workloads.

SpatialRelation does not act in the pipeline The SpatialRelation model is persisted in the catalog, but no pipeline stage uses relations during derivation — which is why they are excluded from spec_hash. Including them would make the cache key sensitive to metadata that does not affect the result, breaking the reproducibility guarantee. Integration with hierarchical grid strategies is reserved for a future version.

purity_threshold reserved The purity_threshold field on Derivation is included in spec_hash but is not applied to the output — purity masking is not implemented. Setting purity_threshold changes the cache key without changing the result.

STAC: reading only disscube.sources.bdc reads Brazil Data Cube cubes through their STAC catalog (search, windowed reads, per-tile composites, mosaics) and writes local GeoTIFFs that are registered as ordinary sources. Derived variables are not published back as STAC, and the valid_from/valid_until and bbox fields on Derivation only follow STAC naming conventions.

Citation

If you use DisSCube in your research, dynamic modeling, or spatial data pipelines, please cite it using the metadata from CITATION.cff or the following BibTeX entry:

@software{costa_disscube_2026,
  author       = {Costa, S{\'e}rgio Souza},
  title        = {{DisSCube: Declarative spatial data cubes}},
  year         = {2026},
  version      = {0.4.0},
  url          = {https://github.com/DisSModel/disscube}
}

License

DisSCube is part of the DisSModel ecosystem and is released under the MIT License. See LICENSE for details.

Metadata

Release files for disscube 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for disscube 0.4.0
File Size Uploaded
disscube-0.4.0.tar.gz 140.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for disscube 0.4.0
File Interpreter ABI Platform
disscube-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 247.0 kB

Release files / disscube-0.4.0.tar.gz

Download URL disscube-0.4.0.tar.gz
Size 140.5 kB
Tags Source
SHA-256 checksum
How to use checksums
316694a9849a91e2c42c819c418360bba0adc91225121f3789c00ea1ad5989cb
BLAKE2b-256 checksum
How to use checksums
ebf6b32975603ddd815b17ab58bba4cbba9150ceda4ababa8127a3c20c416f41
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / disscube-0.4.0-py3-none-any.whl

Download URL disscube-0.4.0-py3-none-any.whl
Size 106.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
552793526afa43d029fa756b8da1aa3883558ff0e9063be2865275350fd75c87
BLAKE2b-256 checksum
How to use checksums
c07e04991945a09ebbbf9892c176c10357446336d6a03edaf2770edc5bb57691
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page