Skip to main content

Declarative spatial query layer for Polars

Project description

PyCanopy

PyPI version Python versions CI License: MIT Docs Open In Colab

A spatial query layer for Polars. Rust core, Python API.


[!NOTE] Highly competitive on Apache SpatialBench (single node spatial query benchmark): fastest on 11/24 testcases, within 5% of winning on 14/24 testcases

PyCanopy vs SedonaDB, DuckDB, and GeoPandas on Apache SpatialBench SF1

Apache SpatialBench SF1 · lower is better · bars past the cap truncated with their value · TIMEOUT / ERROR annotated


Installation

pip install pycanopy

Pre-built wheels for Linux, macOS, and Windows. No Rust toolchain required.

import polars as pl
from pycanopy import SpatialFrame

sf = SpatialFrame(pl.read_parquet("cities.parquet"), x_col="lon", y_col="lat")
result = sf.lazy().filter(pl.col("population") > 100_000).range_query(-10.0, 35.0, 40.0, 70.0).collect()

Why PyCanopy

The driving motivator behind creating this library was to provide the optimizations of relational DBs (query planning, indexing, etc) in a fast, Polars-like interface meant for in-memory spatial work.

Edit [June 19 2026]: Apache SedonaDB released a cool Python DataFrame API. There are similarities between their API and this tool but some key differences are that this library uses (1) a Polars-native query engine and (2) a cost model that decides whether and how to index.

PyCanopy GeoPandas DuckDB SedonaDB Spatial Polars
Polars-native API
Spatial query planner (reorder, pushdown, etc)
Index vs scan decided by cost model
Dynamic index selection

Example Operations

Inspecting the query plan

lf = (
    sf.lazy()
    .range_query(min_x=-10.0, min_y=35.0, max_x=40.0, max_y=70.0)
    .filter(pl.col("population") > 100_000)
)
print(lf.explain())
# RANGE_QUERY [(-10, 35) → (40, 70)]
# FROM
#   FILTER [(col("population")) > (dyn int: 100000)]
#   FROM
#     DF [N=100,000; path: EXPR]

The optimizer moved the scalar filter below the range query. It runs first on all rows, then the spatial index is probed on the smaller survivor set.

kNN join

query_df = pl.DataFrame({"qx": [2.35, 13.4], "qy": [48.85, 52.5]})

result = sf.lazy().knn_join(query_df, x_col="qx", y_col="qy", k=3).collect()

For each row in query_df, returns the 3 nearest rows in the SpatialFrame. Large probes are streamed in morsels automatically.

Proximity join with aggregation

import pycanopy as pc

stats = (
    sf.lazy()
    .within_distance_join(landmarks, x_col="lon", y_col="lat", distance=0.5)
    .group_by(["landmark"])
    .agg(count=pc.agg.count(), avg_fare=pc.agg.mean("fare"))
)

The full pair frame is never materialised. Each probe morsel folds into per-group partials and combines at the end.

Polygon intersects self-join

from shapely.geometry import box

polygons = [box(i, 0, i + 1.5, 1.0) for i in range(10_000)]
sf = SpatialFrame.from_polygons(pl.DataFrame({"id": range(10_000), "geom": polygons}), geometry_col="geom")

overlaps = sf.intersects_pairs(key_col="id")

Returns all intersecting polygon pairs with overlap area and IoU. key_col replaces positional indices with values from that column.

[!NOTE] For the full operation catalogue, index modes, streaming joins, and API reference see the docs site.


Benchmarks

Apache SpatialBench

Run on a single m7i.2xlarge (8 vCPU, 32 GB), the same hardware used by Apache SpatialBench. PyCanopy is measured live with index_mode="auto". Results were produced using the benchmark harness in bench/spatial_bench.

PyCanopy wins a total of 11/24 testcases and lands within 5% of winning 14/24 testcases (there is some variance among benchmark runs).

SF1 (~6M trips)

PyCanopy vs SedonaDB, DuckDB, and GeoPandas on Apache SpatialBench SF1

Apache SpatialBench SF1 · lower is better · linear axis, bars past the cap truncated with their value · TIMEOUT / ERROR annotated

SF10 (~60M trips)

PyCanopy vs SedonaDB, DuckDB, and GeoPandas on Apache SpatialBench SF10

Apache SpatialBench SF10 · lower is better · linear axis, bars past the cap truncated with their value · TIMEOUT / ERROR annotated

All times in seconds. Bold = fastest on that query. SedonaDB, DuckDB, and GeoPandas baselines from published SpatialBench results.

SF1

QueryPyCanopySedonaDBDuckDBGeoPandas
q11.390.660.9612.78
q23.748.079.9520.74
q31.230.801.1713.59
q47.448.419.8325.24
q51.715.101.8047.08
q65.518.599.3624.43
q72.151.661.82137.00
q81.041.101.0816.08
q90.230.2350.150.28
q108.6518.79207.8446.13
q119.9032.98TIMEOUT51.01
q1214.8614.55ERRORTIMEOUT

SF10

QueryPyCanopySedonaDBDuckDBGeoPandas
q18.523.044.58ERROR
q29.398.898.26ERROR
q36.884.095.17TIMEOUT
q417.347.528.51ERROR
q514.6050.8114.40ERROR
q611.079.1110.67ERROR
q722.7314.4414.03ERROR
q87.307.247.57TIMEOUT
q90.340.38942.980.49
q1027.2642.02ERRORERROR
q1137.2197.52ERRORERROR
q12175.31145.66ERRORTIMEOUT

How It Works

The engine has dedicated components for query planning / execution and ultimately returns a Polars DataFrame.

Query flow

flowchart LR
    A[User chain] --> B[SpatialOptimizer] --> C[SpatialExecutor] --> F[pl.DataFrame]

Logical planning

  • Predicate pushdown: scalar filters run first, reducing rows before any spatial work.
  • Fusion: consecutive range/contains predicates get interleaved into a single Rust call.
  • Join side: indexes on the side that makes the join most efficient.
  • Projection pushdown: a terminal .select() narrows both join sides before the gather.
  • IO path: low-selectivity queries return results as a direct slice, bypassing the Polars expression pipeline.
  • EXPR path: runs the spatial engine as a Polars map_batches expression over the query set.

Cost model

index_mode determines how we use the cost model:

Mode Behaviour
auto (default) build index when cost model allows it
eager always build the selected index type, skip the cost check
none always scan

When index_mode="auto", the planner picks the minimum-cost option ($Q$ queries, $N$ items):

$$ \text{winner} = \arg\min \begin{cases} \text{Cost}{\text{probe}}(\text{built index}) & \text{build already paid} \ \text{Cost}{\text{build}} + \text{Cost}{\text{probe}}(\text{best new index}) \ \text{Cost}{\text{probe}}(\text{brute force}) \end{cases} $$


Selectivity (fraction of the dataset expected to match):

$$ \text{sel} = \begin{cases} \text{hist}(\text{bbox}) / N & \text{range (32×32 density histogram)} \ k / N & \text{kNN} \ 1 / N & \text{contains} \end{cases} $$


Probe cost ($Q$ warm queries against a built index):

$$ \text{Cost}{\text{probe}} = Q \times \begin{cases} N \cdot c{\text{scan}} & \text{brute force} \ (\log_2 N + \text{sel} \cdot N) \cdot c_{\text{tree}} & \text{KD-tree or R-tree} \ \text{sel} \cdot N \cdot c_{\text{grid}} & \text{grid} \end{cases} $$


Build cost (paid once):

$$ \text{Cost}{\text{build}} = \begin{cases} 0 & \text{brute force} \ N \cdot c{\text{build}} & \text{grid} \ N \log_2 N \cdot c_{\text{build}} & \text{KD-tree or R-tree} \end{cases} $$

The empirical constants ($c_{\text{scan}}$, $c_{\text{tree}}$, $c_{\text{grid}}$, $c_{\text{build}}$) are calibrated from benchmark runs in bench/ops.

Index selection

select_index is a rule-based pre-filter that picks a candidate index type:

flowchart TD
    A[Query arrives] --> B{N < 500\nor sel > 50%?}
    B -- yes --> BF[Brute force]
    B -- no --> C{kNN with\nk/N > 10%?}
    C -- yes --> BF
    C -- no --> D{Polygon\ndataset?}
    D -- yes --> RT[R-tree]
    D -- no --> E{Range query\nand uniform?}
    E -- yes --> GR[Grid]
    E -- no --> KD[KD-tree]

All index types share the same coordinate arrays with no duplication.

Why Rust

The hot paths need packed immutable index structures, zero-copy array slices at the Python boundary, and loop-level parallelism. C++ would require a separate FFI layer and would lose the native Polars plugin integration that PyO3/Maturin provides for free.


Acknowledgements

Some works that inspired this project:

  • Polars: a columnar DataFrame engine that PyCanopy builds on
  • geo-index: provides packed, immutable, zero-copy KD-tree and R-tree structures used
  • Spatial Polars: an earlier effort to bring spatial functionality to Polars
  • Apache Sedona: state-of-the-art spatial SQL engine + benchmark for evals

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pycanopy-0.3.3.tar.gz (385.8 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

pycanopy-0.3.3-cp310-abi3-win_amd64.whl (747.4 kB view details)

Uploaded CPython 3.10+Windows x86-64

pycanopy-0.3.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (880.9 kB view details)

Uploaded CPython 3.10+manylinux: glibc 2.17+ x86-64

pycanopy-0.3.3-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (832.5 kB view details)

Uploaded CPython 3.10+manylinux: glibc 2.17+ ARM64

pycanopy-0.3.3-cp310-abi3-macosx_11_0_arm64.whl (789.8 kB view details)

Uploaded CPython 3.10+macOS 11.0+ ARM64

pycanopy-0.3.3-cp310-abi3-macosx_10_12_x86_64.whl (839.7 kB view details)

Uploaded CPython 3.10+macOS 10.12+ x86-64

File details

Details for the file pycanopy-0.3.3.tar.gz.

File metadata

  • Download URL: pycanopy-0.3.3.tar.gz
  • Upload date:
  • Size: 385.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pycanopy-0.3.3.tar.gz
Algorithm Hash digest
SHA256 7e2d973d61b9f89d1f0460ffee5da63ef99afdd883a155d8c7e4b8aaeb6c66fe
MD5 5a48da5cf045ad4a3eca44947f4b993f
BLAKE2b-256 2223c366e62c7c8e1fe76947077ee5363e585f4454a8f6a6b719338b0e3d1120

See more details on using hashes here.

Provenance

The following attestation bundles were made for pycanopy-0.3.3.tar.gz:

Publisher: release.yml on pranav-walimbe/PyCanopy

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pycanopy-0.3.3-cp310-abi3-win_amd64.whl.

File metadata

  • Download URL: pycanopy-0.3.3-cp310-abi3-win_amd64.whl
  • Upload date:
  • Size: 747.4 kB
  • Tags: CPython 3.10+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pycanopy-0.3.3-cp310-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 bfe6cb07d2095e1924cd9ef2779409a90ee06985b1588aad02b70e2fe39d0ad4
MD5 162954b139ca99151a8962e8b4373a3d
BLAKE2b-256 1601cb735b3ed2193c3e35becfb00819e1d53a9c72c6891ef43b2f8a099cbb12

See more details on using hashes here.

Provenance

The following attestation bundles were made for pycanopy-0.3.3-cp310-abi3-win_amd64.whl:

Publisher: release.yml on pranav-walimbe/PyCanopy

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pycanopy-0.3.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for pycanopy-0.3.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 7c2c454fd211894fff63417c55cbf91ebdec4382ad0da9bdeec79fdd085d5315
MD5 b22880f81188a77f396ce69e10de0715
BLAKE2b-256 3e9519d72c8c9dc321d128ddcef5d46b5a77a289f8482a4ce0883458bac976c5

See more details on using hashes here.

Provenance

The following attestation bundles were made for pycanopy-0.3.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on pranav-walimbe/PyCanopy

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pycanopy-0.3.3-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for pycanopy-0.3.3-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 1fe2b693bf394b5e411d180ed9022b71327486df3b4ee20562623eb0b369bd1e
MD5 def405cf0a52f8deedcc17686290a5c8
BLAKE2b-256 fa6342e1e79fb1220c4cc35fedc6678f558ce84c2fc8e345ed2d052eb39c6a5a

See more details on using hashes here.

Provenance

The following attestation bundles were made for pycanopy-0.3.3-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on pranav-walimbe/PyCanopy

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pycanopy-0.3.3-cp310-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for pycanopy-0.3.3-cp310-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 bef577a76c3b523df77a4649e084850dca8689941b0e74fb0597a4497e191393
MD5 cbceeed472db5d13acc5d8e88a7d8f7d
BLAKE2b-256 43dbda7cacc842777794eb41e27b34c83d81d3d60c4b761de6c7870dc9d2afde

See more details on using hashes here.

Provenance

The following attestation bundles were made for pycanopy-0.3.3-cp310-abi3-macosx_11_0_arm64.whl:

Publisher: release.yml on pranav-walimbe/PyCanopy

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pycanopy-0.3.3-cp310-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for pycanopy-0.3.3-cp310-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 4048d2f063d552efa3256dca4210fec1e30ee237be7062cd142fafc9fd8a33d2
MD5 3ebff514202c25d7260ee1e268e222a1
BLAKE2b-256 e4c395c7b7766828d197dd90daf88c24d54bb39c0bdc9af40c2bcecc1f2358b5

See more details on using hashes here.

Provenance

The following attestation bundles were made for pycanopy-0.3.3-cp310-abi3-macosx_10_12_x86_64.whl:

Publisher: release.yml on pranav-walimbe/PyCanopy

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page