Skip to main content

Z7

IGEO7 logo

This is a work-in-progress to explore the possibility of neighbor traversal in the IGEO7/Z7. Z7 is an indexing system for the IGEO7 aperture 7 hexagonal discrete global grid DGGRID/Sahr Kmoch et al. (2025). It is based on the Generalized Balanced Ternary (GBT) numeral system described in Lucas, Gibson (1982), van Roessel (1988), Sahr (2019), and Wikipedia (2025).

Installation

Published on PyPI and on prefix.dev.

With pip:

pip install z7py

With pixi:

pixi project channel add https://prefix.dev/allixender/geo
pixi add z7py

With conda or mamba:

conda install -c https://prefix.dev/allixender/geo z7py

Not yet on conda-forge

conda install -c conda-forge z7py is the intended end state, but the package is not published there yet. Until then use PyPI or the channel above.

z7py is deliberately thin: it depends only on NumPy and Numba. It needs neither the DGGRID binary nor a geospatial stack, so it can be imported inside Numba nopython kernels and on bare compute workers.

import numpy as np
from z7py import z7

raw = z7.encode_z7int(3, [1] * 10)        # base cell 3, 10 resolution digits
print(z7.get_resolution(raw))             # 10
print(z7.index_to_z7string(raw))          # '031111111111'

# neighbour traversal is pure integer arithmetic
print(z7.get_neighbours(raw))            # 6 neighbours, UINT64_MAX where invalid

# monotonic mapping preserves sort order -> range compression
mono = z7.z7_to_monotonic_int(raw, 10)
assert z7.monotonic_int_to_z7(mono, 10) == raw

Acknowledgements

2019-2026 Kevin Sahr sahrk@sou.edu DGGRID https://github.com/sahrk/DGGRID/

https://github.com/wrenoud/Z7/

This is a work-in-progress to explore the possibility of neighbor traversal in the Z7 domain. Z7 is an indexing system for aperature 7 hexagonal discrete global grid systems, as coined in Kmoch et al. (2025). It is based on the Generalized Balanced Ternary (GBT) numeral system described in Lucas, Gibson (1982), van Roessel (1988), and Wikipedia (2025).

Rotation Pattern

Lucas and Gibson (1982) gave examples of GBT with exclusively counter-clockwise (CCW) rotation. The disadvantage of this approach is that the orientation of the hexagons at the different hierarchy levels quickly diverge from each other. White et al. (1992) proposed an alternating rotation direction (CW, CCW, CW, ...) to maintain an alignment of hexagon orientations across hierarchy levels. Kmoch et al. (2025) adopted this alternating rotation pattern for Z7 assigning CW to odd resolutions and CCW to even resolutions.

GBT Addition

Make nicer figures?

https://github.com/wrenoud/Z7/?tab=readme-ov-file


High-Performance GBT Implementation

The z7py implementation uses a bit-packed 64-bit representation to enable extremely fast neighbor traversal and range arithmetic.

Bit Layout

Indices are stored as uint64 with the following structure:

  • Bits 63–60: Base Cell ID (0–11).
  • Bits 59–0: 20 refinement digits, each occupying 3 bits.
  • Value 7: A special padding digit indicating the end of the identifier.

Neighbor Traversal and Cascading Carries

Neighbor finding is implemented using GBT arithmetic tables (_GBT_CW_* and _GBT_CCW_*) compiled with Numba.

  1. Alternating Rotations: The implementation automatically switches between Clockwise (CW) and Counter-Clockwise (CCW) addition tables based on whether the resolution level is odd or even.
  2. Cascading Carries: Moving across a cell boundary often triggers a "carry" to the parent level. The algorithm recursively ripples these carries up the hierarchy (similar to carrying a 1 in decimal addition) until the carry is 0 or a base-cell boundary is reached.
  3. Bitwise Efficiency: All digit extractions and updates use raw bit-shifts (>>) and masks (&), ensuring O(1) performance for single-level operations and avoiding expensive string or list allocations.

Sorted Range Index for Z7 (IGEO7) DGGS

Overview

The Z7 indexing scheme (IGEO7) represents a hierarchical discrete global grid. When data is spatially sorted using Z7 identifiers, it naturally clusters along space-filling curves. The Sorted Range Index is a conceptual optimization for Xarray and Zarr that exploits this clustering to provide O(1) or O(log N) data access without loading massive coordinate arrays.

Note: The primary goal of this index is not necessarily to perform exact spatial filtering (such as polygon intersections, or parent-aligned spatial children), but rather to enable highly efficient access to specific Z7 identifiers or larger blocks by mapping them directly to their physical storage locations, reducing chunk loading efforts for close cells.

Key Conceptual Pillars

1. The Base-7 Transition

While Z7 identifiers are natively stored as packed uint64 (Base-2 bit patterns), this representation often contains numerical "gaps" between siblings.

  • Base-7 Encoding: By converting Z7 strings or bit-patterns into Base-7 integers, siblings (children of the same parent) become numerically adjacent (incrementing by 1).
  • Contiguous Ranges: In Base-7, a parent and all its recursive children form a single, contiguous numerical range. This allows representing millions of points with just two values: [start, end].

2. Range Compression & Multi-Range Indexing

A dataset can be represented as a collection of these contiguous ranges. Instead of an index of size $N$ (where $N$ is the number of cells), we maintain an index of size $R$ (where $R$ is the number of contiguous ranges). In well-sorted DGGS datasets, $R \ll N$.

3. Zarr & Dask Alignment

The index is designed to be "chunk-aware":

  • Chunk-Bound Ranges: Ranges are ideally calculated per Zarr chunk using Dask.
  • Efficient Filtering: When querying a spatial region, the index identifies which ranges overlap the query, mapping them directly to specific Zarr chunks. This bypasses the need to read the zone_id coordinate variable from disk for non-relevant chunks.

Conceptual Implementation idea

The general concept of contiguous ranges for DGGS indices:

import pandas as pd

# Convert base-7 strings to integers for easier comparison
def base7_to_int(s):
    return int(s, 7)


# Convert integers back to base-7 strings with proper padding
def int_to_base7(n, length):
    if n == 0:
        return '0' * length
    digits = []
    while n:
        digits.append(str(n % 7))
        n //= 7
    return ''.join(reversed(digits)).zfill(length)


def find_continuous_ranges(z7str_cell_ids):
    """
    Find all continuous ranges in a pandas Series of base-7 digit strings.
    
    Parameters:
    series: pd.Series of strings containing digits 0-6
    
    Returns:
    List of tuples (start, end) representing continuous ranges
    """
    if len(z7str_cell_ids) == 0:
        return []
    
    # Remove duplicates and sort
    unique_sorted = sorted(z7str_cell_ids.unique())
    
    
    # Convert to integers
    int_values = [base7_to_int(x) for x in unique_sorted]
    length = len(unique_sorted[0])
    
    # Find continuous ranges
    ranges = []
    start_idx = 0
    
    for i in range(1, len(int_values)):
        # Check if there's a gap
        if int_values[i] != int_values[i-1] + 1:
            # End of a range
            ranges.append((unique_sorted[start_idx], unique_sorted[i-1]))
            start_idx = i
    
    # Add the last range
    ranges.append((unique_sorted[start_idx], unique_sorted[-1]))
    
    return ranges

# Example usage
data = pd.Series(['00001','00002','00003', '00004', '00012', '00013', '00014', '00030', '00031', '00012'])

ranges = find_continuous_ranges(data)

print(ranges)
# Expected Output: [('00001', '00004'), ('00012', '00014'), ('00030', '00031')]

print(len(ranges))
# Expected Output: 3

Performance and Correctness Considerations

  • Sorting Requirement: The efficiency and correctness of this index rely on the data being lexicographically sorted by Z7 identifier during the dataset creation (regridding) phase. If the data is shuffled, the ranges degrade into individual points.
  • Coordinate Transformation: The index provides a bidirectional mapping:
    • Forward: Maps an absolute position in the data array to a Z7 identifier.
    • Reverse: Maps a Z7 identifier (or range) back to absolute positions/slices in the data array.
  • Lazy Evaluation: By utilizing Dask for range calculation and storage, the index remains lightweight and integrates seamlessly with Xarray's lazy loading mechanisms.
  • Bit-Level Accuracy: Conversions between Base-2 (packed uint64) and Base-7 must be handled carefully (e.g., using np.unpackbits or bitwise shifts) to ensure no loss of precision in the face-ID or refinement levels.

Leveraging High-Performance Kernels

The RangeIndex can be significantly optimized by replacing string-based or byte-unpacking logic with the specialized bitwise kernels found in z7py/z7.py.

1. Performance vs. Memory

The prototype's use of np.unpackbits creates an intermediate boolean representation of the dataset, which is memory-intensive. By using Numba-jitted bitwise operators (such as those in z7py), we can:

  • Zero Allocations: Perform conversions directly on the uint64 stream without creating temporary arrays.
  • CPU Instruction Level Speed: Bitwise shifts and masks map directly to hardware instructions, providing several orders of magnitude faster execution than Python-level loops or numpy sum-reductions.

2. Monotonic Integer Mapping

To ensure that Z7 identifiers are truly contiguous for range arithmetic, we map the hierarchical Z7 structure to a Monotonic Integer: $$Value = (BaseCell \times 7^L) + \sum_{i=1}^L (Digit_i \times 7^{L-i})$$

Why Monotonicity is Critical:

In this context, "monotonic" refers to a strict, one-to-one mapping between the hierarchical grid and a continuous sequence of integers. This property is foundational for several reasons:

  • Hierarchy as a Number Line: Monotonicity ensures that spatial hierarchy translates directly into numerical continuity. If a parent has a monotonic ID of $P$, its children are guaranteed to have perfectly sequential IDs (e.g., $P \times 7 + [0..6]$).
  • Enabling Range Compression: The efficiency of a RangeIndex relies on the ID space being dense. If the mapping were not monotonic, a range [start, end] might contain "ghost" IDs that don't exist in the grid. Monotonicity ensures every integer in a range corresponds to exactly one valid Z7 cell.
  • Constant-Time Position Calculation ($O(1)$): Because the IDs are perfectly sequential, we can calculate the physical offset of any cell within a chunk using simple subtraction: $$\text{Offset} = \text{TargetID} - \text{StartID}$$ This bypasses the need for expensive searches or lookups during data retrieval.

Implementation Specifics:

By using powers of 7 (the refinement factor of Z7), each refinement level occupies exactly its required "slot" in the number line. Placing the BaseCell (0-11) as the most significant component ensures that the index jumps cleanly from one major face of the Earth to the next without overlapping values or numerical "bumps."

3. Numba Implementation Choices

In z7py, several choices maximize performance:

  • @njit(cache=True): Compiles the indexing logic to machine code, bypassing the Python Global Interpreter Lock (GIL) and allowing Dask to execute the range calculation across multiple threads efficiently.
  • Lookup Tables in Memory: Constants like _BASE_CELL_NEIGHBOURS_RAW are stored as contiguous numpy arrays, allowing the Numba compiler to optimize memory access patterns during index traversal.
  • Static Typing: By forcing np.uint64 and np.uint8, we avoid expensive type-checking and boxing/unboxing overhead in the hot path of the RangeIndex creation.

Draft Implementation: Z7MonotonicIndex

The following is a draft implementation of the optimized RangeIndex, leveraging z7py for bitwise speed and dask for scalability.

import numpy as np
import numba as nb
import dask.array as da
import xarray as xr
from z7py import z7_to_monotonic_int, monotonic_int_to_z7

@nb.njit(cache=True)
def get_monotonic_ranges_jit(chunk_uint64, resolution):
    """
    Numba-optimized kernel to find contiguous ranges in a chunk.
    Returns: Array of [length, start_monotonic, end_monotonic]
    """
    n = len(chunk_uint64)
    if n == 0:
        return np.zeros((0, 3), dtype=np.uint64)
    
    # 1. Convert to monotonic integers (O(N) bitwise)
    monotonic = np.empty(n, dtype=np.uint64)
    for i in range(n):
        monotonic[i] = z7_to_monotonic_int(chunk_uint64[i], resolution)
    
    # 2. Identify contiguous segments
    # In a sorted DGGS, most segments will be large blocks
    diffs = np.diff(monotonic)
    gaps = np.where(diffs != 1)[0]
    
    starts = np.concatenate((np.array([0]), gaps + 1))
    ends = np.concatenate((gaps, np.array([n - 1])))
    
    # 3. Build range table
    num_ranges = len(starts)
    result = np.empty((num_ranges, 3), dtype=np.uint64)
    for i in range(num_ranges):
        s_idx, e_idx = starts[i], ends[i]
        result[i, 0] = e_idx - s_idx + 1 # Length
        result[i, 1] = monotonic[s_idx]  # Start ID
        result[i, 2] = monotonic[e_idx]  # End ID
        
    return result

class Z7MonotonicIndex(xr.indexes.Index):
    """
    Xarray Index for Z7 DGGS that uses compressed monotonic ranges.
    Replaces Z7 string/hex identifiers with efficient integer range arithmetic.
    """
    def __init__(self, ranges, resolution, dim="zone_id", coord_name=None):
        # ranges: dask array of [abs_start_pos, start_monotonic_id, end_monotonic_id]
        self.ranges = ranges 
        self.resolution = resolution
        self.dim = dim
        self.coord_name = coord_name or dim

    @classmethod
    def from_z7_array(cls, data_array):
        """Scaffold the index from an existing Xarray DataArray of Z7 uint64."""
        res = data_array.attrs.get("level", 14)
        
        def process_chunk(chunk, block_info=None):
            # Calculate ranges for this specific chunk
            chunk_ranges = get_monotonic_ranges_jit(chunk, res)
            # Offset the local positions to absolute positions
            offset = block_info[0]['array-location'][0][0]
            chunk_ranges[:, 0] += offset 
            return chunk_ranges

        # Lazy calculation of ranges across Zarr chunks
        raw_ranges = data_array.data.map_blocks(
            process_chunk, 
            chunks=(None, 3), 
            dtype=np.uint64
        )
        return cls(raw_ranges, res, dim=data_array.dims[0])

    def sel(self, labels, method=None, tolerance=None):
        """
        The "Reverse Mapping": Z7 ID -> Physical Position
        Uses binary search over ranges + monotonic subtraction (O(log R)).
        """
        target_z7 = labels[self.dim]
        target_m = z7_to_monotonic_int(target_z7, self.resolution)
        
        # 1. Find the range containing target_m (using dask/numpy search)
        # range_idx = search_ranges(self.ranges, target_m)
        
        # 2. Calculate O(1) offset
        # pos = range.abs_start + (target_m - range.start_id)
        
        # ... logic to return IndexSelResult ...
        pass

    def isel(self, indexers):
        """
        The "Forward Mapping": Physical Position -> Z7 ID
        Maps integer slices back to Z7 labels.
        """
        # pos = indexers[self.dim]
        # label_m = ranges.lookup_pos(pos)
        # return monotonic_int_to_z7(label_m, self.resolution)
        pass

Development with Pixi

This project uses Pixi, a modern, high-performance package manager built on the Conda ecosystem. Pixi is particularly well-suited for scientific projects because it provides strict reproducibility through lockfiles and handles complex C/C++ dependencies (like those often found in GIS and numerical libraries) better than standard pip.

Getting Started

  1. Install Pixi:

    curl -fsSL https://pixi.sh/install.sh | bash
    

    (For other platforms, see the installation guide).

  2. Setup and Run Tests: Pixi automatically manages the environment for you. You don't need to pip install anything.

    pixi run test
    
  3. Enter the Development Environment: If you want to run a script or start a REPL within the project environment:

    pixi shell
    python
    

Jupyter Integration

For interactive science and experimentation, you can easily use this Pixi environment with Jupyter:

  1. Add Jupyter to the project:

    pixi add jupyterlab ipykernel
    
  2. Run Jupyter Lab directly:

    pixi run jupyter lab
    
  3. Using with VS Code or external Jupyter: If you prefer using VS Code, simply select the Python interpreter located in .pixi/envs/default/bin/python. VS Code will automatically recognize the environment and its packages.

Why Pixi?

  • Reproducibility: The pixi.lock file ensures that every collaborator uses the exact same versions of all dependencies, including Python itself and system-level libraries.
  • No Activation Needed: Unlike conda or venv, you don't need to manually activate environments to run tasks. pixi run <task> handles it instantly.
  • Unified Conda & Pip: Pixi can install and manage dependencies from both Conda channels (like conda-forge) and PyPI (pip) in the same environment. This solves the "missing package" problem that often plagues tools like poetry or micromamba.
  • Project-Local Environments: Unlike Conda/Micromamba which use global named environments, Pixi stores the environment inside the project directory (.pixi/). This makes it much easier to run isolated experiments with different package versions without polluting your system or forgetting which "env" was for which project.
  • Multi-language: It handles Python, R, C++, and more, making it ideal for projects that bridge high-level analysis and low-level kernels.

License

Apache License 2.0 — see LICENSE and NOTICE.

Citing

If you use z7py in academic work, please cite the IGEO7 paper and the software itself; see CITATION.cff.

References

  1. Kmoch, A., Sahr, K., Chan, W. T., & Uuemaa, E. (2025). IGEO7: A new hierarchically indexed hexagonal equal-area discrete global grid system. AGILE: GIScience Series, 6(32). https://doi.org/10.5194/agile-giss-6-32-2025
  2. Lucas, D., & Gibson, L. (1982). Automated Analysis of imagery. Air Force Office of Scientific Research. https://apps.dtic.mil/sti/tr/pdf/ADA125710.pdf. (No. AFOSRTR830055).
  3. Sahr, K. (2011). Hexagonal discrete global grid systems for geospatial computing. Archives of Photogrammetry Cartography and Remote Sensing, 22, 363–376.
  4. van Roessel, J. W. (1988). Conversion of Cartesian coordinates from and to Generalized Balanced Ternary addresses. Photogrammetric Engineering & Remote Sensing, 54(11), 1565–1570.
  5. White, D., Kimerling, J. A., & Overton, S. W. (1992). Cartographic and Geometric Components of a Global Sampling Design for Environmental Monitoring. Cartography and Geographic Information Systems, 19(1), 5-22. https://doi.org/10.1559/152304092783786636
  6. Wikipedia Contributors. (2025, December 16). Generalized balanced ternary. Wikipedia; Wikimedia Foundation.

Metadata

Release files for z7py 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for z7py 0.1.0
File Size Uploaded
z7py-0.1.0.tar.gz 27.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for z7py 0.1.0
File Interpreter ABI Platform
z7py-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 51.6 kB

Release files / z7py-0.1.0.tar.gz

Download URL z7py-0.1.0.tar.gz
Size 27.1 kB
Tags Source
SHA-256 checksum
How to use checksums
32e1642fe0377b5837b684213a902d32908f18a8ee19fd5dea4fc63f4e776d19
BLAKE2b-256 checksum
How to use checksums
336a690724e6f7dc4c2be99417bf002b3a98cfe0a98a0c479b2e734e5f84a207
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / z7py-0.1.0-py3-none-any.whl

Download URL z7py-0.1.0-py3-none-any.whl
Size 24.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7dacb26174151ca95b6a66a9931eb57aa06efdca989e2d56e0ff779d076e0013
BLAKE2b-256 checksum
How to use checksums
d83cf07ea7c4353d9e38d7ec8e23995038a62101839bdef815a944026c62812c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page