Skip to main content

flatdata-py

Build Status

Python 3 implementation of flatdata.

Running the tests

python3 -m pytest

Basic usage

Once you have created a flatdata schema file, you can generate a Python module to read your existing flatdata archive:

flatdata-generator --gen py --schema locations.flatdata --output-file locations.py

Performance tips

flatdata-py supports two data access patterns with very different performance characteristics on large archives.

Iterating over a vector yields one Python object per element. Each field access unpacks bits from the underlying memory-mapped data. This is fine for accessing individual elements or small ranges, but has significant per-element overhead for bulk operations:

count = sum(1 for x in archive.links if x.speed_limit > 100)

For bulk operations, use the vectorized access methods that read fields directly into NumPy arrays:

# single column access, returns a pandas DataFrame
df = archive.links.speed_limit
count = len(df[df['speed_limit'] > 100])

# full NumPy structured array with all fields
arr = archive.links.to_numpy()
count = int(np.sum(arr['speed_limit'] > 100))

# slices work too
arr = archive.links[1000:2000].to_numpy()
df = archive.links[::10].to_data_frame()
  • Use vector.field_name (column access) when you only need one or a few fields.
  • Use vector.to_numpy() or vector.to_data_frame() when you need all fields at once.
  • Use vector[i].field for random access to individual elements.
  • The underlying data is memory-mapped; the OS pages it from disk on demand. Vectorized results are materialized as NumPy arrays in RAM.

Using the inspector

flatdata-py comes with a handy tool called the flatdata-inspector to inspect the contents of an archive:

  • from the flatdata-py source directory:
./inspector.py
# or
python3 -m flatdata.lib.inspector
  • if you want to install flatdata-py:
pip3 install flatdata-py[inspector]  # the inspector feature requires IPython
flatdata-inspector -p /path/to/my/flatdata.archive

Using the writer

flatdata-writer is an addition to flatdata-py that can create flatdata archives from a flatdata schema, with the following limitations:

  • does not allow adding additional sub-archives to an existing archive

  • supports only bulk-writing (no streaming)

  • not optimized for performance

  • from the flatdata-py source directory

./writer.py --schema archive.flatdata --output-dir testdir --json-file data.json --resource-name resourcename
#or
python3 -m flatdata.lib.writer --schema archive.flatdata --output-dir testdir --json-file data.json --resource-name resourcename

Note that the flatdata-writer CLI tool can only write one resource at a time. For archives that have multiple non-optional resources, the tool has to be executed separately for each resource. Only after all resources have been written can the archive be opened.

  • if you want to install flatdata-py:
pip3 install flatdata-py[writer]
flatdata-writer --schema archive.flatdata --output-dir testdir --json-file data.json --resource-name resourcename

Metadata

Release files for flatdata-py 0.4.12

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for flatdata-py 0.4.12
File Size Uploaded
flatdata_py-0.4.12.tar.gz 16.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for flatdata-py 0.4.12
File Interpreter ABI Platform
flatdata_py-0.4.12-py3-none-any.whl Python 3 none any Details

Total release size: 40.6 kB

Release files / flatdata_py-0.4.12.tar.gz

Download URL flatdata_py-0.4.12.tar.gz
Size 16.8 kB
Tags Source
SHA-256 checksum
How to use checksums
c4f1736270dcea1e4da3f62b5f75b1146d27a372b696a83597a90657e61810f5
BLAKE2b-256 checksum
How to use checksums
d6186836ca307821a6b696df805a43416ae1ca90bf8fd45a92cb1d510e179179
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.8

Release files / flatdata_py-0.4.12-py3-none-any.whl

Download URL flatdata_py-0.4.12-py3-none-any.whl
Size 23.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ec4f2968212ceeeed787fe07e9772193cb6cde2ae99329abef879783769568f3
BLAKE2b-256 checksum
How to use checksums
e78333fb057bd2968440353a9ff515f09f72ac4f3e7289bd9291e698dc9dd0fd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.8

Release history Release notifications | RSS feed

This release

0.4.12 This release

2 release files

0.4.11

3 release files

0.4.10

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page