Skip to main content

GTFS-RT Aggregator

This project provides a pipeline for fetching, storing, and aggregating GTFS-RT (General Transit Feed Specification - Realtime) data from multiple providers into Parquet format.

Features

  • Fetch GTFS-RT data from multiple providers and APIs
  • Store individual data files in Parquet format with multiple storage backends (filesystem, Google Cloud Storage, MinIO)
  • Aggregate data files based on configurable time intervals
  • Run fetcher and aggregator services in parallel
  • Configurable via a single TOML configuration file

Requirements

  • Python 3.11+
  • Required Python packages (see requirements.txt):
    • requests
    • gtfs-realtime-bindings
    • pandas
    • pyarrow
    • schedule
    • pydantic
    • google-cloud-storage (optional, for GCS storage)
    • minio (optional, for MinIO storage)

Installation

From PyPI (recommended)

pip install gtfs-rt-aggregator

From Source

  1. Clone this repository:

    git clone https://github.com/GaspardMerten/gtfs-rt-aggregator
    cd gtfs-rt-aggregator
    
  2. Install in development mode:

    pip install -e .
    

Configuration

The pipeline is configured using a TOML configuration file. Here's an example:

# GTFS-RT Configuration File
[storage]
type = "filesystem"  # Options: "filesystem", "gcs", or "minio"
[storage.params]
base_directory = "data"  # Base directory for filesystem storage

# Provider configurations
[[providers]]
name = "ovapi"
timezone = "Europe/Amsterdam"

  [[providers.apis]]
  url = "https://gtfs.ovapi.nl/nl/vehiclePositions.pb"
  services = ["VehiclePosition"]
  refresh_seconds = 20  # Fetch every 20 seconds
  frequency_minutes = 60  # Group files in 60-minute intervals
  check_interval_seconds = 300  # Check for new files every 5 minutes

  [[providers.apis]]
  url = "https://gtfs.ovapi.nl/nl/tripUpdates.pb"
  services = ["TripUpdate"]
  refresh_seconds = 20  # Fetch every 20 seconds

Storage Backend Examples

Google Cloud Storage

[storage]
type = "gcs"
[storage.params]
bucket_name = "my-gtfs-bucket"
base_path = "gtfs-data"  # Optional: subfolder within the bucket
# Authentication is handled via the GOOGLE_APPLICATION_CREDENTIALS environment variable

MinIO Storage

[storage]
type = "minio"
[storage.params]
endpoint = "minio.example.com:9000"
access_key = "YOUR_ACCESS_KEY"
secret_key = "YOUR_SECRET_KEY"
bucket_name = "gtfs-data"
secure = true  # Use HTTPS
base_path = "gtfs-feeds"  # Optional: subfolder within the bucket

Configuration Options

  • storage: Global storage configuration

    • type: Storage backend type ("filesystem", "gcs", or "minio")
    • params: Backend-specific parameters
  • providers: List of GTFS-RT data providers

    • name: Name of the provider (used for directory structure)
    • timezone: Timezone for the provider's data
    • apis: List of API endpoints for this provider
      • url: URL of the GTFS-RT feed
      • services: List of service types to extract from the feed (VehiclePosition, TripUpdate, Alert)
      • refresh_seconds: How often to fetch data from this API
      • frequency_minutes: The time interval (in minutes) for grouping files
      • check_interval_seconds: How often to check for new files to aggregate

Usage

Command Line

Run the pipeline with a configuration file:

gtfs-rt-pipeline configuration.toml

You can adjust the logging level with the --log-level parameter:

gtfs-rt-pipeline configuration.toml --log-level DEBUG

Programmatic Usage

from gtfs_rt_aggregator import run_pipeline_from_toml

# Run pipeline from a TOML file
run_pipeline_from_toml("configuration.toml")

Or with a configuration object:

from gtfs_rt_aggregator.config.loader import load_config_from_toml
from gtfs_rt_aggregator import run_pipeline

# Load configuration
config = load_config_from_toml("configuration.toml")

# Run pipeline
run_pipeline(config)

Project Structure

src/gtfs_rt_aggregator/
  ├── __init__.py                # Package initialization
  ├── pipeline.py                # Main pipeline implementation
  ├── aggregator/                # Aggregation functionality
  ├── config/                    # Configuration loading and validation
  ├── fetcher/                   # GTFS-RT data fetching functionality
  ├── storage/                   # Storage backend implementations
  └── utils/                     # Utility functions and helpers
      ├── cli.py                 # Command-line interface
      └── ...

License

MIT License

Release files for gtfs-rt-aggregator 0.1.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gtfs-rt-aggregator 0.1.6
File Size Uploaded
gtfs_rt_aggregator-0.1.6.tar.gz 26.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gtfs-rt-aggregator 0.1.6
File Interpreter ABI Platform
gtfs_rt_aggregator-0.1.6-py3-none-any.whl Python 3 none any Details

Total release size: 60.4 kB

Release files / gtfs_rt_aggregator-0.1.6.tar.gz

Download URL gtfs_rt_aggregator-0.1.6.tar.gz
Size 26.2 kB
Tags Source
SHA-256 checksum
How to use checksums
6aea68444f0bd66e7de556833b57fc11b54f15755bf5e48fb07d88ca78721add
BLAKE2b-256 checksum
How to use checksums
050372c6ef20d7a44b5d3631c2bd8fd0baeca62c427f60872bc272da83586587
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 30, 2025.

Transparency log

Release files / gtfs_rt_aggregator-0.1.6-py3-none-any.whl

Download URL gtfs_rt_aggregator-0.1.6-py3-none-any.whl
Size 34.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
50ca42e649fe665d92071731051b590a55c2873322e02887472c8f14d8662162
BLAKE2b-256 checksum
How to use checksums
e2fa94ff7997c5c184a39efc32be3e016438ad9dd3ae02c839cf222b5f5e8bc7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 30, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.6 This release

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page