Skip to main content

microsegments

Where do buses and trams linger? microsegments cuts each line into short stretches (30 m by default) and counts how many times a vehicle position is observed in each one, per hour and per day of the week. It works with GTFS plus any vehicle position feed: GTFS-RT VehiclePositions (lat/lon, as CSV, parquet or .pb) or linear-referenced AVL (last stop + distance, like STIB).

Report page: hour player, map coloured by micro-segment, hotspots, hour x micro-segment matrix

Install

pip install "microsegments[plot]"        # plot extra = matplotlib

Python ≥ 3.11. Core dependencies: polars, numpy, shapely, gtfs-parquet (no pandas).

Quick start (offline, synthetic data)

python examples/make_synth.py                     # simulated tram line S1 -> examples/data/
microsegments run examples/synth.toml -o out/     # result.parquet, analysis.json, report.html, matrix_dir*.png

Linear-referenced feeds (STIB vehicle-distance)

[input]
paths = ["vd55/*.parquet"]           # <YYYYMMDD>.parquet (ts, dir, point, dist) + <YYYYMMDD>.snaps.parquet
kind = "stib"                        # or "linear" for any "stop_id + dist_m" table (map columns below)
events = "punctuality/*.parquet"     # optional stop events: passages from both sources, max taken

[input.linear]
id_normaliser = "leading_digits"     # "05766" -> "5766"
feed_length = "auto"                 # rescale feed metres to GTFS link lengths

[gtfs]
path = "gtfs/"                       # zip, .txt directory or gtfs-parquet directory
# dated = "gtfs/{date}"              # one feed per day instead

[select]
route = "55"
dates = "2025-02-17..2025-05-16"
weekdays = [0, 1, 2, 3, 4]

See examples/stib55.toml (runs on the two-day test fixture).

GTFS-RT positions as CSV (lat/lon)

[input]
paths = ["positions/*.csv"]
kind = "latlon"
ts_unit = "s"                        # "s" | "ms" | "us" | "iso"
[input.columns]                      # canonical name = your column
ts = "timestamp"
route_id = "route_id"
vehicle_id = "vehicle_id"
lat = "latitude"
lon = "longitude"
trip_id = "trip_id"                  # optional, helps map matching

Per-vehicle fixes are thinned to one per tick_s (20 s) so counts stay comparable with a snapshot feed. See examples/gtfsrt_csv.toml for every option.

Python

import microsegments as ms

res = ms.run(ms.Config.from_toml("ms.toml"))   # obs -> coverage -> network -> place -> passages
                                                # -> segments -> count -> analyse -> hotspots
res.result            # polars: one row per segment x hour (bands as hour -1 am, -2 pm, -3 day, -4 evening)
res.hotspots          # ranked stretches
ms.export(res, "report.html", lang="fr")       # standalone page (fr / en)
ms.plot.matrix(res, direction_id=0, metric="obs_per_h", days=[0, 1, 2, 3, 4], stops=False)
ms.plot.map(res, 0, "pm", metric="excess_per_passage")
contract = res.contract()                      # the page JSON (also the platform's /api/analysis)

pre = ms.prepare(cfg)                          # reuse placement to tune the segment length
tr = ms.tune.segment_length(pre.placed, pre.segment_fn(), coverage=pre.coverage, passages=pre.passages,
                            pattern_days=pre.network.pattern_days)
ms.plot.tune_curves(tr)

Stages are usable on their own: read_observations, build_network, segment, count, analyse, hotspots, simulate.

CLI

microsegments run CONFIG -o OUT [--lang en] [--segment-m 20] [--dates A..B] [--source "credit"]
microsegments tune CONFIG [--lengths 10,20,30,50] [--sensitivity] [-o OUT]
microsegments hotspots CONFIG [-o hotspots.geojson|.csv|.parquet]
microsegments inspect CONFIG        # coverage per day, GTFS versions, placement / drop counts

What the numbers mean

Every poll (about every 20 s) puts each vehicle in one micro-segment: one observation. One observation is therefore about 20 s of presence.

Metric Definition Read as
obs_per_h (default) Σ observations / Σ covered hours vehicles seen in the segment per hour of polling; /len × 10 for per 10 m
obs_per_passage Σ observations / Σ passages × 20 s ≈ time each vehicle spends there
excess_per_passage obs per passage − the same at 20–23 h, same days extra observations (≈ seconds / 20) lost per vehicle vs free-flowing evening
excess_obs_per_h, log2_ratio see metrics.py

All values are ratios of sums over the selected days (optionally post-stratified by weekday). Stop zones (30 m before to 60 m after each stop) can be masked to bring out signals and junctions. Colour scales are fixed per metric (p98 over 6–21 h), so hours and weekdays compare directly.

Missing data and GTFS versions

  • Hours with < 80 % of polls, a frozen feed (> 5 min), the line absent, or abnormal vehicle counts are dropped from numerator and denominator. Polling gaps are handled by the exposure (covered hours); counts are never imputed.
  • Days with more than 25 % of their 6–21 h dropped are excluded and listed with the reason (res.analysis.excluded_days, coverage calendar in the page).
  • GTFS changes: patterns are rebuilt per service date; link_key / seg_key are stable across versions. Each segment is aggregated only over the days it belongs to the day's main pattern (the reference too). The page shows the most frequent version; others are selectable. A single feed whose calendar does not cover some dates lends them the patterns of the same weekday.

Choosing the segment length

microsegments tune scores L ∈ {10, 15, 20, 30, 40, 50, 75, 100} m (leave-one-day-out Poisson deviance, split-half reliability, hotspot localisation spread and Jaccard) and picks the smallest L with reliability ≥ 0.8, spread ≤ 30 m and deviance within one standard error of the minimum. 30 m is a good default for 20 s polling at urban speeds; go shorter only with dense GTFS-RT fixes and many days. --sensitivity re-runs phase offsets, gap caps, stop zones, references and passage sources.

Hotspots

Stretches where the excess over the evening has a 95 % bootstrap lower bound above 2 s per vehicle, is positive on ≥ 60 % of days, survives Benjamini–Hochberg (q = 0.1) and lasts ≥ 2 hours or a peak band. Adjacent bins merge, never across a stop / running border. Classes: infrastructure (all day), congestion (peak only), mixed. Needs at least 5 included days. Export as GeoJSON with microsegments hotspots CONFIG -o hotspots.geojson.

License

MIT. The bundled Leaflet CSS is BSD-2-Clause.

Metadata

Release files for microsegments 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for microsegments 0.1.1
File Size Uploaded
microsegments-0.1.1.tar.gz 961.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for microsegments 0.1.1
File Interpreter ABI Platform
microsegments-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 1.1 MB

Release files / microsegments-0.1.1.tar.gz

Download URL microsegments-0.1.1.tar.gz
Size 961.6 kB
Tags Source
SHA-256 checksum
How to use checksums
41365ce03d44d37f3c1867ce8b748a02a9246a77b4280daac7a7ae97ffeb7248
BLAKE2b-256 checksum
How to use checksums
d7d182b4fe3a4fb7b659539486ae49c6c2c73f8ae826fa0a85dfed675d3112b0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / microsegments-0.1.1-py3-none-any.whl

Download URL microsegments-0.1.1-py3-none-any.whl
Size 165.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bfd861bac9d406481a587b8c4baaa6317e7e2db392567f16e4dea1c17610a69c
BLAKE2b-256 checksum
How to use checksums
2a6016a8636b6a76abe75b40b0accb7ae5a775cf732827b276edc016fb218c16
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.0

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page