Skip to main content

Library with Python models derived from the page package world.opensemantic.base

Project description

PyPI-Server Coveralls

opensemantic.base

Python models and controllers for the world.opensemantic.base page package.

Builds on oold-python (BaseController, LinkedBaseModel, cast(), _types registry) and opensemantic (OswBaseModel, compute_scoped_uuid).

Overview

  • Auto-generated Pydantic models (v1 and v2): Database, WebService, DataTool, DataChannel, Person, Organization, etc.
  • DataToolController - generic controller for any DataTool
  • TimeSeriesDatabaseController - SQLite and PostgREST backends for time series storage

Architecture

opensemantic.base/
  _model.py            # auto-generated v2 Pydantic models (DO NOT EDIT)
  _controller_mixin.py # DataToolMixin, TimeSeriesDatabaseController mixins
  _controller.py       # v2 controllers
  __init__.py           # re-exports model + controller classes
  v1/                   # same structure for Pydantic v1

Data model

A DataTool instance is typed by the DataTool category, carries its data channels (each typed by a characteristic), and references its storage:

graph LR
    probeA["probeA<br/>(DataTool instance)"]
    probeA -- HasType --> DataTool["DataTool"]
    probeA -- HasStorageLocation --> archive["archive<br/>(Database)"]
    probeA -- HasDataChannel --> temp[":temp"]
    probeA -- HasDataChannel --> pressure[":pressure"]
    temp -- HasCharacteristic --> Temperature["Temperature"]
    pressure -- HasCharacteristic --> Pressure["Pressure"]

DataToolController

Extends DataTool with channel management, subdevice traversal, and data archiving.

from opensemantic.base import DataToolController

tool = DataToolController(
    name="sensor",
    label=[...],
    data_channels=[ch1, ch2],
    storage_locations=[db],
    auto_archive=True,  # auto-creates archive DB from storage_locations[0]
)

tool.get_all_channels()       # recursive across subdevices
tool.get_channel_owner(ch)    # find which device owns a channel
tool.to_json()                # only model fields (controller fields stripped)

Auto-archive from storage_locations

When auto_archive=True and no explicit archive_database is set, the controller auto-creates a LocalTimeSeriesDatabaseController from the first storage_locations entry (resolved via oold backend).

Typed read/write

Channels with a characteristic IRI enable typed serialization:

# Write: converts to base unit, strips defaults for compact storage
await tool.store_typed_data(DataToolMixin.StoreTypedDataParams(
    tool_osw_id=tool.get_osw_id(),
    rows=[DataToolMixin.TypedDataRow(ts=now, channel=ch, value=Temperature(value=300.0))],
))

# Read: resolves characteristic class via oold's _types registry
results = await tool.read_typed_data(DataToolMixin.ReadTypedDataParams(
    tool_osw_id=tool.get_osw_id(), channel=ch,
))
# results[0] is a Temperature instance with defaults restored

Subobject ID auto-computation

Inline subobject osw_id fields are auto-prefixed with the parent's osw_id:

Parent:  OSW<parent_uuid>
Channel: OSW<parent_uuid>#OSW<channel_uuid>

Fields with range in json_schema_extra are references to separate entities and are not prefixed.

Unloaded characteristic warning

On init, DataToolController checks if channel characteristic IRIs are present in oold's _types registry. Missing entries produce a warning with guidance to import the corresponding package.

TimeSeriesDatabaseController

Abstract base for time series storage, with SQLite and PostgREST implementations.

from opensemantic.base import LocalTimeSeriesDatabaseController

db = LocalTimeSeriesDatabaseController(name="archive", label=[...], db_path="./data.sqlite")
await db.create_tool(params)
await db.write_tool_channel_raw(params)
await db.read_tool_channel_raw(params)

DataToolView (Dashboard UI)

Interactive dashboard for visualizing archived time series data from DataToolControllers.

Features:

  • Wunderbaum TreeGrid for channel selection with characteristic metadata
  • Stacked Bokeh plots grouped by characteristic (temperature, pressure, etc.)
  • Unit conversion via dropdown (e.g. K to C, Pa to hPa)
  • Composite channel splitting (e.g. AirQuality into temperature + pressure sub-plots)
  • Text channel log console with timestamped entries
  • Configurable via JsonEditor (grouping, auto-fetch, row limit, cache)

Archive Demo

Interactive demo: channel selection, plotting, unit switching

Archive Dashboard

Channel selection with stacked plots grouped by characteristic

Unit Switching

Unit conversion via dropdown (K to C)

Log Console

Text channel log console with timestamped entries

from opensemantic.base.view import DataToolView
from opensemantic.base.view._config import DashboardConfig, PlotConfig

view = DataToolView(
    controllers=[ctrl],
    config=DashboardConfig(lang="en", plot=PlotConfig(auto_fetch=True)),
    title="My Dashboard",
)
view.servable()  # for panel serve

See examples/datatool_dashboard.py for a full working example.

To regenerate the screenshots after UI changes, see docs/generate_screenshots.py.

Server-side downsampling

Large time series are downsampled on the server so the plot only transports the points the current zoom level can show. This is driven by PlotConfig.downsample:

from opensemantic.base.view._config import DashboardConfig, PlotConfig, DownsampleConfig

config = DashboardConfig(
    plot=PlotConfig(
        downsample=DownsampleConfig(
            enabled=True,      # downsample when the backend supports it
            max_points=2000,   # target points per channel
            method="auto",     # auto | sample | average | minmax
            edge_anchors=True, # keep the window's first/last real datapoints
        )
    )
)

It only engages on a PostgREST/TimescaleDB backend (the downsample_tool_channel RPC). On a SQLite/local backend, the RPC being absent, or any error, the read silently falls back to the full-resolution path - downsampling never breaks a read. The DataToolView plot also reloads at a finer resolution when you zoom in.

Strategies (N = number of buckets):

  • sample (default): one real datapoint nearest each bucket center. Schema agnostic; works for scalar and composite channels. N rows.
  • average: structure-preserving deep average per bucket (every numeric leaf averaged, non-numeric keys carried), bucket-center timestamp. N rows.
  • minmax: the real min and max datapoint of every numeric sub-characteristic per bucket. Scalar: 2N real rows; composite: the real per-leaf extremes. Best at preserving spikes.
  • auto: minmax for numeric channels, sample for text channels.

average/minmax skip non-numeric leaves (comments, labels) and fall back to sample when a channel has no numeric leaf at all.

Unit normalization caveat: average and minmax compare and combine the bare stored numbers per leaf, so they are only correct when all stored values of a leaf share the same unit. The archive stores base-unit-normalized values, but data ingested without normalization (mixed units in one channel) will produce wrong average/minmax results. sample returns whole real rows and is unaffected.

The RPC lives in pgstack's postgres/config/optional/100_init_tsdb_schema.sql (uses only core, Apache-2 TimescaleDB; no toolkit dependency). It is created at database init; apply it manually (psql -f / pgAdmin) on an already-running cluster. To measure the speedup, see benchmarks/bench_downsample.py.

examples/downsample_demo.py is an interactive demo: four channels carry the same 100k-point signal (with narrow spikes), one per strategy - the channels are named raw, sample, average, minmax. Selecting raw loads slowly with full detail; sample/average load fast but drop the spikes; minmax keeps them. Box-zoom into a flat stretch and click "Load current range" to re-fetch that window at finer resolution - the hidden spikes reappear on sample/average. Needs a running pgstack with the RPC applied; seed once with python examples/downsample_demo.py, then panel serve examples/downsample_demo.py.

Downsampling demo

Strategy comparison

Full window, all four channels: raw and minmax keep the spikes; sample and average smooth them away at the coarse full-window resolution.

Zoomed, before reload

Box-zoomed around a spike on the sample channel: only the coarse full-window points are shown, so the spike is still hidden.

Zoom reveals the peak

After "Load current range": the window is re-fetched at finer buckets and the hidden spike reappears. The toolbar reset returns to the full window.

To regenerate these, see docs/generate_downsample_screenshots.py.

ProcessObjectView (Process/Object Dashboard UI)

Where DataToolView is centered on data tools, ProcessObjectView is centered on the objects (samples, specimens, ... - any Item) that pass through processes. It answers "how did measurement X compare across the runs my objects went through?" by overlaying repeated runs on a common, time-normalized axis.

The view walks these relations - an object is a process input, a process uses data tools, and a data tool has channels:

graph LR
    process[":process"]
    process -- HasInput --> sample["sample (Item)<br/>the 'object'"]
    process -- HasTool --> probeA["probeA<br/>(DataTool)"]
    process -- HasStartDateAndTime --> startTime["start_date_time"]
    process -- HasEndDateAndTime --> endTime["end_date_time"]
    probeA -- HasDataChannel --> channels[":temp, :pressure"]

start_date_time / end_date_time define each run's plot window; the view normalizes every run to its first data point (t=0).

Process Demo

Pick objects (tree 1) and a channel under a process type (tree 2); each run is overlaid from t=0.

Two trees drive the plot:

  1. Objects - the Item instances to compare.
  2. Process types -> channels - for each process type the objects went through, the channels of the data tools used in those runs.

Tree 2 is aggregated for selection only (tick once instead of ticking the same channel on every run and tool). The aggregation is co-presence aware:

  • data tools of the same type that run together in a process stay as separate entries (distinct measurement points);
  • data tools that only ever appear in different runs are treated as drop-in replacements and merged into one … [n channels] entry (the actual channels are listed in its tooltip).

Process trees

Evacuation ran both probes together (separate per-instance entries); Heating swapped the probe between runs (merged DataTool/… [2 channels] entries).

Selecting an object + a channel entry plots every real channel that object has data on - fanning out across process runs and co-present tools. Each line is normalized to its own run (first data point at t=0, x-axis in seconds), grouped by characteristic, and gets a distinct legend entry (object / process / data tool / channel).

Process overlay

One channel selection fans out to two runs of the same sample, overlaid from t=0 for comparison.

Heating drop-in

Adding Heating/DataTool/temp and …/pressure on top of the Evacuation selection overlays both processes, grouped into separate Temperature and Pressure plots; the merged Heating entries resolve to probe A for Sample 1 and probe B for Sample 2 (drop-in replacements compared across objects).

from opensemantic.base.view import ProcessObjectView
from opensemantic.base.view._config import DashboardConfig

view = ProcessObjectView(
    objects=objects,        # list[Item]
    processes=processes,    # list[Process] (filtered to those with start+end
                            # time and >=1 DataTool whose data you can load)
    controllers=controllers,  # list[DataToolController], matched to process tools
    config=DashboardConfig(lang="en"),
    title="Process / Object Archive View",
)
view.servable()  # for panel serve

DataToolView and ProcessObjectView share their plot/unit/config machinery via BaseDataView, and both support embeddable=True (exposing sidebar_cards / main_cards) so a host app can combine them.

See examples/process_dashboard.py for a full working example, and docs/generate_process_screenshots.py to regenerate these screenshots.

Installation

pip install opensemantic.base            # models only
pip install opensemantic.base[controller] # + aiosqlite, postgrest
pip install opensemantic.base[view]       # + panel, bokeh, panelini, pint

Testing

pytest tests/test_controller.py

PostgREST integration tests require a running pgstack instance. To enable them:

  1. Start pgstack: docker compose -f docker-compose.yml -f docker-compose.example-tsdb.override.yml up -d
  2. Copy tests/.env.example to tests/.env and fill in TEST_PGRST_URL and TEST_PGRST_JWT_SECRET
  3. Run tests - PostgREST tests are skipped unless both env vars are set

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

opensemantic_base-0.42.8.post1000002004002.tar.gz (22.3 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

File details

Details for the file opensemantic_base-0.42.8.post1000002004002.tar.gz.

File metadata

File hashes

Hashes for opensemantic_base-0.42.8.post1000002004002.tar.gz
Algorithm Hash digest
SHA256 0277a99cce27f8bf03d5672ef471607264e7027389a4fa0584813856d51813ed
MD5 034f32a445ec0265764ec94767fa0138
BLAKE2b-256 9e2dacd1b7d146b13b3f641cf27dd2a5390cdc1dd6254940233bd57eb6b2aa5b

See more details on using hashes here.

File details

Details for the file opensemantic_base-0.42.8.post1000002004002-py3-none-any.whl.

File metadata

File hashes

Hashes for opensemantic_base-0.42.8.post1000002004002-py3-none-any.whl
Algorithm Hash digest
SHA256 e1e12ed2675ba291e4a626473cc80c653324e333d7140d786b5fb067afd37aa2
MD5 fa41b3e41236991a502920c70293d518
BLAKE2b-256 b05ab02c7155663353fdfcd8c6c52d977c8a699a6063c0221672f83fc3fad1d7

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page