Skip to main content

This repository contains utility functions and interfaces that can be used to interact with the DISCOVER framework.

Project description

DISCOVER-Utils

PyPI version Python License: GPL v3 Documentation

DISCOVER-Utils is a Python utility package for data handling, processing, and annotation of multimedia data. It is designed to work with the DISCOVER framework or as a stand-alone library.

Features

  • Data handling — Unified access to streams (audio, video, sensor data) and annotations (discrete, continuous) via file, MongoDB, or URL backends
  • Multiple video backends — Choose between decord, imageio, moviepy, or pyav for video decoding
  • Dataset management — Iterate over multi-session datasets with DatasetManager and DatasetIterator
  • Processing pipeline — Run DISCOVER server modules from the command line for feature extraction and prediction
  • SSI compatibility — Read and write SSI trainer files and XML configurations

Installation

pip install hcai-discover-utils

Optional video backends

# Fast video decoding with decord
pip install hcai-discover-utils[decord]

# PyAV (FFmpeg bindings)
pip install hcai-discover-utils[pyav]

# MoviePy
pip install hcai-discover-utils[pymovie]

Getting Started

Command-line tools

Process data with DISCOVER server modules:

du-process \
  --dataset "my_dataset" \
  --db_host "127.0.0.1" --db_port "27017" \
  --db_user "user" --db_password "pass" \
  --trainer_file_path "path/to/trainer.trainer" \
  --sessions '["session1", "session2"]' \
  --data '[{"src": "db:anno", "scheme": "transcript", "annotator": "test", "role": "testrole"}]'

File mode (no database)

Read inputs and write outputs directly from/to disk, without a NOVA database. Use file: sources and supply a path via uri (static, single session) or uri_template (per-session paths via {dataset} and {session} placeholders):

du-process \
  --dataset "my_study" \
  --trainer_file_path "path/to/trainer.trainer" \
  --sessions '["session_a", "session_b"]' \
  --data '[
    {
      "id": "video",
      "type": "input",
      "src": "file:stream:video",
      "uri_template": "/data/{dataset}/{session}/video.mp4"
    },
    {
      "id": "valence",
      "type": "output",
      "src": "file:annotation:continuous",
      "uri_template": "/outputs/{dataset}/{session}/valence.annotation",
      "sample_rate": 30,
      "min_val": -1,
      "max_val": 1
    }
  ]'

Each session resolves its own input and output paths. Output annotation descriptors may carry scheme metadata that is used when no annotation file exists yet:

  • file:annotation:continuous: sample_rate, min_val, max_val (defaults: 1, 0, 1).

  • file:annotation:discrete: classes as a map from class id to a dict of per-class XML attributes (typically name, optionally color, etc.). The outer key is the canonical id; the writer injects it into the XML automatically. For example:

    // fragment of a data description entry
    "classes": {
      "0": {"name": "neutral", "color": "#888"},
      "1": {"name": "happiness", "color": "#ffd700"}
    }
    

    Legacy {id: name} strings are also accepted and normalized internally to the canonical form.

This matters for modules that resample continuous outputs to the scheme's sample_rate — without explicit metadata, outputs default to 1 Hz.

Notes:

  • uri and uri_template are filesystem paths (absolute or relative to the working directory). There is no implicit base directory.
  • uri_template placeholders that reference {dataset} or {session} must have non-empty values; otherwise resolve_file_uri raises ValueError.
  • uri_template takes precedence over uri when both are present.

Python API

from discover_utils.data.provider.data_manager import DatasetManager

# Set up a dataset manager for your sessions
dm = DatasetManager(
    dataset="my_dataset",
    db_host="127.0.0.1",
    db_port=27017,
    db_user="user",
    db_password="pass",
    sessions=["session1"],
    data_description=[...],
)

Exploring a database

When you don't already know the dataset, scheme, or annotator names, use the exploration queries on the MongoDB handler to discover what's available. They are read-only and return lightweight metadata (no stream files or annotation labels are loaded):

from discover_utils.data.handler.nova_db_handler import NovaDBHandler

db = NovaDBHandler(
    db_host="127.0.0.1", db_port=27017, db_user="user", db_password="pass",
)

db.list_datasets()                      # datasets you can read
db.list_sessions("my_dataset")          # session names
db.list_scheme_names("my_dataset")      # scheme names
db.list_schemes("my_dataset")           # [{name, type}, ...]
db.list_roles("my_dataset")             # role names
db.list_annotators("my_dataset")        # annotator names
db.list_streams("my_dataset")           # stream metadata (no files read)

# which annotations exist for a session, without loading any labels
db.list_annotations("my_dataset", session="session1")
# -> [{session, annotator, role, scheme, isFinished, isLocked}, ...]

The same methods are inherited by AnnotationHandler, StreamHandler, and SessionHandler, so you can explore and load through a single handler instance.

Documentation

Full API documentation is available at hcmlab.github.io/discover-utils/docbuild/.

Citation

If you use DISCOVER or DISCOVER-Utils in your research, please cite:

@article{hallmen2025discover,
  title     = {DISCOVER: a Data-driven Interactive System for Comprehensive
               Observation, Visualization, and ExploRation of human behavior},
  author    = {Hallmen, Tobias and Schiller, Dominik and others},
  journal   = {Frontiers in Digital Health},
  volume    = {7},
  pages     = {1638539},
  year      = {2025},
  publisher = {Frontiers}
}

License

This project is licensed under the GNU General Public License v3.0.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hcai_discover_utils-1.2.1.tar.gz (87.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hcai_discover_utils-1.2.1-py3-none-any.whl (92.7 kB view details)

Uploaded Python 3

File details

Details for the file hcai_discover_utils-1.2.1.tar.gz.

File metadata

  • Download URL: hcai_discover_utils-1.2.1.tar.gz
  • Upload date:
  • Size: 87.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for hcai_discover_utils-1.2.1.tar.gz
Algorithm Hash digest
SHA256 e14037590cc635e98b03e52b606c0a39046075f9a84e9d105c660c87b53ac325
MD5 116319f1c1b74e2783783c3e8a110baf
BLAKE2b-256 3d9788f4ace769b5e824e41c782b9fdf732c728c4759018262d0927cf759ef39

See more details on using hashes here.

File details

Details for the file hcai_discover_utils-1.2.1-py3-none-any.whl.

File metadata

File hashes

Hashes for hcai_discover_utils-1.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 2d62bfe56159d3a0371b33adcbb0d7b3df04fcf11c46c260d7788e629d46e6c9
MD5 723e3c2e59e732a2d801c76b340ab11e
BLAKE2b-256 fdb84dd32024c2e3aa7b95eabe6d3b1d612384c08571a8fd027795949f86fdba

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page