Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

DAS

DAS 1.0a2 segments and annotates audio with Conformer, TCN, TweetyNet, and WhisperSeg models.

Installation

conda create -y -n das python=3.12 uv
conda activate das
uv pip install "das[gui]==1.0a2"

For development from a source checkout, install the test and docs extras too:

uv pip install -e ".[dev,gui,doc]"

The GUI uses xarray-behave. Run das or das gui to open the single-file annotation view and its Train and Predict dialogs.

Usage

Python API:

import das

checkpoint = das.train(
    data_dir="/path/to/audio",
    output_dir="/path/to/run",
)

written_files = das.predict(
    data_dir="/path/to/audio",
    checkpoint=checkpoint,
    existing_annotations="merge",
)

annotations = das.predict(
    audio=waveform,
    samplerate=16_000,
    checkpoint=checkpoint,
)

Prediction writes each annotation CSV next to its audio file by default. Set output_dir="/path/to/predictions" to collect them in one directory. Set output_dir="", or pass raw audio, to return annotation DataFrames instead of writing prediction CSV files. Training still writes to ./ by default. Prediction handles existing output CSV files with existing_annotations="skip", "overwrite" (default), or "merge". The CLI equivalent is --existing-annotations.

Training reads annotated WAV folders directly; no dataset generation step is needed. It also accepts existing DAS .npy, H5, and Zarr training datasets. Other audio formats supported by the current DAS release remain available through the existing dataset-generation path.

Minimal CLI:

das train \
  --data-dir /path/to/audio \
  --output-dir /path/to/run \
  --config /path/to/train.yaml \
  --frontend mel \
  --batch-size 16
das predict \
  --data-dir /path/to/audio \
  --checkpoint /path/to/run/checkpoints/best-epoch=000.ckpt \
  --output-dir /path/to/predictions \
  --config /path/to/predict.yaml \
  --evaluate \
  --syllable-tolerance-ms 20
das predict \
  --data-dir /path/to/audio \
  --checkpoint /path/to/run/checkpoints/best-epoch=000.ckpt \
  --output-dir /path/to/predictions \
  --config /path/to/predict.yaml \
  --existing-annotations merge \
  --output-suffix _frames.csv

CLI flags and YAML use the same flat field names. Config files can be partial, and --config can be provided multiple times. The merge order is defaults, then config files in the order provided, then explicit CLI flags.

--config also accepts built-in config names: fly-pulse, fly, zebra-finch, and tweetynet. fly-pulse selects the 1 ms raw-waveform ConvResNet plus compact TCN. They load through the same merge path as YAML files:

das train \
  --config fly-pulse \
  --data-dir /path/to/audio \
  --output-dir /path/to/run

Legacy DAS models saved as *_model.h5 plus *_params.yaml can be used with the same predict command by pointing --checkpoint at the shared trunk path:

das predict \
  --data-dir /path/to/audio \
  --checkpoint /path/to/test \
  --output-dir /path/to/predictions

They can also be converted once to a native Torch checkpoint that loads through DASModel.load_from_checkpoint():

das convert-legacy \
  --checkpoint /path/to/test \
  --output /path/to/test.ckpt

The converted checkpoint can then be passed to das predict like any native checkpoint.

WhisperSeg uses DAS .ckpt checkpoints for training and prediction. Training requires --initial-model /path/to/converted-whisperseg.ckpt; original WhisperSeg .pt files and Hugging Face model IDs are not loaded directly.

The self-contained whisperseg-aer/v1 .pt bundles can be repacked once, without retraining:

python -m das.whisperseg.convert /path/to/whisperseg-aer.pt /path/to/whisperseg-aer.ckpt

For native checkpoints, das predict uses the saved chunk length, stride, and class names from the trained model. Legacy DAS models use the saved legacy window size, stride, and class labels from the model params. Evaluation works through das predict --evaluate, with optional --split train|val|test for split-backed datasets.

Print a starter config for a subcommand:

das train --print-config

Save the effective config for a run:

das train \
  --data-dir /path/to/audio \
  --output-dir /path/to/run \
  --encoder tcn \
  --save-config /path/to/run-config.yaml

Segment detection uses adjustable low and high hysteresis thresholds. The CLI, YAML config, and GUI Train and Predict dialogs expose these thresholds, event thresholds and spacing, and manual postprocessing settings.

Built-in TCN encoder example:

das train \
  --data-dir /path/to/audio \
  --output-dir /path/to/run \
  --encoder tcn \
  --encoder-hidden-size 32 \
  --encoder-num-layers 4 \
  --encoder-dilations [1,2,4,8,16]

Example flat config:

mode: train
data_dir: /path/to/audio
output_dir: /path/to/run
batch_size: 16
num_time_steps: 2048
frontend_type: raw
encoder_type: tcn
encoder_hidden_size: 16
encoder_num_layers: 2
encoder_dilations: [1, 2, 4, 8]
decoder_type: linear
learning_rate: 0.01
num_epochs: 20

Release files for das 1.0a2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for das 1.0a2
File Size Uploaded
das-1.0a2.tar.gz 118.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for das 1.0a2
File Interpreter ABI Platform
das-1.0a2-py3-none-any.whl Python 3 none any Details

Total release size: 248.9 kB

Release files / das-1.0a2.tar.gz

Download URL das-1.0a2.tar.gz
Size 118.9 kB
Tags Source
SHA-256 checksum
How to use checksums
aee959b1b88027b6a47afcdef15b378f3f55c05ab31434d1498ddaff6daa0c8e
BLAKE2b-256 checksum
How to use checksums
0e13286ece869283c49fa3b129bd003b0d5774d134ab816e8c1778ea6255bb8a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / das-1.0a2-py3-none-any.whl

Download URL das-1.0a2-py3-none-any.whl
Size 130.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
becb6ae267f0339b4c4f40ec522553a1cf4c56e0b4aca93e22efa9a8f02093da
BLAKE2b-256 checksum
How to use checksums
5ff59d3fdb23e6277597862d75118872575ad1634627ac2cf87a62a84737cf80
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

1.0a2 This release

2 release files

0.33.0

2 release files

0.32.9

2 release files

0.32.8

2 release files

0.32.5

2 release files

0.32.2

2 release files

0.32.0

2 release files

0.31.0

2 release files

0.30.1

2 release files

0.30.0

2 release files

0.29.0

2 release files

0.28.5

2 release files

0.28.4

2 release files

0.28.3

2 release files

0.28.1

2 release files

0.28.0

2 release files

0.27.0

2 release files

0.26.9

2 release files

0.26.8

2 release files

0.26.7

2 release files

0.26.6

2 release files

0.26.5

2 release files

0.26.4

2 release files

0.26.3

2 release files

0.26.2

2 release files

0.26.1

2 release files

0.26.0

2 release files

0.25.9

2 release files

0.25.8

2 release files

0.25.7

2 release files

0.25.6

2 release files

0.25.1

2 release files

0.25.0

2 release files

0.24.1

2 release files

0.23.3

2 release files

0.22.8

2 release files

0.22.7

2 release files

0.22.6

2 release files

0.22.5

2 release files

0.22.4

2 release files

0.22.1

2 release files

0.22.0

2 release files

0.21.7

2 release files

0.21.6

2 release files

0.21.5

2 release files

0.21.1

2 release files

0.21.0

2 release files

0.20.5

2 release files

0.20.2

2 release files

0.20.1

2 release files

0.20.0

2 release files

0.19.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page