Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

DAS

DAS 1.0a1 segments and annotates audio with Conformer, TCN, TweetyNet, and WhisperSeg models.

Installation

conda create -y -n das python=3.12 uv
conda activate das
uv pip install "das[gui]==1.0a1"

For development from a source checkout, install the test and docs extras too:

uv pip install -e ".[dev,gui,doc]"

The GUI uses xarray-behave. Run das or das gui to open the single-file annotation view and its Train and Predict dialogs.

Usage

Python API:

import das

checkpoint = das.train(
    data_dir="/path/to/audio",
    output_dir="/path/to/run",
)

written_files = das.predict(
    data_dir="/path/to/audio",
    checkpoint=checkpoint,
    existing_annotations="merge",
)

annotations = das.predict(
    audio=waveform,
    samplerate=16_000,
    checkpoint=checkpoint,
)

Prediction writes each annotation CSV next to its audio file by default. Set output_dir="/path/to/predictions" to collect them in one directory. Set output_dir="", or pass raw audio, to return annotation DataFrames instead of writing prediction CSV files. Training still writes to ./ by default. Prediction handles existing output CSV files with existing_annotations="skip", "overwrite" (default), or "merge". The CLI equivalent is --existing-annotations.

Training reads annotated WAV folders directly; no dataset generation step is needed. It also accepts existing DAS .npy, H5, and Zarr training datasets. Other audio formats supported by the current DAS release remain available through the existing dataset-generation path.

Minimal CLI:

das train \
  --data-dir /path/to/audio \
  --output-dir /path/to/run \
  --config /path/to/train.yaml \
  --frontend mel \
  --batch-size 16
das predict \
  --data-dir /path/to/audio \
  --checkpoint /path/to/run/checkpoints/best-epoch=000.ckpt \
  --output-dir /path/to/predictions \
  --config /path/to/predict.yaml \
  --evaluate \
  --syllable-tolerance-ms 20
das predict \
  --data-dir /path/to/audio \
  --checkpoint /path/to/run/checkpoints/best-epoch=000.ckpt \
  --output-dir /path/to/predictions \
  --config /path/to/predict.yaml \
  --existing-annotations merge \
  --output-suffix _frames.csv

CLI flags and YAML use the same flat field names. Config files can be partial, and --config can be provided multiple times. The merge order is defaults, then config files in the order provided, then explicit CLI flags.

--config also accepts built-in config names: fly-pulse, fly, zebra-finch, and tweetynet. fly-pulse selects the 1 ms raw-waveform ConvResNet plus compact TCN. They load through the same merge path as YAML files:

das train \
  --config fly-pulse \
  --data-dir /path/to/audio \
  --output-dir /path/to/run

Legacy DAS models saved as *_model.h5 plus *_params.yaml can be used with the same predict command by pointing --checkpoint at the shared trunk path:

das predict \
  --data-dir /path/to/audio \
  --checkpoint /path/to/test \
  --output-dir /path/to/predictions

They can also be converted once to a native Torch checkpoint that loads through DASModel.load_from_checkpoint():

das convert-legacy \
  --checkpoint /path/to/test \
  --output /path/to/test.ckpt

The converted checkpoint can then be passed to das predict like any native checkpoint.

WhisperSeg uses DAS .ckpt checkpoints for training and prediction. Training requires --initial-model /path/to/converted-whisperseg.ckpt; original WhisperSeg .pt files and Hugging Face model IDs are not loaded directly.

The self-contained whisperseg-aer/v1 .pt bundles can be repacked once, without retraining:

python -m das.whisperseg.convert /path/to/whisperseg-aer.pt /path/to/whisperseg-aer.ckpt

For native checkpoints, das predict uses the saved chunk length, stride, and class names from the trained model. Legacy DAS models use the saved legacy window size, stride, and class labels from the model params. Evaluation works through das predict --evaluate, with optional --split train|val|test for split-backed datasets.

Print a starter config for a subcommand:

das train --print-config

Save the effective config for a run:

das train \
  --data-dir /path/to/audio \
  --output-dir /path/to/run \
  --encoder tcn \
  --save-config /path/to/run-config.yaml

Segment detection uses adjustable low and high hysteresis thresholds. The CLI, YAML config, and GUI Train and Predict dialogs expose these thresholds, event thresholds and spacing, and manual postprocessing settings.

Built-in TCN encoder example:

das train \
  --data-dir /path/to/audio \
  --output-dir /path/to/run \
  --encoder tcn \
  --encoder-hidden-size 32 \
  --encoder-num-layers 4 \
  --encoder-dilations [1,2,4,8,16]

Example flat config:

mode: train
data_dir: /path/to/audio
output_dir: /path/to/run
batch_size: 16
num_time_steps: 2048
frontend_type: raw
encoder_type: tcn
encoder_hidden_size: 16
encoder_num_layers: 2
encoder_dilations: [1, 2, 4, 8]
decoder_type: linear
learning_rate: 0.01
num_epochs: 20

Release files for das 1.0a1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for das 1.0a1
File Size Uploaded
das-1.0a1.tar.gz 119.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for das 1.0a1
File Interpreter ABI Platform
das-1.0a1-py3-none-any.whl Python 3 none any Details

Total release size: 249.0 kB

Release files / das-1.0a1.tar.gz

Download URL das-1.0a1.tar.gz
Size 119.0 kB
Tags Source
SHA-256 checksum
How to use checksums
2363b3c0e1210df27dce86aa2503bfeb05a9a81cde6f6e845579211897ec88ed
BLAKE2b-256 checksum
How to use checksums
86e8082fafa4eb18852f647b0618fca88eade642ad854d3421feb75c3c75dca6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / das-1.0a1-py3-none-any.whl

Download URL das-1.0a1-py3-none-any.whl
Size 130.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2d6f2e3b459ce21a0be6bc46257e0ad6d2d847fffe7cea4ba126702169134177
BLAKE2b-256 checksum
How to use checksums
3d5a3f58ac4713aa9619b117afe8cf59c9652ef9c7de88dc75e212517ce8afc3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

1.0a1 This release

2 release files

0.33.0

2 release files

0.32.9

2 release files

0.32.8

2 release files

0.32.5

2 release files

0.32.2

2 release files

0.32.0

2 release files

0.31.0

2 release files

0.30.1

2 release files

0.30.0

2 release files

0.29.0

2 release files

0.28.5

2 release files

0.28.4

2 release files

0.28.3

2 release files

0.28.1

2 release files

0.28.0

2 release files

0.27.0

2 release files

0.26.9

2 release files

0.26.8

2 release files

0.26.7

2 release files

0.26.6

2 release files

0.26.5

2 release files

0.26.4

2 release files

0.26.3

2 release files

0.26.2

2 release files

0.26.1

2 release files

0.26.0

2 release files

0.25.9

2 release files

0.25.8

2 release files

0.25.7

2 release files

0.25.6

2 release files

0.25.1

2 release files

0.25.0

2 release files

0.24.1

2 release files

0.23.3

2 release files

0.22.8

2 release files

0.22.7

2 release files

0.22.6

2 release files

0.22.5

2 release files

0.22.4

2 release files

0.22.1

2 release files

0.22.0

2 release files

0.21.7

2 release files

0.21.6

2 release files

0.21.5

2 release files

0.21.1

2 release files

0.21.0

2 release files

0.20.5

2 release files

0.20.2

2 release files

0.20.1

2 release files

0.20.0

2 release files

0.19.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page