Skip to main content

📦 emota_loader — Python Dataloader for EmoTa Dataset

EmoTa: A Tamil Emotional Speech Dataset (Thevakumar et al., CHiPSAL 2025) is the first open-access emotional speech corpus in Tamil, designed to capture the dialectal diversity of Sri Lankan Tamil speakers1.

Statistic Value
Utterances 936 (22 speakers × 19 sentences × 5 emotions)
Speakers 22 native Sri Lankan Tamil (11 male, 11 female)
Sentences 19 semantically neutral sentences
Emotions angry, happy, sad, fear, neutral
Inter-annotator Agreement Fleiss’ Kappa = 0.74
Baseline F1 Scores XGBoost: 0.91, Random Forest: 0.90

🔧 Installation

You can install the package from PyPI using:

pip install emota_loader

Make sure to download the EmoTa dataset separately and point the loader to its root directory.


🚀 Sample Usage

from emota_loader import EmoTaDataset

# Point to extracted dataset root folder
dataset = EmoTaDataset(root_dir="path/to/EmoTa").samples

print(f"Loaded {len(dataset)} samples")

sample = dataset[0]
print(f"  Audio Path      : {sample.audio_path}")
print(f"  Speaker ID      : {sample.speaker_id}")
print(f"  Speaker Gender  : {sample.speaker_gender}")
print(f"  Speaker Age     : {sample.speaker_age}")
print(f"  Speaker Region  : {sample.speaker_region}")
print(f"  Sentence ID     : {sample.sentence_id}")
print(f"  Transcript      : {sample.transcript}")
print(f"  Emotion         : {sample.emotion}")

Example Output

Loaded 936 samples

  Audio Path      : EmoTa/19_18_ang.wav
  Speaker ID      : 19
  Speaker Gender  : male
  Speaker Age     : 25
  Speaker Region  : northern
  Sentence ID     : 18
  Transcript      : நான் உன்னை சந்திக்க வேண்டும்.
  Emotion         : angry

📄 Citation

Please cite the dataset as:

@inproceedings{thevakumar-etal-2025-emota,
  title = "{E}mo{T}a: A {T}amil Emotional Speech Dataset",
  author = "Thevakumar, Jubeerathan and Thavarasa, Luxshan and Sivatheepan, Thanikan and Kugarajah, Sajeev and Thayasivam, Uthayasanker",
  booktitle = "Proceedings of the First Workshop on Challenges in Processing South Asian Languages (CHiPSAL 2025)",
  year = "2025",
  pages = "193--201",
  address = "Abu Dhabi, UAE",
  publisher = "International Committee on Computational Linguistics"
}

📘 License

Academic use only — see the EmoTa dataset license for details.


  1. Thevakumar, J., Thavarasa, L., et al. (2025). EmoTa: A Tamil Emotional Speech Dataset. Proceedings of CHiPSAL 2025. ↩

Metadata

Release files for emota-loader 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for emota-loader 2.0.0
File Size Uploaded
emota_loader-2.0.0.tar.gz 4.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for emota-loader 2.0.0
File Interpreter ABI Platform
emota_loader-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 10.1 kB

Release files / emota_loader-2.0.0.tar.gz

Download URL emota_loader-2.0.0.tar.gz
Size 4.7 kB
Tags Source
SHA-256 checksum
How to use checksums
fdfbb340584e203df280f72910683e7c1e538dbbac25fc1f19832e3b09aebc96
BLAKE2b-256 checksum
How to use checksums
a08b3b9d8ad3154fa9b52c8af7521f82daf631e11bf8303f57eaf586662657e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 3, 2025.

Transparency log

Release files / emota_loader-2.0.0-py3-none-any.whl

Download URL emota_loader-2.0.0-py3-none-any.whl
Size 5.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
888ce4f0b2d4049f80c32f97a67940652aa835c8a04953d8252c7cf5954e4688
BLAKE2b-256 checksum
How to use checksums
85fef28c98a581bb1036ae8f261c1052dc35329956a7dee60f8852ea3e58f34b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 3, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page