A benchmark dataset for Arabic to Hindi transliteration.
Project description
AH-Translit_Bench: Arabic to Hindi Transliteration Benchmark Dataset
Repository Description
AH-Translit_Bench is a 2000-entry Arabic to Hindi transliteration benchmark dataset covering Al-Quran, bibliographical, and Modern Standard Arabic (MSA) domains. Provided as a pip-installable Python package for model testing and research.
Dataset Overview
A benchmark dataset for Arabic to Hindi transliteration, providing parallel transliterated text pairs across diverse domains.
Dataset Curators
This dataset was proudly curated by:
- Vilal Ali
Research Scholar at Data Sciences And Analytics Centre, IIIT Hyderabad,
Email: vilal.ali@research.iiit.ac.in - Mohd Hozaifa Khan
Research Scholar at Center for Visual Information Technology (CVIT) IIIT Hyderabad Email: mohd.hozaifa@research.iiit.ac.in
Dataset Usage
This dataset is ideal for training, validating, and testing models for Arabic to Hindi transliteration tasks.
Content Type
Text, specifically Arabic source text and its corresponding Hindi transliterated target text.
Dataset Description
The AH-Translit_Bench dataset provides a collection of parallel text pairs for evaluating and developing Arabic to Hindi transliteration systems. It comprises text from three distinct domains: Al-Quran, bibliographical entries, and Modern Standard Arabic (MSA), ensuring a broad coverage for robust model development. Each entry consists of an Arabic string and its manually curated transliteration into Hindi (Devanagari script). The dataset's structure allows for straightforward loading and integration into various machine learning frameworks.
Version Name
AH-Translit_Bench_v1.0
File Type
CSV (Comma Separated Values)
Version Overview
This is the initial release of the AH-Translit_Bench dataset, offering a foundational benchmark for Arabic to Hindi transliteration research.
Version Description
This inaugural version of the AH-Translit_Bench dataset, version 1.0, provides a comprehensive collection of 2000 Arabic-Hindi transliteration pairs. The dataset is segmented into three domain-specific files to cater to varied research needs: al-quran_test_bench_mark_500.csv containing 500 entries from the Al-Quran domain, biblo_test_bench_mark_1000.csv with 1000 entries from the bibliographical domain, and mse_test_bench_mark_500.csv offering 500 entries from the Modern Standard Arabic (MSA) domain. Each entry consists of an Arabic string and its corresponding, carefully transliterated Hindi equivalent, making it suitable for training and evaluating transliteration models.
Dataset Structure
The dataset is provided as a zip file (though usually downloaded directly via pip) containing three separate .csv files, each corresponding to a different domain. The structure within each CSV file is simple: the first column contains the Arabic text, and the second column contains its transliterated Hindi equivalent.
Author's
Vilal Ali 
Mohd Hozaifa Khan 
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ah_translit_bench-1.0.0.tar.gz.
File metadata
- Download URL: ah_translit_bench-1.0.0.tar.gz
- Upload date:
- Size: 117.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
aee0b06b7a3b56d7d7424026126af5f356a972bea2e8f28513a3cf4fb27cf07a
|
|
| MD5 |
2fed051d2e359dbc3170a626ed6c23af
|
|
| BLAKE2b-256 |
355f47368cdfbec7642d444417c95a98bcfdf60c6f526246a47821fcc0184bc0
|
File details
Details for the file ah_translit_bench-1.0.0-py3-none-any.whl.
File metadata
- Download URL: ah_translit_bench-1.0.0-py3-none-any.whl
- Upload date:
- Size: 121.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
33033faa70c4a6e52f4547f4fbaa8888269d9ca538738c174b0c3343a45ae0d3
|
|
| MD5 |
3ae346b5ede459d00e5f6f3304e1831d
|
|
| BLAKE2b-256 |
a01c41029e6c6f63ca1e28b3be2a087c69934daa7f37c31720a985db989a840d
|