datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
everyayah﷽
Dataset Card for Tarteel AI's EveryAyah Dataset
Dataset Summary
This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in Arabic.
Dataset Structure
Data Instances
A typical data point comprises the audio file audio, and its transcription called text.
The duration… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/everyayah.tartanaviation-atc-adsb-utterances
TartanAviation ATC + ADS-B (Utterances)
Speech utterances split from twangodev/tartanaviation-atc-adsb
by voice-activity detection (pyannote/segmentation-3.0).
Each row is one speech segment (16 kHz mono) with the ADS-B from its parent clip.
531,050 utterances · ~398 h speech · 16 kHz mono · 67% carry ADS-B. From 40,899 of 41,823 clips
(silent clips have no utterances). Built with squawk.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-adsb-utterances.tartanaviation-atc-adsb
TartanAviation ATC + ADS-B
Paired ATC audio and ADS-B for Pittsburgh KAGC and KBTP, aligned from CMU
TartanAviation. Each row is one ADS-B-triggered
audio capture (16 kHz mono) plus the aircraft tracks present during it.
41,823 clips · 16 kHz mono · 67% carry ADS-B. Built with squawk.
Usage
from datasets import load_dataset
ds = load_dataset("twangodev/tartanaviation-atc-adsb", split="train", streaming=True)
ex = next(iter(ds))
ex["audio"] # {'array': ...… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-adsb.tarteel-ai-everyayah-Quran﷽
Dataset Card for Tarteel AI's EveryAyah Dataset
Dataset Summary
This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters.
How to download
!pip install -q datasets
from datasets import load_dataset
dataset =load_dataset("Salama1429/tarteel-ai-everyayah-Quran", verification_mode="no_checks")
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in… See the full description on the dataset page: https://huggingface.co/datasets/Salama1429/tarteel-ai-everyayah-Quran.AnimeVox
AnimeVox: Character TTS Corpus
🗣️ Dataset Overview
AnimeVox is an English Text-to-Speech (TTS) dataset featuring 11,020 audio clips from 19 distinct anime characters across popular series. Each clip includes a high-quality transcription, character name, and anime title, making it ideal for voice cloning, custom TTS model fine-tuning, and character voice synthesis research.
The dataset was created and processed using TTSizer, an open-source tool that automates creating… See the full description on the dataset page: https://huggingface.co/datasets/taresh18/AnimeVox.tartanaviation-atc-labels
TartanAviation ATC ASR Labels
Machine transcripts and confidence scores for
twangodev/tartanaviation-atc-adsb-utterances.
A 1:1 labels-only add-on (no audio): one row per source utterance, same shards and row order, keyed
by utterance_id.
531,050 labels · 184 shards · 100% coverage · ensemble ASR + weighted ROVER + ADS-B
callsign snap · ~326 human-reviewed. Built with readback.
Usage
Join 1:1 onto the source. Rows are aligned and in the same order:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/tartanaviation-atc-labels.persian-tarjoman-koochik-16k
Persian Tarjoman — Koochik-labeled (16k mono FLAC, NeMo-tarred, streamable)
74,808 clips / 182.9 hours of clean long-form Persian narration (Tarjoman magazine articles).
Audio: from farsi-asr/PerSets-tarjoman-chunked, resampled to 16kHz mono.
Labels: transcribed with Reza2kn/Shenava-Koochik-v1.0 (CTC head) — the source's own Speechmatics transcripts were empty.
Format: NeMo-tarred / streamable — audio/shard_XXXXX.tar (19 shards) + manifests/train.jsonl (audio_filepath =… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-tarjoman-koochik-16k.Zambezi_ECHO_v1
Zambezi ECHO v1: Shona-English Code-Switched Maternal Health Queries
Dataset Description
Zambezi ECHO v1 SESB (Shona-English Speech Benchmark) is a dataset of short,
simulated patient voice queries in Shona (Zimbabwe), code-switched with
English, covering common maternal and child health concerns — pregnancy
symptoms, danger signs, child illness, and general health questions asked
the way patients actually phrase them in the field, mixing Shona with
English… See the full description on the dataset page: https://huggingface.co/datasets/tarirozw/Zambezi_ECHO_v1.uk-pods
uk-pods - speech datasets of Ukrainian podcasts.
Preparation
Clone the dataset repository and extract the content of clips.tar.gz archive.
git clone https://huggingface.co/datasets/taras-sereda/uk-pods
cd uk-pods && tar -zxvf clips.tar.gz
To use these manifests for training/inference with NeMo [1] modify audio_filepath to absolute locations of audio files extracted in previous step.
# data_root=<clonned_repo_dir> # /home/ubuntu/uk-pods
data_root=$(realpath .)
sed -i… See the full description on the dataset page: https://huggingface.co/datasets/taras-sereda/uk-pods.taras_shau_chenka_vershy_paemy_all
Вершы і паэмы
Аўтар / Author: Тарас ШаўчэнкаМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
253
Працягласць
51 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/taras_shau_chenka_vershy_paemy_all.target-words-geminitts
Target-Word (TW) Evaluation Set
Synthetic speech clips for 101 rare drug terms, intended for evaluation only — measuring how
well an ASR system recognises rare / out-of-vocabulary medical vocabulary (target-word WER / CER /
recall). Each clip reads a real DailyMed sentence containing one target drug name, synthesised with
Google Gemini TTS across multiple voices. This is the frozen evaluation set from the master's thesis
"Audio-free lexical adaptation of Whisper's decoder"… See the full description on the dataset page: https://huggingface.co/datasets/aharalambieva/target-words-geminitts.sna-waxal-unlabeled-tar
WAXAL Shona unlabeled operational TAR dataset
WebDataset packaging of the Shona unlabeled split from
google/WaxalNLP.
config: sna_asr
split: unlabeled
pinned upstream revision: e0a62aaebc61bd5bb8cac17a08d1b42c65551dd2
samples: 85,384
audio: source-encoded bytes preserved without transcoding
The pinned upstream Parquet files remain the recoverable source. This repo is
an operational derivative optimized for sequential streaming. Each example is
a matching audio and .json pair.… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-unlabeled-tar.taras_shau_chenka_vershy_paemy_output_original
Вершы і паэмы — арыгінальнае аўдыё
Аўтар / Author: Тарас ШаўчэнкаМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
taras_shau_chenka_vershy_paemy_output
Доўгасць аўдыё
0h53m
Радкоў у датасеце
208
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/taras_shau_chenka_vershy_paemy_output_original.
