datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nasle-mana-clean-chunked-30s
Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks
Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk.
Configuration
Rows
Columns
Meaning
labeled (train/)
9,886
audio, label
Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.nasser-al-qatami-128kbps
Nasser Al-Qatami
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
nasser-al-qatami-128kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
128 kbps
Ayah files
6350 (1126 MiB)
Verified against the upstream MD5 list
6350
Ayahs absent upstream
0
Upstream folder
Nasser_Alqatami_128kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/nasser-al-qatami-128kbps.nasle-mana-clean
Nasl-e-Mana Speech Corpus
Clean, playable Persian speech audio collected from the public Nasl-e-Mana magazine WordPress site. The export contains two intentionally different collections:
Split
Rows
Columns
Meaning
labeled configuration (train/)
809
audio, label
Audio with recovered article text. These clips were identified as the consistent female source-text voice and are kept together.
to_transcribe configuration (to_transcribe/)
626
audio
Playable audio for which… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean.jordanian_arabic_speech_nasirquran_dataset_afasy_clean
Quranic Dataset by Tanzil Project (Qari: Al Afasy)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification process:… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_afasy_clean.quran_dataset_hani_clean
Quranic Dataset by Tanzil Project (Qari: Hani)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification process:… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_hani_clean.health_domain
Healthcare
Collected with the Collector platform for Ghanaian language data.
Tasks: Translation, Speech recognition
Languages: en → tw
Rows: 1
Audio: 0.1 minutes
Voices: female: 1
Privacy
Rows identify their speaker by an opaque code only — female_01, male_02, or
speaker_01 where no voice was stated. No names, ages, regions, phone numbers
or contact details are included. Codes are numbered per dataset, so codes in
two datasets from this platform cannot be… See the full description on the dataset page: https://huggingface.co/datasets/nasare34/health_domain.quran_dataset_abdulbasit_clean
Quranic Dataset by Tanzil Project (Qari: Abdul Basit)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_abdulbasit_clean.quran_dataset_ajamy_clean
Quranic Dataset by Tanzil Project (Qari: Al Ajamy)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification process:… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_ajamy_clean.quran_dataset_tablawi_clean
Quranic Dataset by Tanzil Project (Qari: At Tablawi)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_tablawi_clean.akan-speech-collect
Akan Speech Collection
Read speech in Akan with transcripts, for training speech recognition.
ASR_En_Ar_CodeSwitchingquran_dataset_basfar_clean
Quranic Dataset by Tanzil Project (Qari: Basfar)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification process:… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_basfar_clean.Nasalization_and_Velum_Control_Preview
Harmonic Frontier Audio – Nasalization and Velum Control (Preview, v0.9)
A high-fidelity human vocal dataset designed for AI training, speech research, and expressive voice modeling.
Nasalization and Velum Control (Preview), created by Harmonic Frontier Audio, provides a compact reference set demonstrating the quality, formatting, and metadata conventions used in the Harmonic Frontier Audio Human Vocality Primitives series.
🔎 Summary
This dataset provides… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Nasalization_and_Velum_Control_Preview.saudi_arabic_accentBanglaS2SNascimentoquran_dataset_akhdar_clean
Quranic Dataset by Tanzil Project (Qari: Al Akhdar)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification process:… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_akhdar_clean.dataset_cpfleurs-eg-cleanbackground-noise-detection-dataset
Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours)
Dataset summary
50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class
Purpose and usage scenarios
Speech… See the full description on the dataset page: https://huggingface.co/datasets/NashAli/background-noise-detection-dataset.persona-voicestake_me_out_season3_ep11_full_audioMoamn-ZvNrDRWRJQkahmed-elsbay-4_mp3pimm_nasdaq_earningscall
Dataset Description (Categorized)
To support the development and evaluation of the Physics-Informed Acoustic Model (PIAM), we construct a multimodal dataset categorized as follows:
1. Sound Earnings Call Data
Number of Companies: 283 NASDAQ-listed corporations
Number of Recordings: 1,795 earnings call sessions
Total Audio Duration: Approximately 1,780 hours
Time Range: January 22, 2021 – June 29, 2025
Speaker Roles Labeled: CEO, CFO, CXO (and others where… See the full description on the dataset page: https://huggingface.co/datasets/soundai2016/pimm_nasdaq_earningscall.nasalomeAboKhaledRevision-aSDdHSW6-bUarabic_mozilla_dataset_clean
