datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hani-ar-rifai-192kbps
Hani Ar-Rifai
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
hani-ar-rifai-192kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
192 kbps
Ayah files
6236 (2053 MiB)
Verified against the upstream MD5 list
6234
Ayahs absent upstream
0
Upstream folder
Hani_Rifai_192kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3 is… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/hani-ar-rifai-192kbps.hani-ar-rifai-64kbps
Hani Ar-Rifai
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
hani-ar-rifai-64kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
64 kbps
Ayah files
6236 (702 MiB)
Verified against the upstream MD5 list
6234
Ayahs absent upstream
0
Upstream folder
Hani_Rifai_64kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3 is Al-Fatihah… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/hani-ar-rifai-64kbps.Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.HowFarAreYou_3DSpeakerTrain_fullquran_dataset_hani_clean
Quranic Dataset by Tanzil Project (Qari: Hani)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification process:… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_hani_clean.SynthaticPipelines
For mor info follow the below link at Github
(SyntheticData@Github)[]
26295 Row
~5.5 GB
~34H:22M
EnvironmentalSoundClassification_ESC50-HumanAndNonSpeechSounds_TTSAudioSkills
AudioSkills-XL Dataset
Project page | Paper | Code
Dataset Description
AudioSkills-XL is a large-scale audio question-answering (AQA) dataset designed to develop (large) audio-language models on expert-level reasoning and problem-solving tasks over short audio clips (≤30 seconds). It expands upon the original AudioSkills collection by adding approximately 4.5 million new QA pairs, resulting in a total of ~10 million diverse examples. The release includes the full dataset… See the full description on the dataset page: https://huggingface.co/datasets/hanishamid/AudioSkills.PronounciationEvaluationFluency_Speechocean762StressDetection_MIRSD_TTSsnips_slu_v1.0DialogueEmotionClassification_DailyTalkSpeakerVerification_LibriSpeech-TestClean_TTSDialogueActClassification_DailyTalkAccentClassification_AccentdbExtended_TTSDialogueEmotionClassification_DailyTalk_testparalinguistic_datasetSarcasmDetection_Mustard_TTSvivos-vi-asrSpoofDetection_ASVspoof2017_TTSEmotionRecognition_MultimodalEmotionlinesDataset_TTSIntentClassification_FluentSpeechCommands-Action_TTSadversarialMultiSpeakerDetection_LibriSpeech-TestClean_TTSDialogueActClassification_DailyTalk_TTSnon_paralinguistic_datasetHowFarAreYou_3DSpeakeryoutube_audio_samples2english_accent_samplesyoutube_audio_samples
