CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01obadx /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train'] وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.audio100K<n<1M7 likes2.7k downloads1y agoHugging Face02obadx /mualem-recitations-annotatedaudio100K<n<1M4 likes2.3k downloads1y agoHugging Face03nour-world /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.audio100K<n<1M0 likes670 downloads9d agoHugging Face04soerenray /speech_commands_enriched_and_annotated Dataset Summary 📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development. 🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways: Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.audio10K<n<100K2 likes626 downloads3y agoHugging Face05mesolitica /Azure-TTS-annotatedaudio100K<n<1M0 likes499 downloads2y agoHugging Face06manassehzw /sna-dataset-annotated manassehzw/sna-dataset-annotated An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline. This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments. Why this annotated release exists The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.audioautomatic-speech-recognition10K<n<100K1 likes437 downloads2mo agoHugging Face07laion /voice-acting-data-annotated Voice Acting Data - Annotated Post-processed version of laion/voice-acting-data. Processing Pipeline RE-USE Speech Enhancement (nvidia/RE-USE) - Applied to non-singing samples for noise reduction LavaSR Super Resolution (YatharthS/LavaSR) - Audio bandwidth extension to 48kHz Whisper Turbo ASR - Full transcript with word-level timestamps Scene Split - Audio split at CUT TO: transition into two parts (Part 1 + Part 2) VoiceCLAP Large Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-data-annotated.audioaudio-classification0 likes382 downloads14d agoHugging Face08manassehzw /sna-waxal-annotated-unlabeled Shona WAXAL annotated-unlabeled checkpoint This is a self-contained operational checkpoint for pseudo-labeling Shona ASR data. It contains 90,253 conservatively segmented FLAC clips (441.585 hours), but intentionally contains no transcripts. Fields transcription is intentionally empty. speaker_id is an approximate source-blind EOM cluster or unknown; speaker_clip_count is zero for unknown assignments. gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.audioautomatic-speech-recognition10K<n<100K0 likes374 downloads2mo agoHugging Face09laion /Emilia-Annotated-WIPStill a WIP, full dataset is still being annotated audio1M<n<10M3 likes367 downloads1y agoHugging Face10Aniemore /resd_annotated RESD (Annotated) RESD with a transcript for every clip. How it was recorded RESD was recorded in a studio by 20 voice actors. There was no script: the actors were not handed lines to read. Instead each actor in a pair was privately given an emotion to play, and the dialogue was improvised from there. So the words are spontaneous while the emotion is deliberate — which is the point, and also the limit. The label describes what the actor was told to convey, not… See the full description on the dataset page: https://huggingface.co/datasets/Aniemore/resd_annotated.audioaudio-classification1K<n<10K11 likes265 downloads2mo agoHugging Face11laion /laion-tts-annotated-v1 LAION TTS Annotated v1 107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens and the annotations, all joined by one key. 283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst detections with timings, and a natural-language caption describing the voice and the delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.tabulartext-to-speech100M<n<1B0 likes250 downloads13d agoHugging Face12ai4b-hf /GLOBE-annotatedaudio100K<n<1M10 likes226 downloads2y agoHugging Face13BrunoHays /multilingual-TEDX-fr-30s-pseudo-annotatedaudio10K<n<100K0 likes192 downloads1y agoHugging Face14kijjjj /audio_data_russian_annotated Dataset Audio Russian Annotated This is a dataset with Russian annotated audio data, split into train for tasks like text-to-speech, speech recognition, and speaker identification. Features text: Audio transcription (string). speaker_name: Speaker identifier (string). audio: Audio file. utterance_pitch_mean: The average pitch of the speech utterance (float64). utterance_pitch_std: The standard deviation of pitch, representing variability in intonation (float64) snr:… See the full description on the dataset page: https://huggingface.co/datasets/kijjjj/audio_data_russian_annotated.audiotext-to-speech100K<n<1M2 likes175 downloads1y agoHugging Face15nour-world /mualem-recitations-annotatedaudio100K<n<1M0 likes119 downloads9d agoHugging Face16laion /laion-tts-annotated-v1-researchgated LAION TTS Annotated v1 — research subsets 29,739,936 annotated speech utterances across three subsets — with the audio, the codec tokens and the complete annotation stack. The audio in this repository comes from podcasts that are openly available on the internet and consists of short snippets only. We cannot redistribute the audio itself, so it is made available here for non-commercial research use by collaboration partners within our TTS research. The other six subsets of this… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1-research.tabulartext-to-speech10M<n<100M0 likes111 downloads13d agoHugging Face17BrunoHays /fleurs_pseudo_annotatedFleurs dataset pseudo annotated using whisper-large-v3-turbo-partial-2000 audio1K<n<10K0 likes71 downloads1y agoHugging Face18mesolitica /nusantara-audiobook-annotatedaudio10K<n<100K1 likes66 downloads2y agoHugging Face19zuhri025 /GLOBE-Annotatedaudio100K<n<1M0 likes63 downloads9mo agoHugging Face20WatsonNT /audio_data_russian_annotated Dataset Audio Russian Annotated This is a dataset with Russian annotated audio data, split into train for tasks like text-to-speech, speech recognition, and speaker identification. Features text: Audio transcription (string). speaker_name: Speaker identifier (string). audio: Audio file. utterance_pitch_mean: The average pitch of the speech utterance (float64). utterance_pitch_std: The standard deviation of pitch, representing variability in intonation (float64)… See the full description on the dataset page: https://huggingface.co/datasets/WatsonNT/audio_data_russian_annotated.audiotext-to-speech100K<n<1M0 likes53 downloads28d agoHugging Face21mesolitica /GCP-TTS-annotatedaudio100K<n<1M1 likes52 downloads2y agoHugging Face22ebellob /annotated_catalan_common_voice_v17_cleaned_enhanced Processed Annotated Catalan Common Voice v17 (CleanUNet + FlashSR) Dataset Summary This dataset is a processed and enhanced version of: projecte-aina/annotated_catalan_common_voice_v17. Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove. However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/annotated_catalan_common_voice_v17_cleaned_enhanced.audiotext-to-speech100K<n<1M2 likes50 downloads5mo agoHugging Face23skilledu /nusantara-audiobook-annotatedaudio10K<n<100K0 likes44 downloads4mo agoHugging Face24lokesh122151 /kannada-tts-annotatedaudio1K<n<10K0 likes36 downloads5mo agoHugging Face25AdrienB134 /Emilia_annotatedaudio100K<n<1M0 likes35 downloads2y agoHugging Face26AmongTheCouch23 /pony-singing-annotatedaudio1K<n<10K0 likes16 downloads3mo agoHugging Face27zuhri025 /OpenDialog_English_Annotatedaudio1K<n<10K0 likes14 downloads7mo agoHugging Face28eduhk-compling /Annotated_Food_Vlog_Dataset_GroupL Dataest Description This project has constructed a multimodal corpus of language strategies for food exploration videos on Chinese social media. The dataset is centered around the videos of the well-known blogger "Diao Yueshe Shi Yu Ji", containing approximately 1,000 entries with a total of 90 minutes of transcribed video content. The dataset is stored in CSV format and meticulously records the original dialogue, synthetic text generated by large language models (LLMs), rhetorical… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/Annotated_Food_Vlog_Dataset_GroupL.audio1K<n<10K0 likes13 downloads5mo agoHugging Face29zuhri025 /multi_round_speech_180k_annotatedaudio1K<n<10K0 likes10 downloads7mo agoHugging Face30amil91 /resd_annotated Dataset Card for "resd_annotated" More Information needed audioaudio-classification1K<n<10K0 likes8 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.