datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']
وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.mualem-recitations-annotatedmuaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.speech_commands_enriched_and_annotated
Dataset Summary
📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways:
Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.Azure-TTS-annotatedsna-dataset-annotated
manassehzw/sna-dataset-annotated
An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline.
This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments.
Why this annotated release exists
The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.voice-acting-data-annotated
Voice Acting Data - Annotated
Post-processed version of laion/voice-acting-data.
Processing Pipeline
RE-USE Speech Enhancement (nvidia/RE-USE) - Applied to non-singing samples for noise reduction
LavaSR Super Resolution (YatharthS/LavaSR) - Audio bandwidth extension to 48kHz
Whisper Turbo ASR - Full transcript with word-level timestamps
Scene Split - Audio split at CUT TO: transition into two parts (Part 1 + Part 2)
VoiceCLAP Large Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-data-annotated.sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.Emilia-Annotated-WIPStill a WIP, full dataset is still being annotated
resd_annotated
RESD (Annotated)
RESD with a transcript for every clip.
How it was recorded
RESD was recorded in a studio by 20 voice actors. There was no script: the actors were not handed lines to read. Instead each actor in a pair was privately given an emotion to play, and the dialogue was improvised from there. So the words are spontaneous while the emotion is deliberate — which is the point, and also the limit. The label describes what the actor was told to convey, not… See the full description on the dataset page: https://huggingface.co/datasets/Aniemore/resd_annotated.laion-tts-annotated-v1
LAION TTS Annotated v1
107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens
and the annotations, all joined by one key.
283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion
intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst
detections with timings, and a natural-language caption describing the voice and the
delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.GLOBE-annotatedmultilingual-TEDX-fr-30s-pseudo-annotatedaudio_data_russian_annotated
Dataset Audio Russian Annotated
This is a dataset with Russian annotated audio data, split into train for tasks like text-to-speech, speech recognition, and speaker identification.
Features
text: Audio transcription (string).
speaker_name: Speaker identifier (string).
audio: Audio file.
utterance_pitch_mean: The average pitch of the speech utterance (float64).
utterance_pitch_std: The standard deviation of pitch, representing variability in intonation (float64)
snr:… See the full description on the dataset page: https://huggingface.co/datasets/kijjjj/audio_data_russian_annotated.mualem-recitations-annotatedlaion-tts-annotated-v1-research
LAION TTS Annotated v1 — research subsets
29,739,936 annotated speech utterances across three subsets — with the audio, the codec tokens
and the complete annotation stack.
The audio in this repository comes from podcasts that are openly available on the internet and
consists of short snippets only. We cannot redistribute the audio itself, so it is made available
here for non-commercial research use by collaboration partners within our TTS research.
The other six subsets of this… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1-research.fleurs_pseudo_annotatedFleurs dataset pseudo annotated using whisper-large-v3-turbo-partial-2000
nusantara-audiobook-annotatedGLOBE-Annotatedaudio_data_russian_annotated
Dataset Audio Russian Annotated
This is a dataset with Russian annotated audio data, split into train for tasks like text-to-speech, speech recognition, and speaker identification.
Features
text: Audio transcription (string).
speaker_name: Speaker identifier (string).
audio: Audio file.
utterance_pitch_mean: The average pitch of the speech utterance (float64).
utterance_pitch_std: The standard deviation of pitch, representing variability in intonation (float64)… See the full description on the dataset page: https://huggingface.co/datasets/WatsonNT/audio_data_russian_annotated.GCP-TTS-annotatedannotated_catalan_common_voice_v17_cleaned_enhanced
Processed Annotated Catalan Common Voice v17 (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of:
projecte-aina/annotated_catalan_common_voice_v17.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/annotated_catalan_common_voice_v17_cleaned_enhanced.nusantara-audiobook-annotatedkannada-tts-annotatedEmilia_annotatedpony-singing-annotatedOpenDialog_English_AnnotatedAnnotated_Food_Vlog_Dataset_GroupL
Dataest Description
This project has constructed a multimodal corpus of language strategies for food exploration videos on Chinese social media. The dataset is centered around the videos of the well-known blogger "Diao Yueshe Shi Yu Ji", containing approximately 1,000 entries with a total of 90 minutes of transcribed video content. The dataset is stored in CSV format and meticulously records the original dialogue, synthetic text generated by large language models (LLMs), rhetorical… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/Annotated_Food_Vlog_Dataset_GroupL.multi_round_speech_180k_annotatedresd_annotated
Dataset Card for "resd_annotated"
More Information needed
