datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']
وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.mualem-recitations-annotatedmuaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.speech_commands_enriched_and_annotated
Dataset Summary
📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways:
Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.Azure-TTS-annotatedsna-dataset-annotated
manassehzw/sna-dataset-annotated
An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline.
This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments.
Why this annotated release exists
The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.Emilia-Annotated-WIPStill a WIP, full dataset is still being annotated
resd_annotated
RESD (Annotated)
RESD with a transcript for every clip.
How it was recorded
RESD was recorded in a studio by 20 voice actors. There was no script: the actors were not handed lines to read. Instead each actor in a pair was privately given an emotion to play, and the dialogue was improvised from there. So the words are spontaneous while the emotion is deliberate — which is the point, and also the limit. The label describes what the actor was told to convey, not… See the full description on the dataset page: https://huggingface.co/datasets/Aniemore/resd_annotated.laion-tts-annotated-v1
LAION TTS Annotated v1
107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens
and the annotations, all joined by one key.
283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion
intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst
detections with timings, and a natural-language caption describing the voice and the
delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.GLOBE-annotatedmultilingual-TEDX-fr-30s-pseudo-annotatedaudio_data_russian_annotated
Dataset Audio Russian Annotated
This is a dataset with Russian annotated audio data, split into train for tasks like text-to-speech, speech recognition, and speaker identification.
Features
text: Audio transcription (string).
speaker_name: Speaker identifier (string).
audio: Audio file.
utterance_pitch_mean: The average pitch of the speech utterance (float64).
utterance_pitch_std: The standard deviation of pitch, representing variability in intonation (float64)
snr:… See the full description on the dataset page: https://huggingface.co/datasets/kijjjj/audio_data_russian_annotated.mualem-recitations-annotatedlaion-tts-annotated-v1-research
Annotated Expressive Speech: Research-Access Sources
29,739,936 annotated utterance records and 103,912 published subset-view hours across three gated sources, with audio, codec tokens and the complete annotation stack. These are not 103,912 unique, independent hours: evasnippets largely reuses podcast source material.
The audio comes largely from publicly accessible podcasts and short snippets, but public accessibility is not an unrestricted audio licence. This repository… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1-research.fleurs_pseudo_annotatedFleurs dataset pseudo annotated using whisper-large-v3-turbo-partial-2000
GLOBE-Annotatednusantara-audiobook-annotatedaudio_data_russian_annotated
Dataset Audio Russian Annotated
This is a dataset with Russian annotated audio data, split into train for tasks like text-to-speech, speech recognition, and speaker identification.
Features
text: Audio transcription (string).
speaker_name: Speaker identifier (string).
audio: Audio file.
utterance_pitch_mean: The average pitch of the speech utterance (float64).
utterance_pitch_std: The standard deviation of pitch, representing variability in intonation (float64)… See the full description on the dataset page: https://huggingface.co/datasets/WatsonNT/audio_data_russian_annotated.GCP-TTS-annotatedannotated_catalan_common_voice_v17_cleaned_enhanced
Processed Annotated Catalan Common Voice v17 (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of:
projecte-aina/annotated_catalan_common_voice_v17.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/annotated_catalan_common_voice_v17_cleaned_enhanced.nusantara-audiobook-annotatedEmilia_annotatedkannada-tts-annotatedOpenDialog_English_Annotatedpony-singing-annotatedmulti_round_speech_180k_annotatedresd_annotated
Dataset Card for "resd_annotated"
More Information needed
GLOBE-annotatedlibritts-r-annotated
