datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sna-dataset-annotated
manassehzw/sna-dataset-annotated
An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline.
This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments.
Why this annotated release exists
The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.audios-lingala-annotatees
Annotated Lingala Dataset – Full Version
Description
This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models.
It includes:
the original audio files (viewable directly in the Hugging Face viewer)
text transcriptions
Mel spectrograms
tokenized labels
Overall statistics
Metric
Value
Total volume
5 h 0 min 18 s
Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.audios-lingala-annotatees-v2
Annotated Lingala Audio — canonical corpus
Annotated Lingala speech for open automatic speech recognition research and for
fine-tuning speech models.
This release is a full reconstruction of the corpus from its source
recordings and annotations. It supersedes
Congo-digital-service/audios-lingala-annotatees,
which is deprecated — see Relationship to the previous release below.
What this dataset contains
Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.annotated_catalan_common_voice_v17_cleaned_enhanced
Processed Annotated Catalan Common Voice v17 (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of:
projecte-aina/annotated_catalan_common_voice_v17.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/annotated_catalan_common_voice_v17_cleaned_enhanced.
