CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Congo-digital-service /audios-lingala-annotatees Annotated Lingala Dataset – Full Version Description This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models. It includes: the original audio files (viewable directly in the Hugging Face viewer) text transcriptions Mel spectrograms tokenized labels Overall statistics Metric Value Total volume 5 h 0 min 18 s Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.audioautomatic-speech-recognition10K<n<100K0 likes322 downloads18d agoHugging Face02KasuleTrevor /Lingala_100hrs Lingala 100hrs 110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated from three publicly available CC-BY-4.0 corpora for ASR research. Composition Counts from a full-pass audit on 2026-07-09: Source Upstream location Rows Splits AfriVoice (Lingala) https://huggingface.co/datasets/DigitalUmuganda/AfriVoice 17,544 train (16,144), validation (915), test (485) LRSC (Lingala Read Speech Corpus)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.audioautomatic-speech-recognition10K<n<100K0 likes310 downloads3mo agoHugging Face03Congo-digital-service /audios-lingala-annotatees-v2 Annotated Lingala Audio — canonical corpus Annotated Lingala speech for open automatic speech recognition research and for fine-tuning speech models. This release is a full reconstruction of the corpus from its source recordings and annotations. It supersedes Congo-digital-service/audios-lingala-annotatees, which is deprecated — see Relationship to the previous release below. What this dataset contains Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.audioautomatic-speech-recognition10K<n<100K0 likes169 downloads14d agoHugging Face04BantuLanguagesInitiative /lingala_real_eval_benchmark_croped Lingala Real Eval Benchmark Cropped Small cropped real-world Lingala audio benchmark for testing BLI ASR 0. The dataset contains short audio clips cropped from longer real-world files, covering different domains such as news, catechesis, comedy, cartoon and interview speech. This dataset is intended for quick qualitative ASR testing and human review. It is not a training dataset. audioautomatic-speech-recognitionn<1K1 likes26 downloads4mo agoHugging Face05Speech-data /Lingala-Speech-Dataset Lingala Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Lingala (ln) 🏷️ Tags Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Lingala 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes14 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.