CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Congo-digital-service /audios-lingala-annotatees Annotated Lingala Dataset – Full Version Description This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models. It includes: the original audio files (viewable directly in the Hugging Face viewer) text transcriptions Mel spectrograms tokenized labels Overall statistics Metric Value Total volume 5 h 0 min 18 s Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.audioautomatic-speech-recognition10K<n<100K0 likes322 downloads14d agoHugging Face02KasuleTrevor /Lingala_100hrs Lingala 100hrs 110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated from three publicly available CC-BY-4.0 corpora for ASR research. Composition Counts from a full-pass audit on 2026-07-09: Source Upstream location Rows Splits AfriVoice (Lingala) https://huggingface.co/datasets/DigitalUmuganda/AfriVoice 17,544 train (16,144), validation (915), test (485) LRSC (Lingala Read Speech Corpus)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.audioautomatic-speech-recognition10K<n<100K0 likes267 downloads3mo agoHugging Face03Congo-digital-service /audios-lingala-annotatees-v2 Annotated Lingala Audio — canonical corpus Annotated Lingala speech for open automatic speech recognition research and for fine-tuning speech models. This release is a full reconstruction of the corpus from its source recordings and annotations. It supersedes Congo-digital-service/audios-lingala-annotatees, which is deprecated — see Relationship to the previous release below. What this dataset contains Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.audioautomatic-speech-recognition10K<n<100K0 likes148 downloads10d agoHugging Face04Regineforte /tts_lingala_maleaudio1K<n<10K0 likes121 downloads11mo agoHugging Face05shunyalabs /lingala-speech-datasetaudio1K<n<10K0 likes87 downloads1y agoHugging Face06Congo-digital-service /qwen-vl-lingala-dataset-augmented Augmentation This dataset derives from dataset-qwen-vl-lingala-qlora-vf (417 train / 50 test) through an augmentation step applied to the training image-text pairs, bringing the training volume to 884 examples. Augmentation method: the 467 additional training examples compared to the source (417 → 884) are obtained mainly through controlled degradation of the input image — noise, brightness/contrast variation, light blur — rather than through synthetic content generation or text… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/qwen-vl-lingala-dataset-augmented.imagen<1K0 likes87 downloads14d agoHugging Face07Congo-digital-service /dataset-qwen-vl-lingala-qlora-vf Qwen-VL Lingala OCR Dataset Description Image/text pairs for training a Qwen2-VL model to perform OCR on Lingala text, including the two special characters absent from the standard Latin alphabet: ɔ (U+0254) and ɛ (U+025B). train: original + augmented images (noise, brightness/contrast, light blur), with targeted oversampling of lines containing ɔ/ɛ. test: original, non-augmented images only, held out before any oversampling to avoid data leakage.… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/dataset-qwen-vl-lingala-qlora-vf.imagen<1K0 likes85 downloads14d agoHugging Face08Congo-digital-service /Lingala-Alpaca-Dataset-Augmented Lingala Alpaca LLM Dataset (Augmented) This dataset is an augmented version of the original dataset available at Congo-digital-service/Dataset-Lingala-Alpaca-LLM, with a larger and more diverse training set than the source, while preserving the original dataset's purpose and structure. Dataset Structure Same fields as the source dataset: instruction, input, output, source_project, split into train and test. See the source dataset card for the category/domain… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/Lingala-Alpaca-Dataset-Augmented.text1K<n<10K0 likes83 downloads14d agoHugging Face09Congo-digital-service /Dataset-Lingala-Alpaca-LLM Lingala Alpaca LLM Dataset This dataset contains human-generated instruction-following data in Lingala, inspired by the Alpaca dataset format, covering five task categories (formal, pedagogical, summary, urban, translation) across five public-interest domains, with a balanced representation of stylistic categories and instruction types. General Statistics Total volume: 6,664 lines in total. Number of unique source texts: 1,772. Structure: The dataset is divided… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/Dataset-Lingala-Alpaca-LLM.text1K<n<10K0 likes82 downloads14d agoHugging Face10Svngoku /wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3texttext-generation10K<n<100K2 likes73 downloads2y agoHugging Face11KasuleTrevor /lingala_20hraudio1K<n<10K1 likes70 downloads2y agoHugging Face12saillab /alpaca-lingala-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-lingala-cleaned.text10K<n<100K0 likes50 downloads2y agoHugging Face13michsethowusu /lingala-sentiments-corpus Lingala Sentiment Corpus Dataset Description This dataset contains sentiment-labeled text data in Lingala for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources. Dataset Statistics Total samples: 427,979 Positive sentiment: 251923 (58.9%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-sentiments-corpus.texttext-classification100K<n<1M0 likes40 downloads1y agoHugging Face14michsethowusu /Code-170k-lingala Dataset Description Code-170k-lingala is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Lingala, making coding education accessible to Lingala speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Lingala language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-lingala.texttext-generation100K<n<1M1 likes38 downloads11mo agoHugging Face15michsethowusu /lingala-emotions-corpus Lingala Emotion Analysis Corpus Dataset Description This dataset contains emotion-labeled text data in Lingala for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources. Dataset Statistics Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-emotions-corpus.texttext-classification100K<n<1M0 likes36 downloads1y agoHugging Face16michsethowusu /english-lingala_sentence-pairs_mt560 English-Lingala Parallel Dataset This dataset contains parallel sentences in English and Lingala (Democratic Republic of the Congo). Dataset Information Language Pair: English ↔ Lingala Language Code: lin Country: Democratic Republic of the Congo Original Source: OPUS MT560 Dataset Dataset Structure The dataset contains parallel sentences that can be used for: Machine translation training Cross-lingual NLP tasks Language model fine-tuning Citation… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-lingala_sentence-pairs_mt560.text100K<n<1M0 likes35 downloads1y agoHugging Face17michsethowusu /french-lingala_sentence-pairs French-Lingala_Sentence-Pairs Dataset This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks. It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: French-Lingala_Sentence-Pairs File Size: 78580422 bytes Languages: French, French Dataset Description The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/french-lingala_sentence-pairs.text100K<n<1M1 likes28 downloads1y agoHugging Face18michsethowusu /kamba-lingala_sentence-pairs Kamba-Lingala_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Kamba-Lingala_Sentence-Pairs Number of Rows: 50317 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-lingala_sentence-pairs.text10K<n<100K0 likes27 downloads1y agoHugging Face19KasuleTrevor /lingala_10hraudio1K<n<10K0 likes24 downloads2y agoHugging Face20KasuleTrevor /Lingala_bmd_testaudio1K<n<10K1 likes24 downloads2y agoHugging Face21michsethowusu /lingala-swahili_sentence-pairs Lingala-Swahili_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Lingala-Swahili_Sentence-Pairs Number of Rows: 391907 Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-swahili_sentence-pairs.text100K<n<1M2 likes24 downloads1y agoHugging Face22KasuleTrevor /lingala_5hraudio1K<n<10K0 likes23 downloads2y agoHugging Face23michsethowusu /bemba-lingala_sentence-pairs Bemba-Lingala_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Bemba-Lingala_Sentence-Pairs Number of Rows: 140378 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-lingala_sentence-pairs.text100K<n<1M0 likes23 downloads1y agoHugging Face24BantuLanguagesInitiative /lingala_real_eval_benchmark_croped Lingala Real Eval Benchmark Cropped Small cropped real-world Lingala audio benchmark for testing BLI ASR 0. The dataset contains short audio clips cropped from longer real-world files, covering different domains such as news, catechesis, comedy, cartoon and interview speech. This dataset is intended for quick qualitative ASR testing and human review. It is not a training dataset. audioautomatic-speech-recognitionn<1K1 likes23 downloads4mo agoHugging Face25michsethowusu /english-lingala_sentence-pairs English-Lingala_Sentence-Pairs Dataset This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks. It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: English-Lingala_Sentence-Pairs File Size: 307050723 bytes Languages: English, English Dataset Description The dataset contains sentence pairs… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-lingala_sentence-pairs.text1M<n<10M1 likes22 downloads1y agoHugging Face26Svngoku /lingala-asr-dataset Lingala Automatic Speech Recognition ( Text To Speech & Speech To Text Dataset ) This dataset card aims to be a base template for new datasets. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] audion<1K2 likes21 downloads2y agoHugging Face27Svngoku /kongo-lingala-swahili-french-english-pairs-bibletabular1M<n<10M1 likes21 downloads1y agoHugging Face28michsethowusu /dyula-lingala_sentence-pairs Dyula-Lingala_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Dyula-Lingala_Sentence-Pairs Number of Rows: 57764 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-lingala_sentence-pairs.text10K<n<100K0 likes20 downloads1y agoHugging Face29saillab /alpaca_lingala_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_lingala_taco.text10K<n<100K3 likes19 downloads2y agoHugging Face30tokossapp /Code-170k-lingala Dataset Description Code-170k-lingala is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Lingala, making coding education accessible to Lingala speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Lingala language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/tokossapp/Code-170k-lingala.texttext-generation100K<n<1M0 likes19 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.