CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LanguageShades /BiasShadesgatedInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab! Dataset Card for BiasShades Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators. Dataset Details Version: 1.0 License: SHADES 1 Montreal Data License Dataset Description 728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.imagetext-classificationn<1K26 likes577 downloads3mo agoHugging Face02agentlans /HuggingFaceFW-finetranslations-100-languages-sample Finetranslations 100 Language Sample Dataset Subset of HuggingFaceFW/finetranslations with the top 100 languages by number of documents. Configurations all: 100 languages combined (100k rows), shuffled 100 individual language configs: 1000 rows each Columns Original columns + language (source language indicator which is the name of the config) Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finetranslations-100-languages-sample.tabulartranslation100K<n<1M0 likes121 downloads7mo agoHugging Face03Mehgoss /sa-languages-corpus South African Languages Text Corpus Plain-text corpus covering all 11 official South African languages, for language modeling. text-generation0 likes89 downloads9d agoHugging Face04itsmebatuhan /bluesky-10m-posts-15-languages Dataset Card: Bluesky 10M Multilingual 📊 Overview Total Posts: 10,099,990 Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi) Collection Period: August 9-12, 2026 Source: Bluesky Jetstream API (public firehose) Format: JSONL Size: ~3 GB 🌍 Language Distribution Language Code Posts % English en 6,843,995 67.8% Japanese ja 1,547,179 15.3% German de 373,626 3.7% Portuguese pt 331,093 3.3% Spanish es 325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.texttext-classification10M<n<100M0 likes80 downloads1mo agoHugging Face05math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes56 downloads3mo agoHugging Face06anrilombard /sa-languages South African Languages Dataset Dataset Overview Language Training Documents Training GPT2 Tokens Avg Tokens/Doc Max Tokens Test Documents Test GPT2 Tokens Test Avg Tokens/Doc Test Max Tokens isiZulu 116,693 192,622,799 1,650.68 335,530 687 1,080,961 1,573.45 15,691 Sesotho 83,329 144,337,938 1,732.15 98,542 841 1,393,086 1,656.4614,071 isiXhosa 99,567 141,484,241 1,421.00 113,710 788 1,161,296 1,473.73 17,220 isiNdebele 21,922 17,533,799 799.83 42,701… See the full description on the dataset page: https://huggingface.co/datasets/anrilombard/sa-languages.texttext-generation100K<n<1M1 likes35 downloads2y agoHugging Face07taresco /big_math_translated_african_languages Big Math Translated -- African Languages This is a set of 41k SynthLabsAI/Big-Math-RL-Verified questions translated into 9 African languages using Azure/GPT-4o. We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language. texttext-generation10K<n<100K0 likes25 downloads9mo agoHugging Face08anrilombard /sa-nguni-languages South African Nguni Languages Dataset Dataset Overview Language Training Documents Training GPT2 Tokens Avg Tokens/Doc Max Tokens Test Documents Test GPT2 Tokens Test Avg Tokens/Doc Test Max Tokens isiZulu 116,693 192,622,799 1,650.68 335,530 687 1,080,961 1,573.45 15,691 isiXhosa 99,567 141,484,241 1,421.00 113,710 788 1,161,296 1,473.7317,220 isiNdebele 21,922 17,533,799 799.83 42,701 222 170,111 766.27 6,615 siSwati 1,668 3,148,007 1,887.29 24,129 17… See the full description on the dataset page: https://huggingface.co/datasets/anrilombard/sa-nguni-languages.texttext-generation100K<n<1M0 likes24 downloads2y agoHugging Face09African-Languages-Lab /proxy-mt-benchmark-scoresgated Proxy-MT Benchmark Scores Multilingual benchmark results for 50 open-weight LLMs, evaluated with the lm-evaluation-harness via a vLLM backend. Covers reasoning, comprehension, and knowledge tasks with an emphasis on African and other lower-resource languages. Layout scores/<model>.csv # parsed per-language scores (tidy, ready to plot) raw/<model>/.../results_*.json # raw lm-eval-harness result files raw/<model>/raw_log.txt # full evaluation… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-benchmark-scores.text-generationn<1K0 likes12 downloads1mo agoHugging Face10MWirelabs /northeast-languages-test-setgated Northeast Languages Test Set A curated test set of 500 deduplicated sentences per language for evaluating language models on Northeast Indian languages. Languages This dataset contains test data for 9 Northeast Indian languages: Assamese (asm) - 500 sentences Garo (grt) - 500 sentences Khasi (kha) - 500 sentences Kokborok (trp) - 500 sentences Meitei (mni) - 500 sentences Mizo (lus) - 500 sentences Naga (nag) - 500 sentences Nyishi (njz) - 500 sentences Pnar (pbv) -… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeast-languages-test-set.texttext-generation1K<n<10K0 likes5 downloads4mo agoHugging Face11sensix-zo /SFT-Paite_Combined-Zo-Languagesgated SFT-Paite-Combined-Zo-Languages This dataset contains curated linguistic data for Continued Pre-Training (CPT) and Supervised Fine-Tuning (SFT). It is specifically structured for the Gemma 31B model to enhance Paite language reasoning while maintaining distinct language boundaries between related dialects. Dataset Composition The dataset follows an 80/20 distribution strategy to prioritize the primary language while providing enough context for language identification and… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/SFT-Paite_Combined-Zo-Languages.text-generation10K<n<100K0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.