datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BiasShadesInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab!
Dataset Card for BiasShades
Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators.
Dataset Details
Version: 1.0
License: SHADES 1 Montreal Data License
Dataset Description
728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.HuggingFaceFW-finetranslations-100-languages-sample
Finetranslations 100 Language Sample Dataset
Subset of HuggingFaceFW/finetranslations with the top 100 languages by number of documents.
Configurations
all: 100 languages combined (100k rows), shuffled
100 individual language configs: 1000 rows each
Columns
Original columns + language (source language indicator which is the name of the config)
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finetranslations-100-languages-sample.sa-languages-corpus
South African Languages Text Corpus
Plain-text corpus covering all 11 official South African languages, for language modeling.
bluesky-10m-posts-15-languages
Dataset Card: Bluesky 10M Multilingual
📊 Overview
Total Posts: 10,099,990
Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi)
Collection Period: August 9-12, 2026
Source: Bluesky Jetstream API (public firehose)
Format: JSONL
Size: ~3 GB
🌍 Language Distribution
Language
Code
Posts
%
English
en
6,843,995
67.8%
Japanese
ja
1,547,179
15.3%
German
de
373,626
3.7%
Portuguese
pt
331,093
3.3%
Spanish
es
325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.gsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.sa-languages
South African Languages Dataset
Dataset Overview
Language
Training Documents
Training GPT2 Tokens
Avg Tokens/Doc
Max Tokens
Test Documents
Test GPT2 Tokens
Test Avg Tokens/Doc
Test Max Tokens
isiZulu
116,693
192,622,799
1,650.68
335,530
687
1,080,961
1,573.45
15,691
Sesotho
83,329
144,337,938
1,732.15
98,542
841
1,393,086
1,656.4614,071
isiXhosa
99,567
141,484,241
1,421.00
113,710
788
1,161,296
1,473.73
17,220
isiNdebele
21,922
17,533,799
799.83
42,701… See the full description on the dataset page: https://huggingface.co/datasets/anrilombard/sa-languages.big_math_translated_african_languages
Big Math Translated -- African Languages
This is a set of 41k SynthLabsAI/Big-Math-RL-Verified questions translated into 9 African languages using Azure/GPT-4o.
We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language.
sa-nguni-languages
South African Nguni Languages Dataset
Dataset Overview
Language
Training Documents
Training GPT2 Tokens
Avg Tokens/Doc
Max Tokens
Test Documents
Test GPT2 Tokens
Test Avg Tokens/Doc
Test Max Tokens
isiZulu
116,693
192,622,799
1,650.68
335,530
687
1,080,961
1,573.45
15,691
isiXhosa
99,567
141,484,241
1,421.00
113,710
788
1,161,296
1,473.7317,220
isiNdebele
21,922
17,533,799
799.83
42,701
222
170,111
766.27
6,615
siSwati
1,668
3,148,007
1,887.29
24,129
17… See the full description on the dataset page: https://huggingface.co/datasets/anrilombard/sa-nguni-languages.proxy-mt-benchmark-scores
Proxy-MT Benchmark Scores
Multilingual benchmark results for 50 open-weight LLMs, evaluated with the
lm-evaluation-harness via a vLLM
backend. Covers reasoning, comprehension, and knowledge tasks with an emphasis on
African and other lower-resource languages.
Layout
scores/<model>.csv # parsed per-language scores (tidy, ready to plot)
raw/<model>/.../results_*.json # raw lm-eval-harness result files
raw/<model>/raw_log.txt # full evaluation… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-benchmark-scores.northeast-languages-test-set
Northeast Languages Test Set
A curated test set of 500 deduplicated sentences per language for evaluating language models on Northeast Indian languages.
Languages
This dataset contains test data for 9 Northeast Indian languages:
Assamese (asm) - 500 sentences
Garo (grt) - 500 sentences
Khasi (kha) - 500 sentences
Kokborok (trp) - 500 sentences
Meitei (mni) - 500 sentences
Mizo (lus) - 500 sentences
Naga (nag) - 500 sentences
Nyishi (njz) - 500 sentences
Pnar (pbv) -… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeast-languages-test-set.SFT-Paite_Combined-Zo-Languages
SFT-Paite-Combined-Zo-Languages
This dataset contains curated linguistic data for Continued Pre-Training (CPT) and Supervised Fine-Tuning (SFT). It is specifically structured for the Gemma 31B model to enhance Paite language reasoning while maintaining distinct language boundaries between related dialects.
Dataset Composition
The dataset follows an 80/20 distribution strategy to prioritize the primary language while providing enough context for language identification and… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/SFT-Paite_Combined-Zo-Languages.
