datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BiasShadesInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab!
Dataset Card for BiasShades
Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators.
Dataset Details
Version: 1.0
License: SHADES 1 Montreal Data License
Dataset Description
728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.gsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.grokipedia-wikipedia-16-languages
Dataset description
This dataset contains a mapping between Grokipedia v0.1 article pages and the corresponding Wikipedia article titles across 16 language editions (based on Wikipedia and Wikidata dumps from 1 November 2025). Each record includes:
The URL of the Grokipedia page: grokipedia_url
Wikipedia titles in the following languages (if they exist): ar (Arabic), de (German), en (English), es (Spanish), fa (Persian), fr (French), it (Italian) , ja (Japanese), nl (Dutch), pl… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/grokipedia-wikipedia-16-languages.African-Languages_Sentiments
African Languages Sentiment Dataset (Hausa, Yorùbá, Swahili)
A stitched multi-source sentiment classification dataset combining three
independently collected sentiment corpora for Hausa, Yorùbá, and Swahili,
built for the Adaption Labs AutoScientist Challenge
(Language category).
Companion model: fine-tuned weights trained on the adapted version of this dataset via
AutoScientist are released separately at… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/African-Languages_Sentiments.classification_Turkic_languages
Description
A dataset with texts and the categories to which these texts belong.
Usage
This dataset can be used to check language models for the correct classification of texts by category.
Dataset structure:
lang: the language to which the text source belongs;
title: the title of the text;
original_text: original text taken from a web page;
processed_text: processed text using preprocessing functions;
category: the category to which the text belongs;… See the full description on the dataset page: https://huggingface.co/datasets/Electrotubbie/classification_Turkic_languages.abdullahkhan70_github-tech-stack-languages-and-frameworks
GitHub Tech Stack Languages & Frameworks
Comprehensive Repository Data: JavaScript, Python, Go, Rust & More
Dataset Info
Source: Kaggle
Original Size: 2.17 MB
Kaggle Downloads: 62
Files: 17
Files
Mirrored from Kaggle
world-languages-dataset
🌍 World Languages Dataset
This dataset contains a list of official and unofficial languages categorized by language families...
languages_datasetThis dataset contains a set of 8612 languages from across the world as well as data such as Glottocode, ISO-639-3 codes, names, language families etc.
Original source: https://glottolog.org/glottolog/language
triplets_Turkic_languages
Triplets for Turkic languages language models
Description
This dataset is designed to test models for working with Next Sentence Prediction (NSP) and Sentence Order Prediction (SOP). It includes two sub-sets with triplets of texts..
Usage
This dataset can be used to train and evaluate models capable of performing NSP and SAP tasks.
Dataset structure:
Each entry in the dataset represents three values:
text: a triplet of text;
flag: a flag indicating… See the full description on the dataset page: https://huggingface.co/datasets/Electrotubbie/triplets_Turkic_languages.Nutribench_subset_with_six_languages
NutriBench Multilingual Extension (English, German, Chinese, Lao)
Dataset Description
This dataset is a 1000 subset with multilingual extension of NutriBench v2, designed for cross-lingual evaluation of nutrition estimation from meal descriptions. It retains the original English meal-description field and adds translated versions in German, Chinese, and Lao.
The dataset was constructed to compare model performance under two settings:
Direct multilingual estimation, where… See the full description on the dataset page: https://huggingface.co/datasets/MikeQian/Nutribench_subset_with_six_languages.proxy-mt-translations
Proxy-MT Translations
English→X machine translations generated with vLLM
across 50 open-weight LLMs on three evaluation benchmarks. This dataset holds the
raw model outputs (one CSV per model × dataset × target language); metric scores
(BLEU / chrF / COMET / MetricX) live in proxy-mt-eval-scores.
Layout
flores-200/<model>/eng-<lang>.csv # 119 target languages
ntrex/<model>/eng-<lang>.csv # 87 target languages
wmt24/<model>/eng-<lang>.csv # 51… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-translations.Corpus-Chadian-languages_shuproxy-mt-eval-scores
Proxy-MT Eval Scores
Corpus-level MT metrics for 50 open-weight LLMs on the translations in
proxy-mt-translations.
Computed by evaluate_mt.py (BLEU, chrF++, ROUGE-L, METEOR, XCOMET-XL, SSA-COMET).
MetricX is backfilled separately and may still be empty in this snapshot.
Layout
<model>/flores-200.csv
<model>/ntrex.csv
<model>/wmt24.csv
Each CSV has one row per eng-<lang> pair:
column
description
translation-pair
e.g. eng-yor
bleu
sacrebleu corpus BLEU… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-eval-scores.Nutribench_full_with_four_languages
NutriBench Multilingual Extension (English, German, Chinese, Lao)
Dataset Description
This dataset is a multilingual extension of NutriBench v2, designed for cross-lingual evaluation of nutrition estimation from meal descriptions. It retains the original English meal-description field and adds translated versions in German, Chinese, and Lao.
The dataset was constructed to compare model performance under two settings:
Direct multilingual estimation, where the model… See the full description on the dataset page: https://huggingface.co/datasets/MikeQian/Nutribench_full_with_four_languages.cross_lingual_languagesprogramming-languages-overviewghana-languagesMerged_Languages
