datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BiasShadesInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab!
Dataset Card for BiasShades
Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators.
Dataset Details
Version: 1.0
License: SHADES 1 Montreal Data License
Dataset Description
728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.abdullahkhan70_github-tech-stack-languages-and-frameworks
GitHub Tech Stack Languages & Frameworks
Comprehensive Repository Data: JavaScript, Python, Go, Rust & More
Dataset Info
Source: Kaggle
Original Size: 2.17 MB
Kaggle Downloads: 62
Files: 17
Files
Mirrored from Kaggle
languages_datasetThis dataset contains a set of 8612 languages from across the world as well as data such as Glottocode, ISO-639-3 codes, names, language families etc.
Original source: https://glottolog.org/glottolog/language
Corpus-Chadian-languages_shuNutribench_subset_with_six_languages
NutriBench Multilingual Extension (English, German, Chinese, Lao)
Dataset Description
This dataset is a 1000 subset with multilingual extension of NutriBench v2, designed for cross-lingual evaluation of nutrition estimation from meal descriptions. It retains the original English meal-description field and adds translated versions in German, Chinese, and Lao.
The dataset was constructed to compare model performance under two settings:
Direct multilingual estimation, where… See the full description on the dataset page: https://huggingface.co/datasets/MikeQian/Nutribench_subset_with_six_languages.proxy-mt-eval-scores
Proxy-MT Eval Scores
Corpus-level MT metrics for 50 open-weight LLMs on the translations in
proxy-mt-translations.
Computed by evaluate_mt.py (BLEU, chrF++, ROUGE-L, METEOR, XCOMET-XL, SSA-COMET).
MetricX is backfilled separately and may still be empty in this snapshot.
Layout
<model>/flores-200.csv
<model>/ntrex.csv
<model>/wmt24.csv
Each CSV has one row per eng-<lang> pair:
column
description
translation-pair
e.g. eng-yor
bleu
sacrebleu corpus BLEU… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-eval-scores.Nutribench_full_with_four_languages
NutriBench Multilingual Extension (English, German, Chinese, Lao)
Dataset Description
This dataset is a multilingual extension of NutriBench v2, designed for cross-lingual evaluation of nutrition estimation from meal descriptions. It retains the original English meal-description field and adds translated versions in German, Chinese, and Lao.
The dataset was constructed to compare model performance under two settings:
Direct multilingual estimation, where the model… See the full description on the dataset page: https://huggingface.co/datasets/MikeQian/Nutribench_full_with_four_languages.programming-languages-overviewMerged_Languages
