gem
Datasets
All datasets matching “gem”wiki_linguaWikiLingua is a large-scale multilingual dataset for the evaluation of
crosslingual abstractive summarization systems. The dataset includes ~770k
article and summary pairs in 18 languages from WikiHow. The gold-standard
article-summary alignments across languages was done by aligning the images
that are used to describe each how-to step in an article.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.UTKFace-geminiUTKFace dataset annotated using Google Gemini.
This dataset only contains annotation and not the image itself. (Json file name corresponds to image file name)
Used model: gemini-pro-vision
Format
{
"sex":male/female,
"attractiveness":very ugly/ugly/normal/attractive/very attractive,
"age":young child/child/adolescent/young adult/adult/young senior/senior/old/very old,
"character":kind/jealous/violent/frienly/playboy/intersting/boring,
"description":string… See the full description on the dataset page: https://huggingface.co/datasets/Aruno/UTKFace-gemini.wiki_auto_asset_turk
Dataset Card for GEM/wiki_auto_asset_turk
Link to Main Data Card
You can find the main data card on the GEM Website.
Dataset Summary
WikiAuto is an English simplification dataset that we paired with ASSET and TURK, two very high-quality evaluation datasets, as test sets. The input is an English sentence taken from Wikipedia and the target a simplified sentence. ASSET and TURK contain the same test examples but have references that are simplified in different… See the full description on the dataset page: https://huggingface.co/datasets/GEM/wiki_auto_asset_turk.xlsumWe present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally
annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
The dataset covers 45 languages ranging from low to high-resource, for many of which no
public dataset is currently available. XL-Sum is highly abstractive, concise,
and of high quality, as indicated by human and intrinsic evaluation.prolong-data-64K-gemma
