datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki_linguaWikiLingua is a large-scale multilingual dataset for the evaluation of
crosslingual abstractive summarization systems. The dataset includes ~770k
article and summary pairs in 18 languages from WikiHow. The gold-standard
article-summary alignments across languages was done by aligning the images
that are used to describe each how-to step in an article.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.UTKFace-geminiUTKFace dataset annotated using Google Gemini.
This dataset only contains annotation and not the image itself. (Json file name corresponds to image file name)
Used model: gemini-pro-vision
Format
{
"sex":male/female,
"attractiveness":very ugly/ugly/normal/attractive/very attractive,
"age":young child/child/adolescent/young adult/adult/young senior/senior/old/very old,
"character":kind/jealous/violent/frienly/playboy/intersting/boring,
"description":string… See the full description on the dataset page: https://huggingface.co/datasets/Aruno/UTKFace-gemini.wiki_auto_asset_turk
Dataset Card for GEM/wiki_auto_asset_turk
Link to Main Data Card
You can find the main data card on the GEM Website.
Dataset Summary
WikiAuto is an English simplification dataset that we paired with ASSET and TURK, two very high-quality evaluation datasets, as test sets. The input is an English sentence taken from Wikipedia and the target a simplified sentence. ASSET and TURK contain the same test examples but have references that are simplified in different… See the full description on the dataset page: https://huggingface.co/datasets/GEM/wiki_auto_asset_turk.xlsumWe present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally
annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
The dataset covers 45 languages ranging from low to high-resource, for many of which no
public dataset is currently available. XL-Sum is highly abstractive, concise,
and of high quality, as indicated by human and intrinsic evaluation.prolong-data-64K-gemmaGPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.emotion-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.gemma2-jailbreaksweb_nlgWebNLG is a bi-lingual dataset (English, Russian) of parallel DBpedia triple sets
and short texts that cover about 450 different DBpedia properties. The WebNLG data
was originally created to promote the development of RDF verbalisers able to
generate short text and to handle micro-planning (i.e., sentence segmentation and
ordering, referring expression generation, aggregation); the goal of the task is
to generate texts starting from 1 to 7 input triples which have entities in common
(so the input is actually a connected Knowledge Graph). The dataset contains about
17,000 triple sets and 45,000 crowdsourced texts in English, and 7,000 triples sets
and 19,000 crowdsourced texts in Russian. A challenging test set section with
entities and/or properties that have not been seen at training time is available.classified_images_gemmaAudioMCQ-StrongAC-GeminiCoT
[ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT
This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly.
Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis.
🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.gemGEM is a benchmark environment for Natural Language Generation with a focus on its Evaluation,
both through human annotations and automated Metrics.
GEM aims to:
- measure NLG progress across 13 datasets spanning many NLG tasks and languages.
- provide an in-depth analysis of data and models presented via data statements and challenge sets.
- develop standards for evaluation of generated text using both automated and human metrics.
It is our goal to regularly update GEM and to encourage toward more inclusive practices in dataset development
by extending existing data or developing datasets for additional languages.v2-outputs-and-scoresGPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/SicariusSicariiStuff/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.gemma-2b-suite-explanations1M_activations_pile_10k_GPT_Gemma_Qwenioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limitemotion-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it (instruct) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: snae/emotion_stories_gemma_4_4B — stories written by gemma-4-4B, a smaller EXTERNAL model (generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it.emotion-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model (whose activations these are): google/gemma-4-31b (base)
Input corpus: snae/emotion_stories_gemma_4_4B — third-person emotion stories written by gemma-4-4B, a smaller EXTERNAL model (the open replication's published corpus; generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b.twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.llama3.1_8b_inst_as_ver_gemma27b_it_math158_32gen_asyncemotion-vectors-gemma-4-31b-postfix
Emotion vectors, google/gemma-4-31b (corrected extraction)
Residual-stream activations for google/gemma-4-31b, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set re-extracts… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-postfix.xwikisThe XWikis Corpus (Perez-Beltrachini and Lapata, 2021) provides datasets with different language pairs and directions for cross-lingual abstractive document summarisation. This current version includes four languages: English, German, French, and Czech. The dataset is derived from Wikipedia. It is based on the observation that for a Wikipedia title, the lead section provides an overview conveying salient information, while the body provides detailed information. It thus assumes the body and lead paragraph as a document-summary pair. Furthermore, as a Wikipedia title can be associated with Wikipedia articles in various languages, 1) Wikipedia’s Interlanguage Links are used to find titles across languages and 2) given any two related Wikipedia titles, e.g., Huile d’Olive (French) and Olive Oil (English), the lead paragraph from one title is paired with the body of the other to derive cross-lingual pairs.xsumThis is the XSUM subset of the GEM benchmark.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.gemma-crafter-five-experiments-20260915gemini-2_5-proCreative-Writing-Gemini3Pro-2700x
Pulitzer Diamond Prose GEMINI Seeds
This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.
