datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-agi3-agy-gemini3.1pro-tr87
ARC-AGI-3 tr87 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game tr87, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-tr87.arc-agi3-agy-gemini3.1pro-g50t
ARC-AGI-3 g50t — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game g50t, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-g50t.testmyCFbench1111arc-agi3-agy-gemini3.1pro-su15
ARC-AGI-3 su15 — Agent Trajectories (agy-gemini3.1pro)
Gameplay trajectories from the harness×model pair agy-gemini3.1pro playing the
ARC-AGI-3 game su15, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-agy-gemini3.1pro-su15.gemma7b-summarize-eval-by-gemini15flashloracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
Synthetic Loracle supervision data generated from FineWeb with OpenRouter.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 6500
synthetic finetunes: 1284
generated finetunes in this shard: 1000
generator backend: openrouter
generator model: google/gemini-3-flash-preview
max docs per finetune: 40
max token budget per finetune: 10000
questions per finetune: 10
Configs… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes.deepsearchqa-gemini2-reasoningNano-SFT-SWE-Gym-gemini-2.5-flashtest_0909_geminiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 4,
"total_frames": 904,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/simheo/test_0909_gemini.GeminiMol-QSARpolaris-acemath-gemini-rubrics-v22026-08-20-difficult-advice-gemini-716-smoke
Difficult-advice SFT corpus, all-gemini arm: the Teaching Claude Why recipe with the entire generator stack swapped from Anthropic (difficult_advice.yaml baseline) to Gemini. google/gemini-3.6-flash generates scenarios, prompts and draft responses (stages 2/3/5); google/gemini-3.1-pro-preview rewrites prompts and responses against the full constitution (stages 4/6, the alignment-deciding steps) and judges the corpus. A behavioural difference vs the baseline corpus is attributable… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-difficult-advice-gemini-716-smoke.Finch-Collection-Gemini-3-Flash
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
A mid-training "practice phase" that teaches small open-source LLMs how to evolve solutions.
👋 This is the Gemini-3-Flash teacher variant of the Finch Collection — evolutionary search trajectories from the paper Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks, but with Gemini-3-Flash as the teacher mutation… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/Finch-Collection-Gemini-3-Flash.vepqa-gemini-1k-correct-balanced-20test_geminiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 1,
"total_frames": 579,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SteveNguyen/test_gemini.med-gemini-medqa-relabeled
Med-Gemini MedQA Relabelling and Analysis
This repository contains data and code corresponding to the MedQA relabelling
performed as part of [1], specifically for the results in Figure 4b and appendix
C.2.
[1] Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn,
Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, Juanma Zambrano Chaves,
Szu-Yeu Hu, Mike Schaekermann, Aishwarya Kamath, Yong Cheng, David G.T. Barrett,
Cathy Cheung, Basil… See the full description on the dataset page: https://huggingface.co/datasets/katielink/med-gemini-medqa-relabeled.emolia-voicenet-gemini-annotations
Emolia VoiceNet Gemini Annotations
468,180 dimension-level annotations over 236,613 Emolia speech clips,
each scored 0-6 (0-2 for the content-safety dimension) on one of 57 perceptual
voice / speech dimensions - arousal, valence, brightness, resonance placement, speaking
styles, genuineness, recording quality, and more - by Gemini 3.5 Flash (non-thinking,
temperature 0). This repository ships the annotations, audio provenance, per-dimension
statistics, and the full scoring… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-voicenet-gemini-annotations.bird-train-gemini3-flash
Dataset Card for Think2SQL-SFT
This dataset is a distilled Supervised Fine-Tuning (SFT) dataset designed to improve the reasoning capabilities of models in Text-to-SQL tasks.
It contains high-quality reasoning traces and SQL queries generated by Gemini 3 Flash.
Paper: Think2SQL: Blueprinting Reward Density and Advantage Scaling for Effective Text-To-SQL Reasoning
Base Benchmark: BIRD-Train
Dataset Description
The dataset consists of 9,428 high-quality traces, of… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-2321/bird-train-gemini3-flash.oasst1-en-hun-gemini
Open assistant 1 dataset hungarian translation (english subset)
This dataset contains hungarian translations for the oasst1 dataset's english subset. The translations were done via gemini pro and the model was instructed to keep stlye, meaning and english entites as they are. I think this produced a higher quality translation than google translate, but even this version is far from perfect.
The exact code used for creating the dataset can be found here.
license:… See the full description on the dataset page: https://huggingface.co/datasets/jazzysnake01/oasst1-en-hun-gemini.multilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.browsecomp-plus-sel-tools-test300-gemini-3p1-pro-v1gemini_meva_full_score_onlygemini_meva_full_analyze_ratecybench-gemini-2.5-protaubench-gemini-traces
taubench-gemini-traces
Complete HTTP-level agentic traces from running taubench_gemini benchmark tasks through an instrumented reverse proxy.
Each trace captures full request/response pairs including system prompts, user messages, assistant responses, tool calls and results, and token usage metadata.
Stats
Total sessions: 115
Multi-turn sessions (2+ LLM calls): 115
Total records: 5744
Total LLM requests: 2872
Format
Raw JSONL traces from the instrumented… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/taubench-gemini-traces.pandalm-gemini-annotatedfsmbench_what_will_be_the_state_geminiGemini_data_badalpaca-cleaned-gemini-hun-ratingsEz az adathalmaz úgy keletkezett, hogy a Bazsalanszky/alpaca-cleaned-gemini-hun-n lefuttattam egy llm által támogatott értékelést.
Az értékelő modell a gemini-pro (az ingyenes) volt. A használt kód az alpagasus módosítása: https://github.com/boapps/alpagasus-hu
jailbreak-gemini-2.5-pro
