datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limitgemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.Taur_CoT_Analysis_Project___google__gemini-1.5-flash-00120260729_mini-v2.2.8_gemini-3-5-flashSWE-smith-rs-gemini-3-flash-trajectories
Trajectories Dataset
Top-level fields:
messages
instance_id
resolved
model
traj_id
patch
Generated at: 2026-02-27 00:16:39Z
Rows: 1449
Shards: 6
Skipped runs (missing/corrupt trajectory): 1
20260731_mini-v2.4.2_gemini-3-6-flashgemini-3.7-flash-ocr-26-aug-2026
gemini-3.7-flash-ocr-26-aug-2026
Page-image → transcription pairs for finetuning a vision-language model to OCR
Devanagari and Tamil printed books.
These labels are not human ground truth. They are the output of a teacher
model, so its accuracy is the ceiling for anything trained on them.
Provenance
Teacher model
google/gemini-3.7-flash (via OpenRouter, reasoning.effort=low)
Page render
PyMuPDF at 200 DPI, grayscale JPEG q90
Sampling
stratified —… See the full description on the dataset page: https://huggingface.co/datasets/kailasa-ngpt/gemini-3.7-flash-ocr-26-aug-2026.Gemini-2.0-Flash-Aoede-Voiceloracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
Synthetic Loracle supervision data generated from FineWeb with OpenRouter.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 6500
synthetic finetunes: 1284
generated finetunes in this shard: 1000
generator backend: openrouter
generator model: google/gemini-3-flash-preview
max docs per finetune: 40
max token budget per finetune: 10000
questions per finetune: 10
Configs… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes.Gemini-2.0-Flash-Fenrir-Voicegemini-3-flash-preview
Gemini 3 Flash Preview
This is a reasoning dataset created using Gemini 3 Flash Preview with a reasoning depth set to high.
The dataset is meant for creating distilled versions of Gemini 3 Flash Preview by fine-tuning already existing open-source LLMs.
This dataset contains a collection of prompts categorized by themes such as benchmarks, psychology, web development, embedded systems, creative writing, design, finance, legal, marketing, and science. No real benchmark questions were… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-flash-preview.Gemini-2.0-Flash-Kore-VoiceNano-SFT-SWE-Gym-gemini-2.5-flashgemini_3.5_flash_distilled_25k
Gemini 3.5 Flash Distilled Dataset (25k)
A 25,000-sample synthetic distilled dataset designed to replicate the core capabilities of Gemini 3.5 Flash: frontier-level agentic execution, rapid multi-step reasoning, dense context analysis, and advanced autonomous coding — all optimized for low-latency inference.
Dataset Summary
This dataset was created via template-based evolutionary synthesis with content-normalized SHA-256 deduplication. Every sample features… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gemini_3.5_flash_distilled_25k.Finch-Collection-Gemini-3-Flash
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
A mid-training "practice phase" that teaches small open-source LLMs how to evolve solutions.
👋 This is the Gemini-3-Flash teacher variant of the Finch Collection — evolutionary search trajectories from the paper Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks, but with Gemini-3-Flash as the teacher mutation… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/Finch-Collection-Gemini-3-Flash.AIME-2021-2025-Perturbed-Solutions-V1.0-Gemini-2.5-FlashHALO-Gemini-3-Flash-AppWorld
Dataset Card: Gemini 3 Flash Traces on AppWorld (test-normal)
Dataset Overview
This dataset contains agent execution traces of Gemini 3 Flash running on the AppWorld benchmark, specifically evaluated on the test-normal dataset split. The traces capture the full span-level execution detail of the model interacting with AppWorld's simulated app ecosystem.
Field
Value
Model
Gemini 3 Flash
Benchmark
AppWorld
Split
test-normal
Total Traces
168
Total Spans
3… See the full description on the dataset page: https://huggingface.co/datasets/inference-net/HALO-Gemini-3-Flash-AppWorld.sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600Dataset for sarvam's entity normalisation task. More detailed information can be found here, in the main model repo: Hugging Face
Detailed Report (Writeup): Google Drive
It also has a gguf variant, with certain additional gguf based innstructions: Hugging Face
Model inference script can be found here: Colab
Model predictions can be found in this dataset and both the repo files. named as:
eval_data_001_predictions.csv and eval_data_001_predictions_excel.csv.
train_data_001_predictions.csvand… See the full description on the dataset page: https://huggingface.co/datasets/Tasmay-Tib/sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600.Gemini-3-Flash-Preview-VIBE
Gemini 3 Flash Preview VIBE
This dataset is our first attempt at an agentic coding SFT dataset.
All of the prompts for this dataset were sourced from MiniMaxAI/VIBE.
Each prompt was given to Gemini 3 Flash Preview with the follow tools and system prompt:
read_file - Read file contents from workspace
write_file - Write content to a file
edit_file - Replace text in a file
list_directory - List files and directories
search_code - Search for patterns in files
run_command - Execute… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Gemini-3-Flash-Preview-VIBE.20260815_mini-v2.4.2_gemini-3-7-flashioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp_free-prompt-mem-limitdummy-ioi-eval-openrouter_google_gemini-2.0-flash-thinking-exp_free-testgemini_flash_ttsallenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
PJMixers-Dev/allenai_WildChat-1M-prompts with responses generated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.20260429_mini-v2.2.6_gemini-3-flashbghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.gemini-2.5-flash-11000x
Gemini 2.5 Flash - 11,000x
This dataset was created by querying Gemini 2.5 Flash over 11,000 times with the goal of compiling the behavior, reasoning traces, output style, and (most importantly) knowledge of the model into a single dataset.
Total Ouput Tokens: 54.4M
Total Dataset Cost: $134 (USD)
stats provided via OpenRouter
If you would like to see more datasets like this for other models, please consider donating to help make that happen, as we are broke college students :)… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-2.5-flash-11000x.gemini-2_5-flashWeyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Weyaxi/HelpSteer-filtered with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.gemini-3.8-flash-lru-tasks-25
Gemini 3.8 Flash LRU Cache Coding Tasks 25
A 25-task synthetic coding evaluation set generated with Gemini 3.8 Flash and
subsequently audited and corrected with ChatGPT.
The tasks focus on difficult LRU-cache and expiring-cache behavior, including
implementation, repair, edge cases, behavioral contracts, and verification.
This is a prompt-only dataset. It does not contain reference answers or
model-generated solutions.
Dataset Summary
The publication artifact… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/gemini-3.8-flash-lru-tasks-25.
