datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
Synthetic Loracle supervision data generated from FineWeb with OpenRouter.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 6500
synthetic finetunes: 1284
generated finetunes in this shard: 1000
generator backend: openrouter
generator model: google/gemini-3-flash-preview
max docs per finetune: 40
max token budget per finetune: 10000
questions per finetune: 10
Configs… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes.Nano-SFT-SWE-Gym-gemini-2.5-flashFinch-Collection-Gemini-3-Flash
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
A mid-training "practice phase" that teaches small open-source LLMs how to evolve solutions.
👋 This is the Gemini-3-Flash teacher variant of the Finch Collection — evolutionary search trajectories from the paper Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks, but with Gemini-3-Flash as the teacher mutation… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/Finch-Collection-Gemini-3-Flash.loracles-finetune-gemini-3-flash
loracles-finetune-gemini-3-flash
Synthetic Loracle supervision data generated from FineWeb.
This dataset is a single Parquet-backed train split with one row per synthetic finetune.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 2800
synthetic finetunes: 587
generated finetunes uploaded: 35
generator backend: openrouter
generator model: google/gemini-3-flash-preview
max docs per finetune: 40
max token budget per finetune: 10000… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracles-finetune-gemini-3-flash.clbench-exploitable-poker-wm_summar-gemini-flash-3-1-liteapollo_english_books_translated_to_dutch_with_geminiflash15
Data description
Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
apollo_english_guidelines_translated_to_dutch_with_geminiflash1.5
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
qgqa-unified-processed-gemini-3-flash-previewgemini-flash-2.0-speech-MOSS-AnnotatedGemini-2.5-Flash-customerservice-Human-evaluator_2_datame-dj-0520-gemini-2.0-flash-001gemini-2.0-flash-responsesdsl_icl_eval-2025_01_21_180629_model-google-gemini-flash-1.5_fewshot-5dsl_icl_eval-2025_01_24_185911_model-google-gemini-flash-15_fewshot-5dsl_icl_eval-2025_01_30_180617_model-google-gemini-flash-15_fewshot-25wikimia24_hard_64_seed2-Paraphrased-Gemini-2.5-Flash-v3.0BookMIA-Paraphrased-Gemini-2.5-FlashWikiMIA-2024-Hard-Paraphrased-Gemini-2.5-Flashgemini-2.5-flash-customerservice-context-summarization-llm-judge-data
customer-service-context-summarization-evaluation-data
Lakshan2003/gemini-2.5-flash-customerservice-context-summarization-llm-judge-data
Dataset updated with context summarization evaluation columns.
This README refresh triggers Hugging Face metadata re-index.
qgqa-gemini-3-flash-20260213-041708dsl_icl_eval-2025_01_21_160647_model-google-gemini-flash-1.5_fewshot-25finewiki-rag-questions-gemini-2.5-flash-preview-09-2025finewiki-rag-answers-gemini-2.5-flash-lite-preview-09-2025frontier-cs-sampled-gemini-3-flash-preview-N8-temp1.0-eval-feedbackdsl_icl_eval-2025_01_25_202916_model-google-gemini-flash-15_fewshot-25dsl_icl_eval-2025_01_26_123531_model-google-gemini-flash-15_fewshot-5dsl_icl_eval-2025_01_26_184531_model-google-gemini-flash-15_fewshot-25dsl_icl_eval-2025_01_30_162524_model-google-gemini-flash-15_fewshot-5gemini__gemini-2.0-flash_eval_5ed6
gemini/gemini-2.0-flash Evaluation Results
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME25
LiveCodeBenchv5
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
HLE
SWEbench
AIME24
Accuracy
27.0
27.5
80.5
88.6
36.6
35.2
64.0
41.9
34.4
8.0
0.0
36.7
AIME25
Average Accuracy: 27.0% ± 0.7%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
26.7%
8
30
2
26.7%
8… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/gemini__gemini-2.0-flash_eval_5ed6.gemini-2.0-flash-lite-pneumonia-dataset
