datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Creative-Writing-Gemini3Pro-2700x
Pulitzer Diamond Prose GEMINI Seeds
This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.gemini-3-pro-preview20260729_mini-v2.2.8_gemini-3-5-flash20260731_mini-v2.4.2_gemini-3-6-flashgoogle-gemini-3-pro-pre-release-model-card
Gemini 3 Pro Model Card
⚠️ This is a mirror to the pre-release model card taht was officially published by Google @ https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf and subsequentially yanked. It will get outdated by the actual release version in a few hours (date of publication Nov 18, 13:16 CET)
Model card published: November, 2025
Gemini 3 Pro - Model Card
Model Cards are intended to provide essential information on… See the full description on the dataset page: https://huggingface.co/datasets/multimodalart/google-gemini-3-pro-pre-release-model-card.20260429_mini-v2.2.6_gemini-3-1-progemini_3_pro_swebench_verified_traj
Live-SWE-agent: live, self-evolving software agent
Live-SWE-agent is the first live, runtime self-evolving software engineering agent that expands and revises its own capabilities on the fly while working on a real-world issue.
Our key insight is that software agents are themselves software systems, and modern LLM-based agents already possess the intrinsic capability to extend or modify their own behavior at runtime.
gemini-3-pro-preview-high-reasoning-1000xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API.
This dataset includes 250x from TeichAI/gemini-3-pro-preview-high-reasoning-250x
Stats
Costs: $ 32.7 (USD)
Tokens: 2.73 M… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-pro-preview-high-reasoning-1000x.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.gemini-3-pro-preview-high-reasoning-250xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API.
Stats
Costs: $ 5.78
Total tokens (input + output): 484 K
20260815_mini-v2.4.2_gemini-3-7-flashGemini-3-Pro-Opus-4.5-Kimi-K2.5-13000x-formatted
Gemini-3-Pro-Opus-4.5-Kimi-K2.5-13000x (TeichAI Format)
Formatted to match TeichAI/claude-4.5-opus-high-reasoning-250x format.
Source
Original: crownelius/Gemini-3-Pro-Opus-4.5-Kimi-K2.5-13000x
Original Examples: 13916
Converted Examples: 13916
Format
{
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
],
"metadata": {...}
}
Files… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Gemini-3-Pro-Opus-4.5-Kimi-K2.5-13000x-formatted.google-gemini-3-pro-pre-release-model-card
Gemini 3 Pro Model Card
⚠️ This is a mirror to the pre-release model card taht was officially published by Google @ https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf and subsequentially yanked. It will get outdated by the actual release version in a few hours (date of publication Nov 18, 13:16 CET)
Model card published: November, 2025
Gemini 3 Pro - Model Card
Model Cards are intended to provide essential information on… See the full description on the dataset page: https://huggingface.co/datasets/Nonkapeni/google-gemini-3-pro-pre-release-model-card.20260429_mini-v2.2.6_gemini-3-flashvisualpuzzles-direct-gemini3progemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/AbderrahmanSkiredj1/gemini-3-pro-10000x-hard-high-reasoning.v3-2k-traj-gemini-3-proswebench-verified-gemini-3-progemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/ofankit/gemini-3-pro-10000x-hard-high-reasoning.gemini-3-pro-preview-high-reasoning-250xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API.
Stats
Costs: $ 5.78
Total tokens (input + output): 484 K
rebuttal_s740_e760_n_paths_1_medium_0.7_16384_gemini-3_1-pro-preview_n10.jsonl__raw
Dataset: connections-dev/connection_queries_jan12
This dataset was generated using the inference script with the following configuration:
Inference Parameters
Model Configuration
Model Name: gemini-3.1-pro-preview
Server URL: Not specified
API Key: Not provided
Request Timeout: 120 seconds
Query Configuration
Query Type: natural
Query Column: query
Sampling Type: n_paths
Generation Parameters
Temperature: 0.7
Max Tokens: 16384
Number… See the full description on the dataset page: https://huggingface.co/datasets/connections-dev/rebuttal_s740_e760_n_paths_1_medium_0.7_16384_gemini-3_1-pro-preview_n10.jsonl__raw.imoproofbench_n4_t8_l32k2k_v03.00step000500_withdelimiter_1767594169_gemini_3_pro_graders700_e720_original_1_high_0.7_16384_gemini-3-pro-preview__raw
Dataset: connections-dev/connection_queries_jan12
This dataset was generated using the inference script with the following configuration:
Inference Parameters
Model Configuration
Model Name: gemini-3-pro-preview
Server URL: Not specified
API Key: Not provided
Request Timeout: 30 seconds
Query Configuration
Query Type: natural
Query Column: query
Sampling Type: original
Generation Parameters
Temperature: 0.7
Max Tokens: 16384
Number of… See the full description on the dataset page: https://huggingface.co/datasets/connections-dev/s700_e720_original_1_high_0.7_16384_gemini-3-pro-preview__raw.Omni-MATH-Qwen3-4B-16k_hard_gemini_iter1-eval-non-zero-promptsconnection_queries_jan12_natural_original_1_medium_0.7_4096_gemini-3-pro-preview
Dataset: connections-dev/connection_queries_jan12
This dataset was generated using the inference script with the following configuration:
Inference Parameters
Model Configuration
Model Name: gemini-3-pro-preview
Server URL: Not specified
API Key: Not provided
Request Timeout: 30 seconds
Query Configuration
Query Type: natural
Query Column: query
Sampling Type: original
Generation Parameters
Temperature: 0.7
Max Tokens: 4096
Number of… See the full description on the dataset page: https://huggingface.co/datasets/connections-dev/connection_queries_jan12_natural_original_1_medium_0.7_4096_gemini-3-pro-preview.s660_e680_original_1_high_0.7_16384_gemini-3-pro-preview__raw
Dataset: connections-dev/connection_queries_jan12
This dataset was generated using the inference script with the following configuration:
Inference Parameters
Model Configuration
Model Name: gemini-3-pro-preview
Server URL: Not specified
API Key: Not provided
Request Timeout: 30 seconds
Query Configuration
Query Type: natural
Query Column: query
Sampling Type: original
Generation Parameters
Temperature: 0.7
Max Tokens: 16384
Number of… See the full description on the dataset page: https://huggingface.co/datasets/connections-dev/s660_e680_original_1_high_0.7_16384_gemini-3-pro-preview__raw.rebuttal_s860_e880_n_paths_1_medium_0.7_16384_gemini-3_1-pro-preview_n10.jsonl__raw
Dataset: connections-dev/connection_queries_jan12
This dataset was generated using the inference script with the following configuration:
Inference Parameters
Model Configuration
Model Name: gemini-3.1-pro-preview
Server URL: Not specified
API Key: Not provided
Request Timeout: 120 seconds
Query Configuration
Query Type: natural
Query Column: query
Sampling Type: n_paths
Generation Parameters
Temperature: 0.7
Max Tokens: 16384
Number… See the full description on the dataset page: https://huggingface.co/datasets/connections-dev/rebuttal_s860_e880_n_paths_1_medium_0.7_16384_gemini-3_1-pro-preview_n10.jsonl__raw.Omni-MATH-Qwen3-4B-16k_hard_gemini_iter1-promptsfusion-dataset-gemini-3-pro-preview-sfts440_e460_original_1_medium_0.7_16384_gemini-3_1-pro-preview__raw_intermediate
