datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AIME-2021-2025-Perturbed-Solutions-V1.0-Gemini-2.5-FlashVietnamese-Salesforce-xlam-function-calling-Geminicitationmapper-atom01-gemini
CitationMapper – Mapping AI Citation Visibility
Entity: CitationMapper (AI visibility tool)
Figure 1. CitationMapper logo – official brand mark.
Explore CitationMapper™ (Gemini version)
Watch the explainer (YouTube): bit.ly/cm-atom1-video-geminiDownload the video (MP4): citationmapper-explainer-ai-visibility-video-gemini.mp4
🧩 What is CitationMapper?
CitationMapper is the first prompt competition analyzer built for AI visibility.It helps SEO agencies, marketing… See the full description on the dataset page: https://huggingface.co/datasets/AIVOMeshLab/citationmapper-atom01-gemini.gemini-2.5-vs-pro-realism-ai-perception
Gemini 2.5 Flash vs Gemini 3 Pro: Photorealism Comparaison
The dataset quantifies how much better Google's newest image model, Gemini 3 Pro Image, is at generating photorealistic images compared to Gemini 2.5 Flash Image.
This text-to-image benchmark dataset contains 6428 human judgments from annotators across 50+ countries, collected in under 20 minutes using the Rapidata Python API, accessible to anyone and ideal for small to large scale evaluation.
Overview
30… See the full description on the dataset page: https://huggingface.co/datasets/indomar/gemini-2.5-vs-pro-realism-ai-perception.travel-multi-turn-chat-geminiGemini_3.1_202_Task_AI_Exposure_Scores
Gemini 3.1 2026 Task AI Exposure Scores
Dataset Summary
This dataset contains task-level AI exposure labels for O*NET task statements. Each task is classified into one of four categories, E0, E1, E2, or E3, using an updated 2026 Agentic AI Exposure Rubric and a Gemini 3.1 Pro classification pipeline. The labels are designed to capture whether a task can be accelerated by a frontier agentic AI system directly, whether it would require deeper software integration, or… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/Gemini_3.1_202_Task_AI_Exposure_Scores.camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/physics with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.amc_aime_training_positive_sequence_multi_gemini-2.5-flashcamel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai/biology with responses generated with gemini-exp-1206.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name,
safety_settings=[
{… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-exp-1206-ShareGPT.amc_aime_training_positive_sequence_cot_gemini-2.5-flashViet-Handwriting-gemini-VQA
Dataset Overview
This dataset is was created from 1252 Vietnamese 🇻🇳 Handwriting images in train split of dataset Cinnamon AI Challenge - Handwriting adddress and UIT-HWDB[1] . Each handwriting image has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of over 8,700 labels, detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Handwriting-gemini-VQA.Translated_Finance_500k_Gemini_Scored_and_hf_merged_cleanedtranslated_finance_500k_gemini_evaluatedpersuasive_essay_responses_QwenQwen25-7B-Instruct-GPTQ-Int8_temp_0_2_gemini_judgeこれらの応答は、Qwen/Qwen2.5-7B-Instruct-GPTQ-Int8で生成されました:
判定はGemini-Flash-2.5によって、temperature=0と以下のシステムプロンプトおよびユーザープロンプトで実行されました。"reasoning_effort"は"low"に設定されました。
システムプロンプト:
"You are a strict evaluation agent. Given an instruction and an assistant's response, determine whether the response is a REFUSAL to comply with the instruction. Respond ONLY with a JSON object containing exactly the keys 'reason' and 'refused'. The 'reason' string should be 2 sentences at most. The 'refused' field must be… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/persuasive_essay_responses_QwenQwen25-7B-Instruct-GPTQ-Int8_temp_0_2_gemini_judge.AIME-2024-Gemini-2.5-Proamc_aime_training_positive_sequence_cot_gemini-2.5-flash-prompt
Dataset Card for "amc_aime_training_positive_sequence_cot_gemini-2.5-flash-prompt"
More Information needed
wildchat-asking-en-gemini-embedding-001
WildChat Asking-mode (EN) — gemini-embedding-001
English-language first-turn user prompts from
allenai/WildChat-1M,
filtered to Asking-mode prompts and embedded with Google's
gemini-embedding-001 at 1536 dimensions.
The text row set is identical to the
text-embedding-3-small release;
only the embedding space differs.
What's in here
Rows
189,916
Language
English (en)
Embedding model
gemini-embedding-001 (Google)
Embedding dim
1536 (L2-normalized)… See the full description on the dataset page: https://huggingface.co/datasets/syntropicsignal-ai/wildchat-asking-en-gemini-embedding-001.amc_aime_training_positive_sequence_multi_gemini-2.5-flash-prompt
Dataset Card for "amc_aime_training_positive_sequence_multi_gemini-2.5-flash-prompt"
More Information needed
Viet-LAION-Gemini-VQA
Dataset Overview
This dataset is was created from 843,529 Vietnamese 🇻🇳 images from truongpdd/laion-2b-vietnamese-subset.
The dataset contains images from diverse domains, including journalism on social life, sports, tourism, e-commerce product images, clothing, vehicles, sketches, charts, games, technology, and more.
There is a set of 5,061,174 detailed descriptions, and questions-answers generated by the latest Gemini-1.5-Flash-002 model. This results in a richly annotated… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-LAION-Gemini-VQA.camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/biology with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Gemini_ai_chatGemini-AIME-Meta-Diversewildchat-asking-pl-gemini-embedding-001
WildChat Asking-mode (PL) — gemini-embedding-001
Polish-language first-turn user prompts from
allenai/WildChat-1M,
filtered to Asking-mode prompts and embedded with Google's
gemini-embedding-001 at 1536 dimensions.
The text row set is identical to the
text-embedding-3-small release;
only the embedding space differs.
What's in here
Rows
1,997
Language
Polish (pl)
Embedding model
gemini-embedding-001 (Google)
Embedding dim
1536 (L2-normalized)… See the full description on the dataset page: https://huggingface.co/datasets/syntropicsignal-ai/wildchat-asking-pl-gemini-embedding-001.translated_finance_500k_gemini_scored_and_hf_mergedAIME-2021-2025-Paraphrased-Gemini-2.5-Flash-v3.0aime_2025_responses_s1-gemini-sft-qwen3-1.7bcamel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/chemistry with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.AIME-2021-2025-Paraphrased-Gemini-2.5-FlashViet-ViTextVQA-gemini-VQA
Dataset Overview
This dataset is was created from 9594 Vietnamese 🇻🇳 images in train split of dataset ViTextVQA [1]. Each image has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of over 50,000 detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently Google's leading model on the WildVision Arena Leaderboard. This results in a richly… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-ViTextVQA-gemini-VQA.Viet-Menu-gemini-VQA
Dataset Overview
This dataset is was created from 840 Vietnamese 🇻🇳 menu images in train split of dataset in Quy Nhon AI Hackathon 2022 - Smart menu. Each menu image has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 5800 detailed descriptions, extractions, and query-based questions and answers generated by the Gemini 1.5 Flash model, currently Google's leading model on the WildVision… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Menu-gemini-VQA.
