datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
med-gemini-medqa-relabeled
Med-Gemini MedQA Relabelling and Analysis
This repository contains data and code corresponding to the MedQA relabelling
performed as part of [1], specifically for the results in Figure 4b and appendix
C.2.
[1] Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn,
Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, Juanma Zambrano Chaves,
Szu-Yeu Hu, Mike Schaekermann, Aishwarya Kamath, Yong Cheng, David G.T. Barrett,
Cathy Cheung, Basil… See the full description on the dataset page: https://huggingface.co/datasets/katielink/med-gemini-medqa-relabeled.gemini-finetune-datasetmultilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.gemini_testGemini_3.1_202_Task_AI_Exposure_Scores
Gemini 3.1 2026 Task AI Exposure Scores
Dataset Summary
This dataset contains task-level AI exposure labels for O*NET task statements. Each task is classified into one of four categories, E0, E1, E2, or E3, using an updated 2026 Agentic AI Exposure Rubric and a Gemini 3.1 Pro classification pipeline. The labels are designed to capture whether a task can be accelerated by a frontier agentic AI system directly, whether it would require deeper software integration, or… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/Gemini_3.1_202_Task_AI_Exposure_Scores.travel-multi-turn-chat-geminiGemini3_India_legal_Benchmark
📘 Indian Law Benchmark — 107 Questions (Strict LLM Evaluation)
📝 Dataset Summary
This dataset contains 107 Indian law questions designed to evaluate factual recall, statutory grounding, and calibrated confidence in large language models.The questions span:
Constitution of India
Indian Penal Code (IPC)
Bharatiya Nyaya Sanhita (BNS, 2023)
Code of Criminal Procedure (CrPC)
Bharatiya Nagarik Suraksha Sanhita (BNSS, 2023)
Indian Evidence Act (IEA)
Bharatiya… See the full description on the dataset page: https://huggingface.co/datasets/adhithyakiran/Gemini3_India_legal_Benchmark.medical-google-chatgpt-gemini-source-overlap
Google and AI Source Overlap Across 12 Medical Niches
An open, reproducible US dataset comparing explicit ChatGPT and Gemini citations with paired Google organic Top 20 results across 12 medical niches and 432 frozen questions.
Full study: https://rotgar.com/medical/resources/google-top-20-chatgpt-gemini-source-overlap
Version DOI: https://doi.org/10.5281/zenodo.21850734
Version: 1.0
Fieldwork: August 7, 2026
Publication date: August 8, 2026
Market and language: United States… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/medical-google-chatgpt-gemini-source-overlap.IMRAD-sections-clf-gemini-augmented
Dataset Card for IMRAD Classification Dataset (100k Rows)
Dataset Name: IMRAD Classification Dataset (100k Rows)
Dataset Description:
This dataset contains approximately 100,000 sentences extracted from scientific research papers and labeled according to their corresponding IMRAD (Introduction, Methods, Results, and Discussion) sections. The data was initially sourced from the unarXive_imrad_clf dataset on Hugging Face and expanded using data augmentation techniques. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/stormsidali2001/IMRAD-sections-clf-gemini-augmented.AIME-2024-Gemini-2.5-ProHC3-Gemini-Flash-Responses
Single-class dataset. Every row is AI-generated (Is_AI = 1). It is
designed to measure generator shift and must be paired with human text, e.g.
the human side of HC3 PLUS, to form a balanced detection benchmark.
HC3 Gemini 2.0 Flash Responses
Dataset Description
This dataset contains {len(df_upload):,} AI-generated text samples produced by
Google Gemini 2.0 Flash in response to questions from the
HC3 (Human ChatGPT Comparison Corpus) benchmark.
It was… See the full description on the dataset page: https://huggingface.co/datasets/mohamedmady/HC3-Gemini-Flash-Responses.Synthetic_Story_Generated_By_Geminiacl-ocl-fork-gemini-power-responsesfineweb-edu-gemini-annotations-portuguese-regressiongemini_filtered_sft_traces_simplified_reasoninggemini-2.0-flash-lite-pneumonia-datasettravel-multi-turn-chat-geminigemini_fake_newsA few fake news articles I created with Gemini for research purposes.
ts-detect-test-smell-gemini-explainedVideoUFO_Caption_Gemini_2.5_Pro
The detailed captions for our VideoUFO dataset generated by Gemini 2.5 pro
Gemini_1.5_flashECT-Gemini-summarization-datasetgemini3PhotoEditBattleResults-Gemini-EXT
