datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mbpp
Dataset Card for Mostly Basic Python Problems (mbpp)
Dataset Summary
The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us.
Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.paws
Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling
Dataset Summary
PAWS: Paraphrase Adversaries from Word Scrambling
This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset.
For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.natural_questions
Dataset Card for Natural Questions
Dataset Summary
The NQ corpus contains questions from real users, and it requires QA systems to
read and comprehend an entire Wikipedia article that may or may not contain the
answer to the question. The inclusion of real user questions, and the
requirement that solutions should read an entire page to find the answer, cause
NQ to be a more realistic and challenging task than prior QA datasets.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.nq_open
Dataset Card for nq_open
Dataset Summary
The NQ-Open task, introduced by Lee et.al. 2019,
is an open domain question answering benchmark that is derived from Natural Questions.
The goal is to predict an English answer string for an input English question.
All questions can be answered using the contents of English Wikipedia.
Supported Tasks and Leaderboards
Open Domain Question-Answering,
EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.tydiqa
Dataset Card for "tydiqa"
Dataset Summary
TyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs.
The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language
expresses -- such that we expect models performing well on this set to generalize across a large number of the languages
in the world. It contains language phenomena that would not be found in… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/tydiqa.go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.conceptual_captions
Dataset Card for Conceptual Captions
Dataset Summary
Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.paws-x
Dataset Card for PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification
Dataset Summary
This dataset contains 23,659 human translated PAWS evaluation pairs and
296,406 machine translated training pairs in six typologically distinct
languages: French, Spanish, German, Chinese, Japanese, and Korean. All
translated pairs are sourced from examples in
PAWS-Wiki.
For further details, see the accompanying paper:
PAWS-X: A Cross-lingual Adversarial Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws-x.regent-subset-of-jat-dataset-tokenizedThis is the dataset for REGENT: A Retrieval-Augmented Generalist Agent That Can Act In-Context In New Environments.
The REGENT dataset includes a subset of the JAT (tokenized) dataset (from https://huggingface.co/datasets/jat-project) for the REGENT training environments.
It has around a 100k transitions from each of the 145 training environments (45 metaworld, 52 atari, 9 mujoco, 39 babyai).
Please find this in the *_subset folders.
It also has distance values for input sequences used in… See the full description on the dataset page: https://huggingface.co/datasets/regent-research/regent-subset-of-jat-dataset-tokenized.control-pretraining-datasets-smoke
geodesic-research/control-pretraining-datasets-smoke
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/control-pretraining-datasets-smoke", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/control-pretraining-datasets-smoke.Defactify_Image_Dataset
Defactify_Image_Dataset
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection.
📝 Dataset Description
Dataset Summary
The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.open-economic-quant-research-data
Open Economic & Quant Research Data
Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation.
Repository structure
CasualLab/: causal inference and policy-simulation research content.
Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.poem_sentiment
Dataset Card for Gutenberg Poem Dataset
Dataset Summary
Poem Sentiment is a sentiment dataset of poem verses from Project Gutenberg.
This dataset can be used for tasks such as sentiment classification or style transfer for poems.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The text in the dataset is in English (en).
Dataset Structure
Data Instances
Example of one instance in the dataset.
{'id': 0… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/poem_sentiment.cif-dataset
Cracks in the Foundation
A civil-infrastructure visual inspection dataset for instance segmentation with 6 defect/condition categories:
Algae · Crack · Net-Crack · Crack with Precipitation · Rust · Spalling
Each sample is either a full-resolution inspection image or a 1024×1024 tile derived from one.
Tiled samples carry extra fields (tile_row, tile_col, file_name_original, …) that are None for full-resolution samples.
Splits
Each split is its own parquet shard and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/cif-dataset.discofuse
Dataset Card for "discofuse"
Dataset Summary
DiscoFuse is a large scale dataset for discourse-based sentence fusion.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
discofuse-sport
Size of downloaded dataset files: 4.33 GB
Size of the generated dataset: 15.04 GB
Total amount of disk used: 19.36 GB
An example of 'train' looks as follows.
{… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.circa
Dataset Card for CIRCA
Dataset Summary
The Circa (meaning ‘approximately’) dataset aims to help machine learning systems to solve the problem of interpreting indirect answers to polar questions.
The dataset contains pairs of yes/no questions and indirect answers, together with annotations for the interpretation of the answer. The data is collected in 10 different social conversational situations (eg. food preferences of a friend).
The following are the situational… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/circa.cfq
Dataset Card for "cfq"
Dataset Summary
The Compositional Freebase Questions (CFQ) is a dataset that is specifically designed to measure compositional
generalization. CFQ is a simple yet realistic, large dataset of natural language questions and answers that also
provides for each question a corresponding SPARQL query against the Freebase knowledge base. This means that CFQ can
also be used for semantic parsing.
Supported Tasks and Leaderboards
More Information… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/cfq.xquad_r
Dataset Card for [Dataset Name]
Dataset Summary
XQuAD-R is a retrieval version of the XQuAD dataset (a cross-lingual extractive
QA dataset). Like XQuAD, XQUAD-R is an 11-way parallel dataset, where each
question appears in 11 different languages and has 11 parallel correct answers
across the languages.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The dataset can be found with the following languages:
Arabic: xquad-r/ar.json… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/xquad_r.flux-kontext-ipa-datasetDefactify_Text_Dataset
A Comprehensive Dataset for Human vs. AI Generated Text Detection
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Text Detection.
Dataset Summary
This comprehensive dataset comprises over 73,193 text samples designed for the detection and attribution of AI-generated text. It combines authentic New York Times articles with synthetic versions generated by several state-of-the-art Large Language Models (LLMs). The goal of the… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Text_Dataset.disfl_qa
Dataset Card for DISFL-QA: A Benchmark Dataset for Understanding Disfluencies in Question Answering
Dataset Summary
Disfl-QA is a targeted dataset for contextual disfluencies in an information seeking setting, namely question answering over Wikipedia passages. Disfl-QA builds upon the SQuAD-v2 (Rajpurkar et al., 2018) dataset, where each question in the dev set is annotated to add a contextual disfluency using the paragraph as a source of distractors.
The final dataset… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/disfl_qa.aquamuse
Dataset Card for AQuaMuSe
Dataset Summary
AQuaMuSe is a novel scalable approach to automatically mine dual query based multi-document summarization datasets for extractive and abstractive summaries using question answering dataset (Google Natural Questions) and large document corpora (Common Crawl)
This dataset contains versions of automatically generated datasets for abstractive and extractive query-based multi-document summarization as described in AQuaMuSe paper.… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/aquamuse.jupyter-errors-dataset
Dataset Summary
The presented dataset contains 10000 Jupyter notebooks,
each of which contains at least one error. In addition to the notebook content,
the dataset also provides information about the repository where the notebook is stored.
This information can help restore the environment if needed.
Getting Started
This dataset is organized such that it can be naively loaded via the Hugging Face datasets library. We recommend using streaming due to the large size… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/jupyter-errors-dataset.dr-tulu-rl-data
[!NOTE]
For full information, go check out the Dr Tulu paper here.
DR Tulu RL Data
This dataset contains the RL training data for DR Tulu, containing prompts and search-based rubrics generated from OpenScholar and SearchArena prompts, with rubrics generated using GPT-4.1.
Important: This does not contain the RaR datasets we use in final RL training, but only the OpenScholar and SearchArena subsets. For the RaR data, we use data from:
anisha2102/RaR-Science-20k-o3-mini… See the full description on the dataset page: https://huggingface.co/datasets/rl-research/dr-tulu-rl-data.coarse_discourse
Dataset Card for "coarse_discourse"
Dataset Summary
A large corpus of discourse annotations and relations on ~10K forum threads.
We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.AST-Music-Data-45KAST-Music-Data-82Kglobal-research-output-mobility-dataset
Global Research Output & Mobility (OpenAlex)
Institution-level research output profiles and aggregated author mobility flows built from the official OpenAlex snapshot: institutions, sources, topics and funders.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/global-research-output-mobility-dataset
Packages in this repo
Package
Tier
Rows
Size… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/global-research-output-mobility-dataset.us-research-grants-nih-nsf-dataset
US Research Grants (NIH + NSF)
Federal research funding as analysis-ready tables: NIH ExPORTER project grants and NSF awards with amounts, institutions, investigators and program metadata (no abstract full text in open packages).
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/us-research-grants-nih-nsf-dataset
Formats & how to load
Native parquet… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-research-grants-nih-nsf-dataset.research_model_dataset_v1
