CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /mbpp Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.text1K<n<10K259 likes487k downloads3y agoHugging Face02google-research-datasets /paws Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.texttext-classification100K<n<1M40 likes239k downloads3y agoHugging Face03google-research-datasets /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.textquestion-answering10K<n<100K127 likes79k downloads3y agoHugging Face04google-research-datasets /nq_open Dataset Card for nq_open Dataset Summary The NQ-Open task, introduced by Lee et.al. 2019, is an open domain question answering benchmark that is derived from Natural Questions. The goal is to predict an English answer string for an input English question. All questions can be answered using the contents of English Wikipedia. Supported Tasks and Leaderboards Open Domain Question-Answering, EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.textquestion-answering10K<n<100K36 likes33k downloads3y agoHugging Face05google-research-datasets /tydiqa Dataset Card for "tydiqa" Dataset Summary TyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language expresses -- such that we expect models performing well on this set to generalize across a large number of the languages in the world. It contains language phenomena that would not be found in… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/tydiqa.textquestion-answering100K<n<1M38 likes14k downloads2y agoHugging Face06google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M267 likes14k downloads3y agoHugging Face07google-research-datasets /conceptual_captions Dataset Card for Conceptual Captions Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.imageimage-to-text1M<n<10M111 likes12k downloads2y agoHugging Face08google-research-datasets /paws-x Dataset Card for PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification Dataset Summary This dataset contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean. All translated pairs are sourced from examples in PAWS-Wiki. For further details, see the accompanying paper: PAWS-X: A Cross-lingual Adversarial Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws-x.texttext-classification100K<n<1M52 likes7.3k downloads3y agoHugging Face09regent-research /regent-subset-of-jat-dataset-tokenizedThis is the dataset for REGENT: A Retrieval-Augmented Generalist Agent That Can Act In-Context In New Environments. The REGENT dataset includes a subset of the JAT (tokenized) dataset (from https://huggingface.co/datasets/jat-project) for the REGENT training environments. It has around a 100k transitions from each of the 145 training environments (45 metaworld, 52 atari, 9 mujoco, 39 babyai). Please find this in the *_subset folders. It also has distance values for input sequences used in… See the full description on the dataset page: https://huggingface.co/datasets/regent-research/regent-subset-of-jat-dataset-tokenized.timeseries10M<n<100M0 likes6.7k downloads2y agoHugging Face10geodesic-research /control-pretraining-datasets-smoke geodesic-research/control-pretraining-datasets-smoke Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/control-pretraining-datasets-smoke", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/control-pretraining-datasets-smoke.text10K<n<100K0 likes4.1k downloads18d agoHugging Face11Rajarshi-Roy-research /Defactify_Image_Dataset Defactify_Image_Dataset This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection. 📝 Dataset Description Dataset Summary The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.imageimage-classification10K<n<100K22 likes2.5k downloads4mo agoHugging Face12ShawnChamberlain /open-economic-quant-research-data Open Economic & Quant Research Data Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation. Repository structure CasualLab/: causal inference and policy-simulation research content. Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.documenttabular-classificationn<1K0 likes2.2k downloads1mo agoHugging Face13google-research-datasets /poem_sentiment Dataset Card for Gutenberg Poem Dataset Dataset Summary Poem Sentiment is a sentiment dataset of poem verses from Project Gutenberg. This dataset can be used for tasks such as sentiment classification or style transfer for poems. Supported Tasks and Leaderboards [More Information Needed] Languages The text in the dataset is in English (en). Dataset Structure Data Instances Example of one instance in the dataset. {'id': 0… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/poem_sentiment.texttext-classification1K<n<10K18 likes1.4k downloads2y agoHugging Face14ibm-research /cif-dataset Cracks in the Foundation A civil-infrastructure visual inspection dataset for instance segmentation with 6 defect/condition categories: Algae · Crack · Net-Crack · Crack with Precipitation · Rust · Spalling Each sample is either a full-resolution inspection image or a 1024×1024 tile derived from one. Tiled samples carry extra fields (tile_row, tile_col, file_name_original, …) that are None for full-resolution samples. Splits Each split is its own parquet shard and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/cif-dataset.imageobject-detection100K<n<1M8 likes1.3k downloads5mo agoHugging Face15google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.2k downloads3y agoHugging Face16google-research-datasets /circa Dataset Card for CIRCA Dataset Summary The Circa (meaning ‘approximately’) dataset aims to help machine learning systems to solve the problem of interpreting indirect answers to polar questions. The dataset contains pairs of yes/no questions and indirect answers, together with annotations for the interpretation of the answer. The data is collected in 10 different social conversational situations (eg. food preferences of a friend). The following are the situational… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/circa.texttext-classification10K<n<100K7 likes1.2k downloads3y agoHugging Face17google-research-datasets /cfq Dataset Card for "cfq" Dataset Summary The Compositional Freebase Questions (CFQ) is a dataset that is specifically designed to measure compositional generalization. CFQ is a simple yet realistic, large dataset of natural language questions and answers that also provides for each question a corresponding SPARQL query against the Freebase knowledge base. This means that CFQ can also be used for semantic parsing. Supported Tasks and Leaderboards More Information… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/cfq.textquestion-answering100K<n<1M7 likes628 downloads3y agoHugging Face18google-research-datasets /xquad_r Dataset Card for [Dataset Name] Dataset Summary XQuAD-R is a retrieval version of the XQuAD dataset (a cross-lingual extractive QA dataset). Like XQuAD, XQUAD-R is an 11-way parallel dataset, where each question appears in 11 different languages and has 11 parallel correct answers across the languages. Supported Tasks and Leaderboards [More Information Needed] Languages The dataset can be found with the following languages: Arabic: xquad-r/ar.json… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/xquad_r.textquestion-answering10K<n<100K3 likes402 downloads3y agoHugging Face19obvious-research /flux-kontext-ipa-datasetimage10K<n<100K2 likes380 downloads1y agoHugging Face20Rajarshi-Roy-research /Defactify_Text_Dataset A Comprehensive Dataset for Human vs. AI Generated Text Detection This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Text Detection. Dataset Summary This comprehensive dataset comprises over 73,193 text samples designed for the detection and attribution of AI-generated text. It combines authentic New York Times articles with synthetic versions generated by several state-of-the-art Large Language Models (LLMs). The goal of the… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Text_Dataset.texttext-classification10K<n<100K1 likes360 downloads4mo agoHugging Face21google-research-datasets /disfl_qa Dataset Card for DISFL-QA: A Benchmark Dataset for Understanding Disfluencies in Question Answering Dataset Summary Disfl-QA is a targeted dataset for contextual disfluencies in an information seeking setting, namely question answering over Wikipedia passages. Disfl-QA builds upon the SQuAD-v2 (Rajpurkar et al., 2018) dataset, where each question in the dev set is annotated to add a contextual disfluency using the paragraph as a source of distractors. The final dataset… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/disfl_qa.textquestion-answering10K<n<100K7 likes324 downloads2y agoHugging Face22google-research-datasets /aquamuse Dataset Card for AQuaMuSe Dataset Summary AQuaMuSe is a novel scalable approach to automatically mine dual query based multi-document summarization datasets for extractive and abstractive summaries using question answering dataset (Google Natural Questions) and large document corpora (Common Crawl) This dataset contains versions of automatically generated datasets for abstractive and extractive query-based multi-document summarization as described in AQuaMuSe paper.… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/aquamuse.textother10K<n<100K12 likes311 downloads3y agoHugging Face23JetBrains-Research /jupyter-errors-dataset Dataset Summary The presented dataset contains 10000 Jupyter notebooks, each of which contains at least one error. In addition to the notebook content, the dataset also provides information about the repository where the notebook is stored. This information can help restore the environment if needed. Getting Started This dataset is organized such that it can be naively loaded via the Hugging Face datasets library. We recommend using streaming due to the large size… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/jupyter-errors-dataset.text1K<n<10K3 likes301 downloads3y agoHugging Face24rl-research /dr-tulu-rl-data [!NOTE] For full information, go check out the Dr Tulu paper here. DR Tulu RL Data This dataset contains the RL training data for DR Tulu, containing prompts and search-based rubrics generated from OpenScholar and SearchArena prompts, with rubrics generated using GPT-4.1. Important: This does not contain the RaR datasets we use in final RL training, but only the OpenScholar and SearchArena subsets. For the RaR data, we use data from: anisha2102/RaR-Science-20k-o3-mini… See the full description on the dataset page: https://huggingface.co/datasets/rl-research/dr-tulu-rl-data.text1K<n<10K14 likes259 downloads10mo agoHugging Face25google-research-datasets /coarse_discourse Dataset Card for "coarse_discourse" Dataset Summary A large corpus of discourse annotations and relations on ~10K forum threads. We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.texttext-classification100K<n<1M6 likes250 downloads3y agoHugging Face26Vyvo-Research /AST-Music-Data-45Kaudio10K<n<100K0 likes243 downloads11mo agoHugging Face27Vyvo-Research /AST-Music-Data-82Kaudio10K<n<100K0 likes221 downloads11mo agoHugging Face28zalizedata /global-research-output-mobility-dataset Global Research Output & Mobility (OpenAlex) Institution-level research output profiles and aggregated author mobility flows built from the official OpenAlex snapshot: institutions, sources, topics and funders. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/global-research-output-mobility-dataset Packages in this repo Package Tier Rows Size… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/global-research-output-mobility-dataset.tabulartabular-classification10M<n<100M0 likes206 downloads1mo agoHugging Face29zalizedata /us-research-grants-nih-nsf-dataset US Research Grants (NIH + NSF) Federal research funding as analysis-ready tables: NIH ExPORTER project grants and NSF awards with amounts, institutions, investigators and program metadata (no abstract full text in open packages). Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/us-research-grants-nih-nsf-dataset Formats & how to load Native parquet… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-research-grants-nih-nsf-dataset.tabulartext-classification10M<n<100M0 likes190 downloads1mo agoHugging Face30Layasaran /research_model_dataset_v1100K<n<1M0 likes184 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.