datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai2_arc
Dataset Card for "ai2_arc"
Dataset Summary
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in
advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains
only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also
including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.openbookqa
Dataset Card for OpenBookQA
Dataset Summary
OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic
(with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In
particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge,
and rich text comprehension.
OpenBookQA is a new kind of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/openbookqa.winogrande
Dataset Card for "winogrande"
Dataset Summary
WinoGrande is a new collection of 44k problems, inspired by Winograd Schema Challenge (Levesque, Davis, and Morgenstern
2011), but adjusted to improve the scale and robustness against the dataset-specific bias. Formulated as a
fill-in-a-blank task with binary options, the goal is to choose the right option for a given sentence which requires
commonsense reasoning.
Supported Tasks and Leaderboards
More… See the full description on the dataset page: https://huggingface.co/datasets/allenai/winogrande.sciq
Dataset Card for "sciq"
Dataset Summary
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/sciq.swag
Dataset Card for Situations With Adversarial Generations
Dataset Summary
Given a partial description like "she opened the hood of the car,"
humans can reason about the situation and anticipate what might come
next ("then, she examined the engine"). SWAG (Situations With Adversarial Generations)
is a large-scale dataset for this task of grounded commonsense
inference, unifying natural language inference and physically grounded reasoning.
The dataset consists of 113k… See the full description on the dataset page: https://huggingface.co/datasets/allenai/swag.quartz
Dataset Card for "quartz"
Dataset Summary
QuaRTz is a crowdsourced dataset of 3864 multiple-choice questions about open domain qualitative relationships. Each
question is paired with one of 405 different background sentences (sometimes short paragraphs).
The QuaRTz dataset V1 contains 3864 questions about open domain qualitative relationships. Each question is paired with
one of 405 different background sentences (sometimes short paragraphs).
The dataset is split into… See the full description on the dataset page: https://huggingface.co/datasets/allenai/quartz.qasc
Dataset Card for "qasc"
Dataset Summary
QASC is a question-answering dataset with a focus on sentence composition. It consists of 9,980 8-way multiple-choice
questions about grade school science (8,134 train, 926 dev, 920 test), and comes with a corpus of 17M sentences.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/qasc.scitail
Dataset Card for "scitail"
Dataset Summary
The SciTail dataset is an entailment dataset created from multiple-choice science exams and web sentences. Each question
and the correct answer choice are converted into an assertive statement to form the hypothesis. We use information
retrieval to obtain relevant text from a large text corpus of web sentences, and use these sentences as a premise P. We
crowdsource the annotation of such premise-hypothesis pair as supports… See the full description on the dataset page: https://huggingface.co/datasets/allenai/scitail.whisper_transcriptions.reazon_speech_all.wer_10.0.vectorizedMolmoAct-Midtraining-Mixture
MolmoAct - Midtraining Mixture
Data Mixture used for MolmoAct Midtraining. Contains MolmoAct Dataset formulated as Action Reasoning Data.
MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique manipulation tasks in both home and tabletop environments. It has… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Midtraining-Mixture.tulu-3-sft-mixture
Tulu 3 SFT Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
The Tulu 3 SFT mixture was used to train the Tulu 3 series of models.
It contains 939,344 samples from the following sets:
CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024)
FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture.whisper_transcriptions.reazon_speech_allmolmobot-data
MolmoBot-data
Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms:
DoorOpeningDataGenConfig
RBY1OpenDataGenConfig
RBY1PickDataGenConfig
FrankaPickOmniCamConfig
RBY1PickAndPlaceDataGenConfig
FrankaPickAndPlaceOmniCamConfig
FrankaPickAndPlaceColorOmniCamConfig
FrankaPickAndPlaceNextToOmniCamConfig
Please note that every package indexed by the parquet files can contain several instances of episode data.
We also provide an… See the full description on the dataset page: https://huggingface.co/datasets/allenai/molmobot-data.WildChat-1M
Dataset Card for WildChat
Dataset Description
Paper: https://arxiv.org/abs/2405.01470
Interactive Search Tool: https://wildvisualizer.com (paper)
License: ODC-BY
Language(s) (NLP): multi-lingual
Point of Contact: Yuntian Deng
Dataset Summary
WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat by… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-1M.MolmoAct2-BimanualYAM-DatasetThis dataset was created using LeRobot.
MolmoAct2-BimanualYAM Dataset
This repository is the merged ckpt / merged LeRobot dataset artifact for the MolmoAct2-BimanualYAM Dataset, a large-scale collection of bimanual robot manipulation demonstrations collected for MolmoAct2. Across the full collection, MolmoAct2-BimanualYAM contains more than 720 hours of training demonstrations spanning diverse tabletop manipulation tasks.
Language Annotations
This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct2-BimanualYAM-Dataset.MolmoAct-DatasetThis dataset was created using LeRobot.
Dataset Description
This dataset contains MolmoAct Dataset in lerobot format. All contents in this dataset were collected in-house by Ai2.
Quick links:
📂 All Models
📂 All Data
📃 Paper
🎥 Blog Post
🎥 Video
Code
License and Use
This dataset is licensed under CC BY-4.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Citation
@misc{molmoact2025… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Dataset.IFBench_test
License
This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use.
Citation
Please cite:
@misc{pyatkin2025generalizing,
title={Generalizing Verifiable Instruction Following},
author={Valentina Pyatkin and Saumya Malik and Victoria Graf and Hamish Ivison and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_test.art
Dataset Card for "art"
Dataset Summary
ART consists of over 20k commonsense narrative contexts and 200k explanations.
The Abductive Natural Language Inference Dataset from AI2.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
anli
Size of downloaded dataset files: 5.12 MB
Size of the generated dataset: 34.36 MB
Total amount of disk used: 39.48… See the full description on the dataset page: https://huggingface.co/datasets/allenai/art.tulu-3-sft-personas-instruction-following
Dataset Descriptions
This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset.
To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.vigil-jailbreak-all-MiniLM-L6-v2
Vigil: LLM Jailbreak all-MiniLM-L6-v2
Repo: github.com/deadbits/vigil-llm
Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs.
This repository contains all-MiniLM-L6-v2 embeddings for all "jailbreak" prompts used by Vigil.
You can use the parquet2vdb.py utility to load the embeddings in the Vigil chromadb instance, or use them in your own… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-jailbreak-all-MiniLM-L6-v2.ropes
Dataset Card for ROPES
Dataset Summary
ROPES (Reasoning Over Paragraph Effects in Situations) is a QA dataset which tests a system's ability to apply knowledge from a passage of text to a new situation. A system is presented a background passage containing a causal or qualitative relation(s) (e.g., "animal pollinators increase efficiency of fertilization in flowers"), a novel situation that uses this background, and questions that require reasoning about effects of the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ropes.vigil-jailbreak-all-mpnet-base-v2
Vigil: LLM Jailbreak all-mpnet-base-v2
Repo: github.com/deadbits/vigil-llm
Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs.
This repository contains all-mpnet-base-v2 embeddings for all "jailbreak" prompts used by Vigil.
You can use the parquet2vdb.py utility to load the embeddings in the Vigil chromadb instance, or use them in your own… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-jailbreak-all-mpnet-base-v2.MolmoAct-Pretraining-Mixture
MolmoAct - Pretraining Mixture
Data Mixture used for MolmoAct Pretraining. Contains a subset of OXE formulated as Action Reasoning Data along with auxiliary robot data and link to Multimodal Web data.
MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Pretraining-Mixture.wildguardmix
Dataset Card for WildGuardMix
Disclaimer:
The data includes examples that might be disturbing, harmful or upsetting. It includes a range of harmful topics such as discriminatory language and discussions
about abuse, violence, self-harm, sexual content, misinformation among other high-risk categories. The main goal of this data is for advancing research in building safe LLMs.
It is recommended not to train a LLM exclusively on the harmful examples.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/allenai/wildguardmix.CoSyn-400K
CoSyn-400k
CoSyn-400k is a collection of synthetic question-answer pairs about very diverse range of computer-generated images.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
and using GPT-4o mini to generate Q/A pairs based on the code (without using the rendered image).
The code used to generate this data is open source.
Synthetic pointing data is available in a seperate repo.
Quick links:
📃 CoSyn… See the full description on the dataset page: https://huggingface.co/datasets/allenai/CoSyn-400K.hellaswagscirepevalreward-bench
Code | Leaderboard | Prior Preference Sets | Results | Paper
Reward Bench Evaluation Dataset Card
The RewardBench evaluation dataset evaluates capabilities of reward models over the following categories:
Chat: Includes the easy chat subsets (alpacaeval-easy, alpacaeval-length, alpacaeval-hard, mt-bench-easy, mt-bench-medium)
Chat Hard: Includes the hard chat subsets (mt-bench-hard, llmbar-natural, llmbar-adver-neighbor, llmbar-adver-GPTInst, llmbar-adver-GPTOut… See the full description on the dataset page: https://huggingface.co/datasets/allenai/reward-bench.WildChat-4.8M
Dataset Card for WildChat-4.8M
Dataset Description
Interactive Search Tool: https://wildvisualizer.com
WildChat paper: https://arxiv.org/abs/2405.01470
WildVis paper: https://arxiv.org/abs/2409.03753
Point of Contact: Yuntian Deng
Dataset Summary
WildChat-4.8M is a collection of 3,199,860 conversations between human users and ChatGPT. This version only contains non-toxic user inputs and ChatGPT responses, as flagged by the OpenAI Moderations API or… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-4.8M.Molmo2-SynMultiImageQA
Molmo2-SynMultiImageQA
Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images, including charts, tables, documents, diagrams, etc.
The synthetic data is generated by extending the CoSyn framework into multi-image settings,
with Claude-sonnet-4-5 as the coding LLM to generate code that can be executed to render an image.
Then, we use GPT-5 to generate question-answer pairs with code (without using the rendered… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-SynMultiImageQA.
