datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai2_arc
Dataset Card for "ai2_arc"
Dataset Summary
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in
advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains
only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also
including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.openbookqa
Dataset Card for OpenBookQA
Dataset Summary
OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic
(with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In
particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge,
and rich text comprehension.
OpenBookQA is a new kind of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/openbookqa.sciq
Dataset Card for "sciq"
Dataset Summary
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/sciq.quartz
Dataset Card for "quartz"
Dataset Summary
QuaRTz is a crowdsourced dataset of 3864 multiple-choice questions about open domain qualitative relationships. Each
question is paired with one of 405 different background sentences (sometimes short paragraphs).
The QuaRTz dataset V1 contains 3864 questions about open domain qualitative relationships. Each question is paired with
one of 405 different background sentences (sometimes short paragraphs).
The dataset is split into… See the full description on the dataset page: https://huggingface.co/datasets/allenai/quartz.qasc
Dataset Card for "qasc"
Dataset Summary
QASC is a question-answering dataset with a focus on sentence composition. It consists of 9,980 8-way multiple-choice
questions about grade school science (8,134 train, 926 dev, 920 test), and comes with a corpus of 17M sentences.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/qasc.math_qaOur dataset is gathered by using a new representation language to annotate over the AQuA-RAT dataset. AQuA-RAT has provided the questions, options, rationale, and the correct options.WildChat-1M
Dataset Card for WildChat
Dataset Description
Paper: https://arxiv.org/abs/2405.01470
Interactive Search Tool: https://wildvisualizer.com (paper)
License: ODC-BY
Language(s) (NLP): multi-lingual
Point of Contact: Yuntian Deng
Dataset Summary
WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat by… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-1M.ropes
Dataset Card for ROPES
Dataset Summary
ROPES (Reasoning Over Paragraph Effects in Situations) is a QA dataset which tests a system's ability to apply knowledge from a passage of text to a new situation. A system is presented a background passage containing a causal or qualitative relation(s) (e.g., "animal pollinators increase efficiency of fertilization in flowers"), a novel situation that uses this background, and questions that require reasoning about effects of the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ropes.reward-bench
Code | Leaderboard | Prior Preference Sets | Results | Paper
Reward Bench Evaluation Dataset Card
The RewardBench evaluation dataset evaluates capabilities of reward models over the following categories:
Chat: Includes the easy chat subsets (alpacaeval-easy, alpacaeval-length, alpacaeval-hard, mt-bench-easy, mt-bench-medium)
Chat Hard: Includes the hard chat subsets (mt-bench-hard, llmbar-natural, llmbar-adver-neighbor, llmbar-adver-GPTInst, llmbar-adver-GPTOut… See the full description on the dataset page: https://huggingface.co/datasets/allenai/reward-bench.WildChat-4.8M
Dataset Card for WildChat-4.8M
Dataset Description
Interactive Search Tool: https://wildvisualizer.com
WildChat paper: https://arxiv.org/abs/2405.01470
WildVis paper: https://arxiv.org/abs/2409.03753
Point of Contact: Yuntian Deng
Dataset Summary
WildChat-4.8M is a collection of 3,199,860 conversations between human users and ChatGPT. This version only contains non-toxic user inputs and ChatGPT responses, as flagged by the OpenAI Moderations API or… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-4.8M.qasperA dataset containing 1585 papers with 5049 information-seeking questions asked by regular readers of NLP papers, and answered by a separate set of NLP practitioners.WildChat
Dataset Card for WildChat
Note: a newer version with 4.8 million conversations and demographic information can be found here.
Dataset Description
Paper: https://arxiv.org/abs/2405.01470
Interactive Search Tool: https://wildvisualizer.com (paper)
License: ODC-BY
Language(s) (NLP): multi-lingual
Point of Contact: Yuntian Deng
Dataset Summary
WildChat is a collection of 650K conversations between human users and ChatGPT. We collected WildChat… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat.reward-bench-2Code | Leaderboard | Results | Paper
RewardBench 2 Evaluation Dataset Card
The RewardBench 2 evaluation dataset is the new version of RewardBench that is based on unseen human data and designed to be substantially more difficult! RewardBench 2 evaluates capabilities of reward models over the following categories:
Factuality (NEW!): Tests the ability of RMs to detect hallucinations and other basic errors in completions.
Precise Instruction Following (NEW!): Tests the ability of RMs… See the full description on the dataset page: https://huggingface.co/datasets/allenai/reward-bench-2.preference-test-sets
Preference Test Sets
Very few preference datasets have heldout test sets for validation of reward model accuracy results.
In this dataset, we curate the test sets from popular preference datasets into a common schema for easy loading and evaluation.
Anthropic HH (Helpful & Harmless Agent and Red Teaming), test set in full is 8552 samples
Anthropic HHH Alignment (Helpful, Honest, & Harmless), formatted from Big Bench for standalone evaluation.
Learning to summarize, downsampled from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/preference-test-sets.tulu-v2-sft-mixture
Dataset Card for Tulu V2 Mix
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
Tulu is a series of language models that are trained to act as helpful assistants.
The dataset consists of a mix of :
FLAN (Apache 2.0): We use 50,000 examples sampled from FLAN v2. To emphasize CoT-style reasoning, we sample another 50,000 examples… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture.quorefQuoref is a QA dataset which tests the coreferential reasoning capability of reading comprehension systems. In this
span-selection benchmark containing 24K questions over 4.7K paragraphs from Wikipedia, a system must resolve hard
coreferences before selecting the appropriate span(s) in the paragraphs for answering questions.quacQuestion Answering in Context is a dataset for modeling, understanding,
and participating in information seeking dialog. Data instances consist
of an interactive dialog between two crowd workers: (1) a student who
poses a sequence of freeform questions to learn as much as possible
about a hidden Wikipedia text, and (2) a teacher who answers the questions
by providing short excerpts (spans) from the text. QuAC introduces
challenges not found in existing machine comprehension datasets: its
questions are often more open-ended, unanswerable, or only meaningful
within the dialog context.opus-doctor-patient-conversations-all-human-diseases
Opus-4.8-High-Thinking generated Doctor-Patient Conversations for All Human Diseases
Covers every human disease listed on my previous work here: nisten/all-human-diseases
The dataset strictly used Opus 4.8 - High and was cleaned over 3 times via Opus 4.8, 4.7 and 4.6. Minor corrections were needed upon each pass mainly to bypass single word safety filters like i.e. monkeypox.
The main hallucination noticed during generation was that Opus would make up wrong PMID ( PubMed ID )… See the full description on the dataset page: https://huggingface.co/datasets/nisten/opus-doctor-patient-conversations-all-human-diseases.ALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V
├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.master-dataset-all-V2
Master Dataset All V2
Google NQ Sequentially Sharded Dataset.
klej-dyk
klej-dyk
Description
The Czy wiesz? (eng. Did you know?) the dataset consists of almost 5k question-answer pairs obtained from Czy wiesz... section of Polish Wikipedia. Each question is written by a Wikipedia collaborator and is answered with a link to a relevant Wikipedia article. In huggingface version of this dataset, they chose the negatives which have the largest token overlap with a question.
Tasks (input, output, and metrics)
The task is to predict if… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-dyk.tulu-v3.1-mix-preview-4096-OLMoE
OLMoE SFT Mix
The SFT mix used is an expanded version of the Tulu v2 SFT mix with new additions for code, CodeFeedback-Filtered-Instruction, reasoning, MetaMathQA, and instruction following, No Robots and a subset of Daring Anteater.
Please see the referenced datasets for the multiple licenses used in subsequent data.
We do not introduce any new data with this dataset.
Config for creation via open-instruct:
dataset_mixer:
allenai/tulu-v2-sft-mixture-olmo-4096: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v3.1-mix-preview-4096-OLMoE.assist-llm-function-calling
Function Calling dataset for Assist LLM for Home Assistant
This dataset is generated by using other conversation agent pipelines as teachers
from the deivce-actions-v2 dataset.
This dataset is used to support fine tuning of llama based models.
See Device Actions for a notebook for construction of this dataset and the device-actions dataset.
prescience
PreScience: A Benchmark for Forecasting Scientific Contributions
Dataset Summary
Can AI systems trained on the scientific record up to a fixed point in time forecast the scientific advances that follow? Such a capability could help researchers identify collaborators and impactful research directions, and anticipate which problems and methods will become central next. We introduce PreScience, a scientific forecasting benchmark that decomposes the research process into four… See the full description on the dataset page: https://huggingface.co/datasets/allenai/prescience.zestZEST tests whether NLP systems can perform unseen tasks in a zero-shot way, given a natural language description of
the task. It is an instantiation of our proposed framework "learning from task descriptions". The tasks include
classification, typed entity extraction and relationship extraction, and each task is paired with 20 different
annotated (input, output) examples. ZEST's structure allows us to systematically test whether models can generalize
in five different ways.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.alloy-sovereign-eval-runs
Alloy Sovereign Eval Runs · the honest first measured run
Append-only measured eval runs produced by routing SZL's
K-Verify Benchmark v1
through the live Alloy governed-inference stack on SZL's own sovereign
metal (provider: sovereign, zero cloud, zero spend). Each row is one
inference: its verdict, latency, NVML-measured energy, and a
signed receipt id that is re-checkable against the live Alloy receipt chain.
Built and maintained by SZL Holdings. Apache-2.0.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/alloy-sovereign-eval-runs.letterboxd-all-movie-data
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/letterboxd-all-movie-data.allenai-WildChat
AllenAI WildChat Combined Dataset
This unofficial repository provides the AllenAI WildChat Combined Dataset, which merges the WildChat-4.8M and WildChat-1M collections of human–ChatGPT conversations.
WildChat-1M contains 1 million chats, of which 25.53% are from GPT‑4 and the remainder from GPT‑3.5. These conversations cover a wide range of complex interactions, including code-switching, ambiguity, and political topics.
WildChat-4.8M originally comprised 4.8 million conversations.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-WildChat.ALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/ALLaVA-4V.
