datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hellaswag
Dataset Card for "hellaswag"
Dataset Summary
HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 71.49 MB
Size of the generated dataset: 65.32 MB
Total… See the full description on the dataset page: https://huggingface.co/datasets/Rowan/hellaswag.hellaswagm_hellaswag
Multilingual HellaSwag
Dataset Summary
This dataset is a machine translated version of the HellaSwag dataset.
The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
HC3Human ChatGPT Comparison Corpus (HC3)HC3-ChineseHuman ChatGPT Comparison Corpus (HC3) Chinese Versionhellaswag-multilingualOpenDebateEvidence-Anonymized
Dataset Card for OpenDebateEvidence (Anonymized)
A collection of evidence used in collegiate and high school debate competitions,
with all debater-identifying columns removed.
This is an anonymized redistribution of
Yusuf5/OpenCaselist. The
argumentative content is byte-for-byte unchanged. 26 of the original 45 columns
have been dropped. See Anonymization for exactly what was
removed and why.
Dataset Details
Dataset Description
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.okapi_hellaswag
okapi_hellaswag
Multilingual translation of Hellaswag.
Dataset Details
Dataset Description
Hellaswag is a commonsense inference challenge dataset. Though its questions are
trivial for humans (>95% accuracy), state-of-the-art models struggle (<48%). This is
achieved via Adversarial Filtering (AF), a data collection paradigm wherein a
series of discriminators iteratively select an adversarial set of machine-generated
wrong answers. AF proves to be surprisingly… See the full description on the dataset page: https://huggingface.co/datasets/jon-tow/okapi_hellaswag.EMO-transcribed-1lineNemogrounding_datasetFor desktop and web datasets in GUI grounding, the data is generally collected via screenshots alongside accessibility tools like A11y or HTML parsers to extract element structure and bounding boxes. However, these bounding boxes may sometimes be misaligned with the visual rendering due to UI animations or timing inconsistencies. In our work, we primarily rely on datasets curated from Aria-UI and OS-Atlas, which we found to be cleaner and better aligned than alternative data collections.
To… See the full description on the dataset page: https://huggingface.co/datasets/HelloKKMe/grounding_dataset.FinQA_finetung_datset
FinQA Fine-tuning Dataset
Dataset Description
This dataset is a cleaned version of the FinQA dataset prepared for financial language model fine-tuning.
Each example contains:
context: Financial report text and tables.
question: Financial reasoning question.
answer: Expected answer.
Dataset Structure
Train: 6251 examples
Test: 1147 examples
Features
Column
Description
context
Financial document context including tables… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_finetung_datset.hellaswag_de
HellaSwag (DE) — Boldt German Evaluation Suite
Improved German translation of the HellaSwag benchmark (Zellers et al., 2019), part of the Boldt German Evaluation Suite. HellaSwag is a commonsense natural language inference benchmark in which models must select the most plausible continuation of a short activity or situation description from four candidates.
Translation
This version was translated from the English original using Tower+ 72B by translating complete… See the full description on the dataset page: https://huggingface.co/datasets/Boldt/hellaswag_de.testset_hellaswagFinQA_TAT-QA_financial_finetuning_dataset
Dataset Summary
This dataset provides a unified, flattened context / question / answer format for
question answering over financial documents that combine tabular and textual data. It is
built to support training and evaluating models on numerical and discrete reasoning
tasks in the finance domain, drawing on the structure and style of established
finance-QA benchmarks such as TAT-QA and FinQA.
Each example pairs a passage of financial context (derived from a table and/or… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_TAT-QA_financial_finetuning_dataset.EMO-MEAD-TranscribedDebateSum
DebateSum
Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset"
Arxiv pre-print available here: https://arxiv.org/abs/2011.07251
Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9
Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/
Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.ro_hellaswag
Dataset Description
Hellaswag is a commonsense inference challenge dataset.
Here we provide the Romanian translation of the Hellaswag from the paper "Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback" (Lai et al., 2023).
This dataset is used as a benchmark and is part of the evaluation protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_hellaswag.opengpt-x_hellaswagxThis is a copy of the translations from openGPT-X/hellaswagx, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the HellaSwag dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_hellaswagx.SWE-bench-all__style-3__fs-oraclehellaswagCC-FilteredCorpus
English Cleaned Common Crawl Markdown Dataset
An English-focused dataset created from Common Crawl, cleaned and converted to Markdown.
The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting.
Features
English-focused
Cleaned and filtered web content
HTML converted to Markdown
Exact and near-duplicate filtering
GPT-2 perplexity filtering
Stored as compressed Parquet shards
Source
The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.SWE-bench__style-3__fs-oraclehellaswagultra
🤯HellaSwagUltra
📘 Overview
HellaSwagUltra is a large-scale multilingual commonsense reasoning benchmark that covers 60+ languages and contains over 160k+ test instances, grounded in local cultural knowledge.It aims to address the saturation of existing commonsense benchmarks (e.g., HellaSwag, StoryCloze) and the lack of culturally diverse, multilingual evaluation datasets.
Unlike conventional reasoning tests, HellaSwagUltra embeds two implicit commonsense or… See the full description on the dataset page: https://huggingface.co/datasets/aialt/hellaswagultra.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.enron_emails_parsedhellaswagultra
🤯HellaSwagUltra
📘 Overview
HellaSwagUltra is a large-scale multilingual commonsense reasoning benchmark that covers 60+ languages and contains over 160k+ test instances, grounded in local cultural knowledge.It aims to address the saturation of existing commonsense benchmarks (e.g., HellaSwag, StoryCloze) and the lack of culturally diverse, multilingual evaluation datasets.
Unlike conventional reasoning tests, HellaSwagUltra embeds two implicit commonsense or… See the full description on the dataset page: https://huggingface.co/datasets/hellaswagultra/hellaswagultra.hellaswagXTransEnV_hellaswag
Version 2 (2026-08): Dialect configs regenerated with a stronger pipeline
The 18 dialect configs (AAVE, AppE, AuE, AuE_V, BahE, EAngE, IrE, Manx, NZE,
N_Eng, NfE, OzE, SE_AmE, SE_Eng, SW_Eng, ScE, TdCE, WaE) were regenerated with an
upgraded Trans-EnV pipeline. The ESL configs (A_*/B_*) are unchanged (v1).
Previous versions of all files remain available via git revisions of this repo.
What changed
Transformation model: google/gemma-2-27b-it → google/gemma-4-31B-it,
with a… See the full description on the dataset page: https://huggingface.co/datasets/jiyounglee0523/TransEnV_hellaswag.HellaSwag_DPO_FewShot
Dataset Card for "HellaSwag_DPOP_FewShot"
HellaSwag is a dataset containing commonsense inference questions known to be hard for LLMs.
In the original dataset, each instance consists of a prompt, with one correct completion and three incorrect completions.
We create a paired preference-ranked dataset by creating three pairs for each correct response in the training split.
An example prompt is "Then, the man writes over the snow covering the window of a car, and a woman wearing… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/HellaSwag_DPO_FewShot.Hellenic-greek-parliamentary-speech
HParl: Hellenic Parliamentary Speech Corpus
Dataset Description
Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors.
Link to the original source: https://inventory.clarin.gr/corpus/1602
HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.
