datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
m_hellaswag
Multilingual HellaSwag
Dataset Summary
This dataset is a machine translated version of the HellaSwag dataset.
The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
grounding_datasetFor desktop and web datasets in GUI grounding, the data is generally collected via screenshots alongside accessibility tools like A11y or HTML parsers to extract element structure and bounding boxes. However, these bounding boxes may sometimes be misaligned with the visual rendering due to UI animations or timing inconsistencies. In our work, we primarily rely on datasets curated from Aria-UI and OS-Atlas, which we found to be cleaner and better aligned than alternative data collections.
To… See the full description on the dataset page: https://huggingface.co/datasets/HelloKKMe/grounding_dataset.ro_hellaswag
Dataset Description
Hellaswag is a commonsense inference challenge dataset.
Here we provide the Romanian translation of the Hellaswag from the paper "Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback" (Lai et al., 2023).
This dataset is used as a benchmark and is part of the evaluation protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_hellaswag.opengpt-x_hellaswagxThis is a copy of the translations from openGPT-X/hellaswagx, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the HellaSwag dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_hellaswagx.hellaswagultra
🤯HellaSwagUltra
📘 Overview
HellaSwagUltra is a large-scale multilingual commonsense reasoning benchmark that covers 60+ languages and contains over 160k+ test instances, grounded in local cultural knowledge.It aims to address the saturation of existing commonsense benchmarks (e.g., HellaSwag, StoryCloze) and the lack of culturally diverse, multilingual evaluation datasets.
Unlike conventional reasoning tests, HellaSwagUltra embeds two implicit commonsense or… See the full description on the dataset page: https://huggingface.co/datasets/aialt/hellaswagultra.hellaswagultra
🤯HellaSwagUltra
📘 Overview
HellaSwagUltra is a large-scale multilingual commonsense reasoning benchmark that covers 60+ languages and contains over 160k+ test instances, grounded in local cultural knowledge.It aims to address the saturation of existing commonsense benchmarks (e.g., HellaSwag, StoryCloze) and the lack of culturally diverse, multilingual evaluation datasets.
Unlike conventional reasoning tests, HellaSwagUltra embeds two implicit commonsense or… See the full description on the dataset page: https://huggingface.co/datasets/hellaswagultra/hellaswagultra.hellaswag-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language.
Dataset Card for Hellaswag-Turkish
malhajar/hellaswag-turkish is a translated version of hellaswag aimed specifically to be used in the OpenLLMTurkishLeaderboard
This Dataset contains rigid tests extracted from the paper Can a Machine Really Finish Your Sentence? published at ACL2019.… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/hellaswag-tr.chat_dataset_bilibilidata borrows from https://github.com/linyiLYi/bilibot
Indic-Hellaswag
Hellaswag Translated
Citation:
@inproceedings{zellers2019hellaswag,
title={HellaSwag: Can a Machine Really Finish Your Sentence?},
author={Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin},
booktitle ={Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
year={2019}
}
Contributions:Thanks to @Srinidhi9113 for adding the dataset.
hellaswag_italian
HellaSwag - Italian (IT)
This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence.
Dataset Details
The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.ko_hellaswag
Korean HellaSwag
hellaswag 영어 데이터셋을 한국어로 번역
https://huggingface.co/datasets/Rowan/hellaswag
Structure
{
"ind": 24,
"activity_label": "지붕 슁글 제거",
"ctx_a": "한 남자가 지붕 위에 앉아 있다.",
"ctx_b": "그",
"ctx": "한 남자가 지붕 위에 앉아 있다. 그",
"endings": [
"스키 한 켤레를 감싸기 위해 랩을 사용하고 있습니다.",
"레벨 타일을 뜯어내고 있습니다.",
"루빅스 큐브를 들고 있습니다.",
"지붕에 지붕을 올리기 시작합니다."
],
"source_id": "activitynet~v_-JhWjGDPHMY",
"split": "val",
"split_type": "indomain",
"label": "3"
}
{...}
HelloBenchHelloBench is an open-source benchmark designed to evaluate the long text generation capabilities of large language models (LLMs) from HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models.
hello-ot-imagenet-pca
HELLO representative ImageNet VA-VAE PCA features
This dataset contains four independently generated float32 NumPy matrices,
with shapes (262144, d) for d = 4, 32, 256, 2048, totaling about 2.29 GiB.
All use seed=42 and serve the public main-scaling and parameter-sensitivity
experiments. Larger sample counts are outside the published data scope.
The artifact contains numeric PCA-projected features only. It contains no
images, labels, captions, filenames, or ImageNet identifiers.… See the full description on the dataset page: https://huggingface.co/datasets/WenzhouXia/hello-ot-imagenet-pca.AWPCD
Arithmetic Word Problem Compendium Dataset (AWPCD)
Dataset Description
The dataset is a comprehensive collection of mathematical word problems spanning multiple domains with rich metadata and natural language variations. The problems contain 1 - 5 steps of mathematical operations that are specifically designed to encourage showing work and maintaining appropriate decimal precision throughout calculations.
The available data is a sample of 1,000 problems, and commerical… See the full description on the dataset page: https://huggingface.co/datasets/HelloCephalopod/AWPCD.yoga_posesm_hellaswag
Multilingual HellaSwag
Dataset Summary
This dataset is a machine translated version of the HellaSwag dataset.
The languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
The NUS Deep Learning Lab contributed to this effort by standardizing the dataset, ensuring consistent question formatting and alignment across all languages. This standardization enhances cross-linguistic… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/m_hellaswag.LuauDataset1.0Mercy_Tracker
Mercy Tracker Bot
A Discord bot for tracking the Mercy system in Raid: Shadow Legends using slash commands. Enhanced with improved error handling, data backup, and user-friendly features.
Features
🔮 Track summons for Legendary and Mythical champions
📊 Support for Ancient, Void, Sacred, Primal, and Remnant shards
💾 Automatic data backup and recovery
🎯 Visual progress bars and detailed mercy information
⚡ Slash command interface with validation
🛡️ Comprehensive error… See the full description on the dataset page: https://huggingface.co/datasets/HellscythePT/Mercy_Tracker.hellaswag-mk
Hellaswag MK version
This dataset is a Macedonian adaptation of the hellaswag dataset, originally curated (English -> Serbian) by Aleksa Gordić. It was translated from Serbian to Macedonian using the Google Translate API.
You can find this dataset as part of the macedonian-llm-eval GitHub and HuggingFace.
This dataset is used for training and evaluating models as described in Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language
Why… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/hellaswag-mk.repro-sharp-inequalities-between-total-variation-and-hellinger-distances-for-gaussian-traces
Agent traces
Agent sessions published from a Trackio Logbook.
indus-script-synthetic
Synthetic Indus Script Dataset
This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions.
Stage 1 — Train on real inscriptions:
Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.ro_hellaswag
Dataset Description
Hellaswag is a commonsense inference challenge dataset.
Here we provide the Romanian translation of the Hellaswag from the paper "Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback" (Lai et al., 2023).
This dataset is used as a benchmark and is part of the evaluation protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_hellaswag.cs_hellaswag
Introduction
This is an automatic translation of HellaSwag dataset in Czech.
The dataset was translated using LINDAT Translation Service available as an online-API.
Licensing Information
Please follow the HellaSwag repository for licensing information. CZLC members do not own the dataset, nor are they responsible for its contents.
hellaswaggmedical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Hellrabbit/medical-o1-reasoning-SFT.hellaswag_italian
HellaSwag - Italian (IT)
This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence.
Dataset Details
The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/hellaswag_italian.HellaSwag_HT_eu_sample
HellaSwag Human Translated Sample for Basque
A subset of 250 samples manually translated to Basque from the HellaSwag dataset (Zellers et al., 2019).
The corresponding 250 English samples are also provided.
The HellaSwag dataset is a dataset for commonsense NLI.
Dataset Creation
Source Data
A subset of 250 samples manually translated to Basque from the HellaSwag dataset (Zellers et al., 2019).
Annotations
Annotation process
A subset of 250… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/HellaSwag_HT_eu_sample.AraDICE-HellaSwag
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the HellaSwag split of the data.
Evaluation
We have used lm-harness eval framework to for… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-HellaSwag.hellaswagQwen3-32B_16episodes_comparisons_full
Qwen3-32B_16episodes_comparisons_full
This is a pairwise comparison dataset created from SWE-bench evaluation results.
Files
Qwen3-32B_16episodes_comparisons_full_comparison_pairs.jsonl: JSONL file containing comparison pairs
Metadata
{
"dataset_name": "Qwen3-32B_16episodes_comparisons_full",
"model_name": "Qwen3-32B",
"num_episodes": 16,
"episodes_used": [
1,
2,
3,
4,
5,
6,
7,
8,
9,
10,
11,
12,
13… See the full description on the dataset page: https://huggingface.co/datasets/helloelwin/Qwen3-32B_16episodes_comparisons_full.
