datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DebateSum
DebateSum
Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset"
Arxiv pre-print available here: https://arxiv.org/abs/2011.07251
Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9
Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/
Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/helloerikaaa/cbis-ddsm-r.HellaSwag_TH
Dataset Card for HellaSwag_TH
Dataset Description
The cleaned version is available at Patt/HellaSwag_TH_cleanned.
This dataset is Thai translated version of hellaswag using google translate with Multilingual Universal Sentence Encoder to calculate score for Thai translation.
Languages
EN
TH
Citation
@misc{HellaSwag_TH,
author = {Triamamornwooth Patteera},
title = {HellaSwag_TH},
year = {2023},
publisher = {Hugging Face}… See the full description on the dataset page: https://huggingface.co/datasets/Patt/HellaSwag_TH.DynaMem-DynaBenchhellomtkinit_hellothere
mtkinit/hellothere
Created from AIOD platform
helloworld_datasetHellaSwag_PT-PT
Dataset Card for HellaSwag_PT-PT
Dataset Summary
This repository provides a European Portuguese (pt-PT) translation of HellaSwag, an adversarial benchmark for grounded commonsense inference. Given a short scenario, the model must choose the most plausible continuation from four options.Load with:
import datasets
data = datasets.load_dataset("ruibrogandrade/HellaSwag_PT-PT")
Supported Tasks and Leaderboards
Multiple-Choice Question Answering (commonsense… See the full description on the dataset page: https://huggingface.co/datasets/ruibrogandrade/HellaSwag_PT-PT.hello_worldhelloworld_datasethellaswag_trhellaswag-fiaivivn_testHello-world-21312
Hello-world-21312
Created from AIOD platform
helloWorld_datasethelloFriendDatasethellaswag-fi-google-translatepermuted-letter-string-analogies
Permuted Letter-String Analogies
This repository contains datasets introduced in Hellwig et al. (2026). The datasets are an extension of the letter-string analogies introduced in Lewis & Mitchell (2025).
Each folder contains a dataset with different data attributes, and contains a train, validation, and test set.
The naming convention is:
all_transformations_<copy>_study<N>_perm<N>
all_transformations:
Below are illustrations for each transformation on the standard alphabet.… See the full description on the dataset page: https://huggingface.co/datasets/philipp-hellwig/permuted-letter-string-analogies.test_helloworldDatasetvlsp_testen_Swahili_datasetsocial_pretrain_data
