datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xstory_cloze
Dataset Card for XStoryCloze
Dataset Summary
XStoryCloze consists of the professionally translated version of the English StoryCloze dataset (Spring 2016 version) to 10 non-English languages. This dataset is released by Meta AI.
Supported Tasks and Leaderboards
commonsense reasoning
Languages
en, ru, zh (Simplified), es (Latin America), ar, hi, id, te, sw, eu, my.
Dataset Structure
Data Instances
Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/xstory_cloze.story_clozeCloze-SSH
SSH Cloze Benchmark
A Cloze-style benchmark for evaluating language models on Social Sciences and Humanities (SSH) text understanding. The benchmark measures whether a model can choose between two equivalent candidate tokens (e.g. higher vs. lower, positive vs. negative) in the context of an academic abstract, where the correct choice requires domain knowledge rather than general English fluency.
This dataset was introduced in the technical report SHARE: Social-Humanities AI for… See the full description on the dataset page: https://huggingface.co/datasets/Joaoffg/Cloze-SSH.KI_simple_cloze_for_finetuning_LLMs
