CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01juletxara /xstory_cloze Dataset Card for XStoryCloze Dataset Summary XStoryCloze consists of the professionally translated version of the English StoryCloze dataset (Spring 2016 version) to 10 non-English languages. This dataset is released by Meta AI. Supported Tasks and Leaderboards commonsense reasoning Languages en, ru, zh (Simplified), es (Latin America), ar, hi, id, te, sw, eu, my. Dataset Structure Data Instances Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/xstory_cloze.textother10K<n<100K16 likes11k downloads1y agoHugging Face02MoE-UNC /story_clozetext1K<n<10K1 likes1.3k downloads3y agoHugging Face03gimmaru /story_cloze-2016 Dataset Card for "story_cloze-2016" More Information needed Note: This dataset was utilized for the evaluation of probability-based prompt selection techniques in the paper 'Improving Probability-based Prompt Selection Through Unified Evaluation and Analysis'. It differs from the actual benchmark dataset. text1K<n<10K1 likes1k downloads3y agoHugging Face04lecslab /story_clozetext1K<n<10K2 likes550 downloads2y agoHugging Face05EleutherAI /wmdp_bio_robust_clozetext1K<n<10K0 likes369 downloads11mo agoHugging Face06BoJack /Omni-Cloze Omni-Cloze Benchmark 📖 Paper | 🕵️ Omni-Detective Pipeline | 🧑‍🏫 Omni-Cloze Benchmark Guides Omni-Cloze frames detailed captioning evaluation as a cloze-style multiple-choice proxy task. Omni-Cloze is a unified benchmark for evaluating detailed captioning across audio-only, visual-only, and audio–visual settings. The dataset spans 9 main domains and 47 sub-categories covering diverse topics such as education, entertainment, sports, news, science, and lifestyle, with a… See the full description on the dataset page: https://huggingface.co/datasets/BoJack/Omni-Cloze.text1K<n<10K2 likes361 downloads6mo agoHugging Face07EleutherAI /wmdp_bio_clozetext1K<n<10K0 likes351 downloads1y agoHugging Face08sasha /wino_bias_cloze1textn<1K0 likes305 downloads4y agoHugging Face09google /code_x_glue_cc_cloze_testing_all Dataset Card for "code_x_glue_cc_cloze_testing_all" Dataset Summary CodeXGLUE ClozeTesting-all dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-all Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem. Here we… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_all.texttext-generation100K<n<1M6 likes300 downloads3y agoHugging Face10google /code_x_glue_cc_cloze_testing_maxmin Dataset Card for "code_x_glue_cc_cloze_testing_maxmin" Dataset Summary CodeXGLUE ClozeTesting-maxmin dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-maxmin Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_maxmin.texttext-generation1K<n<10K3 likes236 downloads3y agoHugging Face11Muennighoff /story_clozegatedtext10K<n<100K0 likes230 downloads4y agoHugging Face12qikp /rocstories-cloze ROCStories & Story Cloze Test Dataset This dataset contains the ROCStories corpus and the Story Cloze Test evaluation sets, originally released by the Story Cloze Test team at the University of Rochester. texttext-generation100K<n<1M0 likes187 downloads2mo agoHugging Face13portuguese-benchmark-datasets /story_cloze_pt Dataset Card for "story_cloze_pt" This is a portuguese translation of the xstory_cloze dataset. The translation was performed using the Google Translate API. This dataset follows the same structure as the original. text1K<n<10K1 likes186 downloads3y agoHugging Face14AlekseyKorshuk /lmeh-story-cloze-2016 Dataset Card for "lmeh-story-cloze-2016" More Information needed text1K<n<10K0 likes152 downloads3y agoHugging Face15sasha /wino_bias_cloze2textn<1K1 likes123 downloads4y agoHugging Face16Lots-of-LoRAs /task105_story_cloze-rocstories_sentence_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task105_story_cloze-rocstories_sentence_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task105_story_cloze-rocstories_sentence_generation.texttext-generation1K<n<10K0 likes104 downloads2y agoHugging Face17mikaberidze /xstory-cloze-ftp XStoryCloze-FTP A first-token-prediction (FTP) reframing of juletxara/xstory_cloze. Each example is a single text sequence ending in Answer: (fullwidth U+FF1A, no trailing space) so a model can predict the answer as one token. Format matches lm-evaluation-harness's MCQA template byte-for-byte. Format Example (en): <sentence 1> <sentence 2> <sentence 3> <sentence 4> A: <ending 1> B: <ending 2> Answer: Schema: question_id: int, text: str, answer_label: str (one of… See the full description on the dataset page: https://huggingface.co/datasets/mikaberidze/xstory-cloze-ftp.textmultiple-choice10K<n<100K1 likes84 downloads4mo agoHugging Face18hon9kon9ize /yue_xstory_cloze Dataset Card for Cantonese XStoryCloze This dataset is a Cantonese translation of the Simplified Chinese subset of juletxara/xstory_cloze. For more detailed information about the original dataset, please refer to the provided link. This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset. Sample { "input_sentence_1":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue_xstory_cloze.text1K<n<10K1 likes51 downloads3y agoHugging Face19elidek-themis /crows_pairs_clozeA subset of Crows-Pairs (total 1340). Dropped instances where sentence pairs differ by more than one word. Normalization fixes: - Punctuation correction (periods, commas) - Capitalization - Typos - Whitespace cleanup - Hyphenation standardization to support cloze pairs Source: https://gitlab.inria.fr/french-crows-pairs/acl-2022-paper-data-and-code @inproceedings{neveol-etal-2022-french, title = "{F}rench {C}row{S}-Pairs: Extending a challenge dataset for measuring social bias in… See the full description on the dataset page: https://huggingface.co/datasets/elidek-themis/crows_pairs_cloze.text1K<n<10K0 likes51 downloads8mo agoHugging Face20yuanxin112 /morphbench-verb-cloze MorphBench — contextual verb-inflection cloze (EN / DE / FR) A verb in a natural sentence is replaced by the marker [x]. The prompt gives the lemma and a partial feature bundle; the sentence supplies exactly the missing dimension. The model generates the surface form. prompt : cloze lemma=<L> context=<sentence with [x]> feats=<partial feats> -> gold : the inflected surface form The benchmark exists to test whether a morphology-aware tokenizer helps a small LM inflect words it… See the full description on the dataset page: https://huggingface.co/datasets/yuanxin112/morphbench-verb-cloze.tabulartext-generation10K<n<100K0 likes46 downloads1mo agoHugging Face21rifoag /javanese_sundanese_story_clozetext1K<n<10K0 likes37 downloads2y agoHugging Face22Joaoffg /Cloze-SSH SSH Cloze Benchmark A Cloze-style benchmark for evaluating language models on Social Sciences and Humanities (SSH) text understanding. The benchmark measures whether a model can choose between two equivalent candidate tokens (e.g. higher vs. lower, positive vs. negative) in the context of an academic abstract, where the correct choice requires domain knowledge rather than general English fluency. This dataset was introduced in the technical report SHARE: Social-Humanities AI for… See the full description on the dataset page: https://huggingface.co/datasets/Joaoffg/Cloze-SSH.texttext-generationn<1K1 likes37 downloads6mo agoHugging Face23CGU-Widelab /Cloze_QA_Dataset_Wikitext2 Cloze QA Dataset (WikiText-2) Dataset Description The Cloze QA Dataset is automatically generated from the WikiText-2 corpus. It contains fill-in-the-blank (cloze) style questions derived directly from sentences in Wikipedia articles. This dataset is particularly useful for evaluating local recall, reading comprehension, and contextual understanding. Each document produces exactly three unique QA pairs, preserving document structure and sentence alignment while… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Cloze_QA_Dataset_Wikitext2.textquestion-answering1 likes29 downloads2mo agoHugging Face24Kyle1668 /mmlu_auxiliary_train_formatted_clozetext10K<n<100K0 likes26 downloads1y agoHugging Face25turkhukuk /hukukbert-cloze-benchmark Turkish Legal Cloze Benchmark (JSONL) A small-scale Cloze-style multiple-choice benchmark for evaluating Turkish legal-domain language models. The dataset is designed to test whether models understand legal terminology, doctrinal structure, and domain-specific phrasing in Turkish law. Dataset Format Each example is stored as one JSON object per line (JSONL). Schema { "id": "string", "sentence": "string with [MASK] placeholder", "options": ["choice1"… See the full description on the dataset page: https://huggingface.co/datasets/turkhukuk/hukukbert-cloze-benchmark.textmultiple-choicen<1K2 likes26 downloads7mo agoHugging Face26Hplm /historical-clozetabular10K<n<100K0 likes25 downloads2y agoHugging Face27elidek-themis /wino_bias_clozetext1K<n<10K0 likes21 downloads1y agoHugging Face28EleutherAI /mmlu_auxiliary_train_formatted_cloze_20250619-1406text10K<n<100K0 likes19 downloads1y agoHugging Face29Hplm /historical-cloze-max-filteredtabular10K<n<100K0 likes18 downloads2y agoHugging Face30EleutherAI /mmlu_auxiliary_train_formatted_cloze_20250619-1417text10K<n<100K0 likes17 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.