CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-mll /blimp Dataset Card for "blimp" Dataset Summary BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars. Supported Tasks and Leaderboards More Information Needed Languages More Information… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/blimp.texttext-classification10K<n<100K40 likes150k downloads3y agoHugging Face02WillHeld /blimp Dataset Card for "blimp" HuggingFace Hub Upload of BLiMP: The Benchmark of Linguistic Minimal Pairs from https://github.com/alexwarstadt/blimp If you use this dataset in your work, please cite the original authors and paper. @article{warstadt2020blimp, author = {Warstadt, Alex and Parrish, Alicia and Liu, Haokun and Mohananey, Anhad and Peng, Wei and Wang, Sheng-Fu and Bowman, Samuel R.}, title = {BLiMP: The Benchmark of Linguistic Minimal Pairs for English}, journal =… See the full description on the dataset page: https://huggingface.co/datasets/WillHeld/blimp.text10K<n<100K0 likes1.2k downloads4y agoHugging Face03juletxara /blimp-nl BLiMP-NL: Dutch BLIMP Dataset Description BLiMP-NL is a dataset of Dutch linguistic minimal pairs for evaluating language models' syntactic knowledge. It contains minimal pairs for 22 grammatical phenomena in Dutch, further divided into 84 paradigms. Dataset Structure The dataset contains minimal pairs of grammatical and ungrammatical sentences in Dutch, organized into subsets testing various linguistic phenomena. Each minimal pair tests a specific grammatical… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/blimp-nl.text1K<n<10K0 likes955 downloads1y agoHugging Face04liu-nlp /unimorph-blimpThis is a automatically corrupted, raw dataset that may contain many errors. More sophisticated and larger variants will be released soon. text10K<n<100K0 likes229 downloads6mo agoHugging Face05liu-nlp /blimp-single-errorDataset for probing model preferences for linguistically acceptable sentences. Generated by introducing automatic corruptions into sentences from Wikipedia, based on UniMorph minimal tag pairs. More info coming soon! @misc{glocker2025growmergescalingstrategies, title={Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation}, author={Kevin Glocker and Kätriin Kukk and Romina Oji and Marcel Bollmann and Marco Kuhlmann and Jenny Kunz}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/blimp-single-error.text100K<n<1M0 likes193 downloads10mo agoHugging Face06BabyLM-community /BabyLM-BLIMP-Filteredtext10K<n<100K0 likes175 downloads5mo agoHugging Face07liu-nlp /unimorph-blimp-200 This experimental BLiMP-style dataset contains morphological corruptions based on minimal tag pairs from UniMorph. A minimal tag pair is a pair differing in exactly one variable, e.g., N;DEF;GEN;PL --> N;DEF;NOM;PL. The source corpora are Wikipedia articles, most of them having some sort of quality tag tags (e.g., 'excellent articles'). The approach is inspired by MultiBLiMP, which also uses UniMorph to generate subject–verb agreement minimal pairs, but generalizes it to all possible tag… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/unimorph-blimp-200.text100K<n<1M0 likes170 downloads6mo agoHugging Face08liu-nlp /icelandic-blimp-single-errortext10K<n<100K0 likes157 downloads1y agoHugging Face09liu-nlp /unimorph-blimp-posvalidatedtext10K<n<100K0 likes152 downloads10mo agoHugging Face10liu-nlp /german-blimp-single-errortext100K<n<1M0 likes147 downloads1y agoHugging Face11Lots-of-LoRAs /task1559_blimp_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1559_blimp_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1559_blimp_binary_classification.texttext-generation1K<n<10K0 likes130 downloads2y agoHugging Face12liu-nlp /unimorph-blimp-growuptext10K<n<100K0 likes113 downloads5mo agoHugging Face13liu-nlp /swedish-blimp-single-errortext10K<n<100K0 likes98 downloads1y agoHugging Face14tasksource /blimptext10K<n<100K0 likes94 downloads6mo agoHugging Face15juletxara /lindsea-blimp LINDSEA BLIMP Dataset Description LINDSEA BLIMP is a dataset of Indonesian linguistic minimal pairs for evaluating language models' syntactic knowledge. The dataset is based on the BHASA project's Indonesian syntax data. Dataset Structure The dataset contains minimal pairs of grammatical and ungrammatical sentences in Indonesian, organized into separate subsets testing various linguistic phenomena: argument_structure (160 pairs): Tests word order, modals, and… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/lindsea-blimp.textn<1K0 likes68 downloads1y agoHugging Face16liu-nlp /estonian-blimp-single-errortext10K<n<100K0 likes60 downloads10mo agoHugging Face17ReliableAI /irish_blimpgatedtabular1K<n<10K3 likes54 downloads1y agoHugging Face18hyungjikim /blimp-with-hexatagstexttext-classification10K<n<100K0 likes48 downloads9mo agoHugging Face19PatrickHaller /blimp_synthtext10K<n<100K0 likes36 downloads1y agoHugging Face20liu-nlp /swedish-blimp-nouns-def-to-indeftext1K<n<10K0 likes32 downloads1y agoHugging Face21liu-nlp /faroese-blimp-verbs-3rd-to-2nd-pers-experimentaltext1K<n<10K0 likes31 downloads1y agoHugging Face22hriaz /blimp-hexatagged-incremental configs: text10K<n<100K0 likes29 downloads5mo agoHugging Face23liu-nlp /swedish-blimp-adjectives-neut_and_adverbs_to_masfemBLiMP-style dataset based on the following 45 articles from the Swedish Wikipedia: "Svenska", 'Fucking Åmål', 'Norrland', 'Nikola Tesla', 'Radiohead', 'Förbjudna staden', 'Danska fälttåget i Skåne och Blekinge 1709–1710', 'Gamla Göta landsväg', 'Regalskeppet Vasa', 'The Last of Us', 'H.C. Andersen', 'Medeltidens mat', 'Kända träd i Stockholm', 'Jupiters ringar', 'Tvättbjörn', 'Umeå stads kyrka', 'Örebro slott', 'Skogssamer', 'Norrlands nation', 'Euro', 'Fri vilja', 'Älvdalska', 'Bohus… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/swedish-blimp-adjectives-neut_and_adverbs_to_masfem.textn<1K0 likes24 downloads1y agoHugging Face24liu-nlp /faroese-blimp-nouns-dat-to-nom-experimentaltext1K<n<10K0 likes22 downloads1y agoHugging Face25liu-nlp /icelandic-blimp-nouns-def-to-indf-experimentaltext1K<n<10K0 likes19 downloads1y agoHugging Face26liu-nlp /german-blimp-verbs-sg-to-pl-experimentaltext1K<n<10K0 likes19 downloads1y agoHugging Face27Lots-of-LoRAs /task1560_blimp_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1560_blimp_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1560_blimp_binary_classification.texttext-generation1K<n<10K0 likes18 downloads2y agoHugging Face28liu-nlp /icelandic-blimp-verbs-ind-to-sbjv-experimentaltext1K<n<10K0 likes18 downloads1y agoHugging Face29liu-nlp /estonian-blimp-nom-sg-to-gen-sg-experimentaltext1K<n<10K0 likes18 downloads1y agoHugging Face30liu-nlp /swedish-blimp-adjectives-pl-to-sgBLiMP-style dataset based on the following 45 articles from the Swedish Wikipedia: "Svenska", 'Fucking Åmål', 'Norrland', 'Nikola Tesla', 'Radiohead', 'Förbjudna staden', 'Danska fälttåget i Skåne och Blekinge 1709–1710', 'Gamla Göta landsväg', 'Regalskeppet Vasa', 'The Last of Us', 'H.C. Andersen', 'Medeltidens mat', 'Kända träd i Stockholm', 'Jupiters ringar', 'Tvättbjörn', 'Umeå stads kyrka', 'Örebro slott', 'Skogssamer', 'Norrlands nation', 'Euro', 'Fri vilja', 'Älvdalska', 'Bohus… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/swedish-blimp-adjectives-pl-to-sg.text1K<n<10K0 likes17 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.