CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-mll /blimp Dataset Card for "blimp" Dataset Summary BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars. Supported Tasks and Leaderboards More Information Needed Languages More Information… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/blimp.texttext-classification10K<n<100K40 likes151k downloads3y agoHugging Face02WillHeld /blimp Dataset Card for "blimp" HuggingFace Hub Upload of BLiMP: The Benchmark of Linguistic Minimal Pairs from https://github.com/alexwarstadt/blimp If you use this dataset in your work, please cite the original authors and paper. @article{warstadt2020blimp, author = {Warstadt, Alex and Parrish, Alicia and Liu, Haokun and Mohananey, Anhad and Peng, Wei and Wang, Sheng-Fu and Bowman, Samuel R.}, title = {BLiMP: The Benchmark of Linguistic Minimal Pairs for English}, journal =… See the full description on the dataset page: https://huggingface.co/datasets/WillHeld/blimp.text10K<n<100K0 likes1.1k downloads4y agoHugging Face03juletxara /blimp-nl BLiMP-NL: Dutch BLIMP Dataset Description BLiMP-NL is a dataset of Dutch linguistic minimal pairs for evaluating language models' syntactic knowledge. It contains minimal pairs for 22 grammatical phenomena in Dutch, further divided into 84 paradigms. Dataset Structure The dataset contains minimal pairs of grammatical and ungrammatical sentences in Dutch, organized into subsets testing various linguistic phenomena. Each minimal pair tests a specific grammatical… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/blimp-nl.text1K<n<10K0 likes968 downloads1y agoHugging Face04liu-nlp /unimorph-blimpThis is a automatically corrupted, raw dataset that may contain many errors. More sophisticated and larger variants will be released soon. text10K<n<100K0 likes239 downloads6mo agoHugging Face05jmichaelov /blimp_nl BLiMP-NL: A Corpus of Dutch Minimal Pairs and Acceptability Judgments for Language Model Evaluation [A] corpus of 8400 Dutch sentence pairs, intended primarily for the grammatical evaluation of language models. Each pair consists of a grammatical sentence and a minimally different ungrammatical sentence. The corpus covers 84 paradigms, classified into 22 syntactic phenomena. Ten sentence pairs of each paradigm were created by hand, while the remaining 90 were generated… See the full description on the dataset page: https://huggingface.co/datasets/jmichaelov/blimp_nl.textmultiple-choice1K<n<10K0 likes235 downloads1y agoHugging Face06liu-nlp /blimp-single-errorDataset for probing model preferences for linguistically acceptable sentences. Generated by introducing automatic corruptions into sentences from Wikipedia, based on UniMorph minimal tag pairs. More info coming soon! @misc{glocker2025growmergescalingstrategies, title={Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation}, author={Kevin Glocker and Kätriin Kukk and Romina Oji and Marcel Bollmann and Marco Kuhlmann and Jenny Kunz}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/blimp-single-error.text100K<n<1M0 likes205 downloads10mo agoHugging Face07BabyLM-community /BabyLM-BLIMP-Filteredtext10K<n<100K0 likes186 downloads5mo agoHugging Face08liu-nlp /unimorph-blimp-200 This experimental BLiMP-style dataset contains morphological corruptions based on minimal tag pairs from UniMorph. A minimal tag pair is a pair differing in exactly one variable, e.g., N;DEF;GEN;PL --> N;DEF;NOM;PL. The source corpora are Wikipedia articles, most of them having some sort of quality tag tags (e.g., 'excellent articles'). The approach is inspired by MultiBLiMP, which also uses UniMorph to generate subject–verb agreement minimal pairs, but generalizes it to all possible tag… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/unimorph-blimp-200.text100K<n<1M0 likes178 downloads6mo agoHugging Face09liu-nlp /unimorph-blimp-posvalidatedtext10K<n<100K0 likes159 downloads10mo agoHugging Face10liu-nlp /icelandic-blimp-single-errortext10K<n<100K0 likes158 downloads1y agoHugging Face11liu-nlp /german-blimp-single-errortext100K<n<1M0 likes148 downloads1y agoHugging Face12Lots-of-LoRAs /task1559_blimp_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1559_blimp_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1559_blimp_binary_classification.texttext-generation1K<n<10K0 likes133 downloads2y agoHugging Face13bbunzeck /phoneme-blimptext10K<n<100K0 likes116 downloads2y agoHugging Face14liu-nlp /unimorph-blimp-growuptext10K<n<100K0 likes111 downloads5mo agoHugging Face15liu-nlp /swedish-blimp-single-errortext10K<n<100K0 likes98 downloads1y agoHugging Face16tasksource /blimptext10K<n<100K0 likes92 downloads6mo agoHugging Face17NeTSlab /BLiMP-IT BLiMP-IT Dataset Summary BLiMP-IT is a linguistically motivated benchmark for evaluating Italian language models through minimal pairs. Each example consists of a grammatical sentence paired with a minimally different ungrammatical counterpart that isolates a single morphosyntactic contrast. The benchmark is designed to evaluate whether language models assign higher probability to the grammatical sentence than to the ungrammatical one. The benchmark is inspired by… See the full description on the dataset page: https://huggingface.co/datasets/NeTSlab/BLiMP-IT.texttext-classification1K<n<10K0 likes69 downloads1mo agoHugging Face18juletxara /lindsea-blimp LINDSEA BLIMP Dataset Description LINDSEA BLIMP is a dataset of Indonesian linguistic minimal pairs for evaluating language models' syntactic knowledge. The dataset is based on the BHASA project's Indonesian syntax data. Dataset Structure The dataset contains minimal pairs of grammatical and ungrammatical sentences in Indonesian, organized into separate subsets testing various linguistic phenomena: argument_structure (160 pairs): Tests word order, modals, and… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/lindsea-blimp.textn<1K0 likes68 downloads1y agoHugging Face19liu-nlp /estonian-blimp-single-errortext10K<n<100K0 likes56 downloads10mo agoHugging Face20ReliableAI /irish_blimpgatedtabular1K<n<10K3 likes54 downloads11mo agoHugging Face21hyungjikim /blimp-with-hexatagstexttext-classification10K<n<100K0 likes47 downloads9mo agoHugging Face22PatrickHaller /blimp_synthtext10K<n<100K0 likes38 downloads1y agoHugging Face23liu-nlp /faroese-blimp-verbs-3rd-to-2nd-pers-experimentaltext1K<n<10K0 likes32 downloads1y agoHugging Face24liu-nlp /swedish-blimp-nouns-def-to-indeftext1K<n<10K0 likes31 downloads1y agoHugging Face25hriaz /blimp-hexatagged-incremental configs: text10K<n<100K0 likes30 downloads5mo agoHugging Face26Lots-of-LoRAs /task1560_blimp_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1560_blimp_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1560_blimp_binary_classification.texttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face27liu-nlp /swedish-blimp-adjectives-neut_and_adverbs_to_masfemBLiMP-style dataset based on the following 45 articles from the Swedish Wikipedia: "Svenska", 'Fucking Åmål', 'Norrland', 'Nikola Tesla', 'Radiohead', 'Förbjudna staden', 'Danska fälttåget i Skåne och Blekinge 1709–1710', 'Gamla Göta landsväg', 'Regalskeppet Vasa', 'The Last of Us', 'H.C. Andersen', 'Medeltidens mat', 'Kända träd i Stockholm', 'Jupiters ringar', 'Tvättbjörn', 'Umeå stads kyrka', 'Örebro slott', 'Skogssamer', 'Norrlands nation', 'Euro', 'Fri vilja', 'Älvdalska', 'Bohus… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/swedish-blimp-adjectives-neut_and_adverbs_to_masfem.textn<1K0 likes22 downloads1y agoHugging Face28liu-nlp /faroese-blimp-nouns-dat-to-nom-experimentaltext1K<n<10K0 likes22 downloads1y agoHugging Face29liu-nlp /german-blimp-verbs-sg-to-pl-experimentaltext1K<n<10K0 likes20 downloads1y agoHugging Face30liu-nlp /icelandic-blimp-nouns-def-to-indf-experimentaltext1K<n<10K0 likes19 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.