datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blimp
Dataset Card for "blimp"
Dataset Summary
BLiMP is a challenge set for evaluating what language models (LMs) know about
major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each
containing 1000 minimal pairs isolating specific contrasts in syntax,
morphology, or semantics. The data is automatically generated according to
expert-crafted grammars.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/blimp.blimp
Dataset Card for "blimp"
HuggingFace Hub Upload of BLiMP: The Benchmark of Linguistic Minimal Pairs from https://github.com/alexwarstadt/blimp
If you use this dataset in your work, please cite the original authors and paper.
@article{warstadt2020blimp,
author = {Warstadt, Alex and Parrish, Alicia and Liu, Haokun and Mohananey, Anhad and Peng, Wei and Wang, Sheng-Fu and Bowman, Samuel R.},
title = {BLiMP: The Benchmark of Linguistic Minimal Pairs for English},
journal =… See the full description on the dataset page: https://huggingface.co/datasets/WillHeld/blimp.blimp-nl
BLiMP-NL: Dutch BLIMP
Dataset Description
BLiMP-NL is a dataset of Dutch linguistic minimal pairs for evaluating language models' syntactic knowledge. It contains minimal pairs for 22 grammatical phenomena in Dutch, further divided into 84 paradigms.
Dataset Structure
The dataset contains minimal pairs of grammatical and ungrammatical sentences in Dutch, organized into subsets testing various linguistic phenomena. Each minimal pair tests a specific grammatical… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/blimp-nl.unimorph-blimpThis is a automatically corrupted, raw dataset that may contain many errors. More sophisticated and larger variants will be released soon.
blimp_nl
BLiMP-NL: A Corpus of Dutch Minimal Pairs and Acceptability Judgments for Language Model Evaluation
[A] corpus of 8400 Dutch sentence pairs, intended primarily for the grammatical evaluation of language models. Each pair consists of a grammatical sentence and a minimally different ungrammatical sentence. The corpus covers 84 paradigms, classified into 22 syntactic phenomena. Ten sentence pairs of each paradigm were created by hand, while the remaining 90 were generated… See the full description on the dataset page: https://huggingface.co/datasets/jmichaelov/blimp_nl.blimp-single-errorDataset for probing model preferences for linguistically acceptable sentences.
Generated by introducing automatic corruptions into sentences from Wikipedia, based on UniMorph minimal tag pairs.
More info coming soon!
@misc{glocker2025growmergescalingstrategies,
title={Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation},
author={Kevin Glocker and Kätriin Kukk and Romina Oji and Marcel Bollmann and Marco Kuhlmann and Jenny Kunz},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/blimp-single-error.BabyLM-BLIMP-Filteredunimorph-blimp-200
This experimental BLiMP-style dataset contains morphological corruptions based on minimal tag pairs from UniMorph.
A minimal tag pair is a pair differing in exactly one variable, e.g., N;DEF;GEN;PL --> N;DEF;NOM;PL.
The source corpora are Wikipedia articles, most of them having some sort of quality tag tags (e.g., 'excellent articles').
The approach is inspired by MultiBLiMP, which also uses UniMorph to generate subject–verb agreement minimal pairs, but generalizes it to all possible tag… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/unimorph-blimp-200.unimorph-blimp-posvalidatedicelandic-blimp-single-errorgerman-blimp-single-errortask1559_blimp_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1559_blimp_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1559_blimp_binary_classification.phoneme-blimpunimorph-blimp-growupswedish-blimp-single-errorblimpBLiMP-IT
BLiMP-IT
Dataset Summary
BLiMP-IT is a linguistically motivated benchmark for evaluating Italian language models through minimal pairs.
Each example consists of a grammatical sentence paired with a minimally different ungrammatical counterpart that isolates a single morphosyntactic contrast.
The benchmark is designed to evaluate whether language models assign higher probability to the grammatical sentence than to the ungrammatical one.
The benchmark is inspired by… See the full description on the dataset page: https://huggingface.co/datasets/NeTSlab/BLiMP-IT.lindsea-blimp
LINDSEA BLIMP
Dataset Description
LINDSEA BLIMP is a dataset of Indonesian linguistic minimal pairs for evaluating language models' syntactic knowledge. The dataset is based on the BHASA project's Indonesian syntax data.
Dataset Structure
The dataset contains minimal pairs of grammatical and ungrammatical sentences in Indonesian, organized into separate subsets testing various linguistic phenomena:
argument_structure (160 pairs): Tests word order, modals, and… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/lindsea-blimp.estonian-blimp-single-errorirish_blimpblimp-with-hexatagsblimp_synthfaroese-blimp-verbs-3rd-to-2nd-pers-experimentalswedish-blimp-nouns-def-to-indefblimp-hexatagged-incremental
configs:
task1560_blimp_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1560_blimp_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1560_blimp_binary_classification.swedish-blimp-adjectives-neut_and_adverbs_to_masfemBLiMP-style dataset based on the following 45 articles from the Swedish Wikipedia: "Svenska", 'Fucking Åmål', 'Norrland', 'Nikola Tesla', 'Radiohead', 'Förbjudna staden', 'Danska fälttåget i Skåne och Blekinge 1709–1710', 'Gamla Göta landsväg', 'Regalskeppet Vasa', 'The Last of Us', 'H.C. Andersen', 'Medeltidens mat', 'Kända träd i Stockholm', 'Jupiters ringar', 'Tvättbjörn', 'Umeå stads kyrka', 'Örebro slott', 'Skogssamer', 'Norrlands nation', 'Euro', 'Fri vilja', 'Älvdalska', 'Bohus… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/swedish-blimp-adjectives-neut_and_adverbs_to_masfem.faroese-blimp-nouns-dat-to-nom-experimentalgerman-blimp-verbs-sg-to-pl-experimentalicelandic-blimp-nouns-def-to-indf-experimental
