juletxara/lindsea-blimp
LINDSEA BLIMP Dataset Description LINDSEA BLIMP is a dataset of Indonesian linguistic minimal pairs for evaluating language models' syntactic knowledge. The dataset is based on the BHASA project's Indonesian syntax data. Dataset Structure The dataset contains minimal pairs of grammatical and ungrammatical sentences in Indonesian, organized into separate subsets testing various linguistic phenomena: argument_structure (160 pairs): Tests word order… See the full description on the dataset page: https://huggingface.co/datasets/juletxara/lindsea-blimp.
068
