datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AraDiCE-BoolQ
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the BoolQ split of the data.
Evaluation
We have used lm-harness eval framework to for the… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE-BoolQ.BoolQuestions
BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language?
Official repository for BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language?
GitHub Repository: https://github.com/zmzhang2000/boolean-dense-retrieval
HuggingFace Hub: https://huggingface.co/datasets/ustc-zhangzm/BoolQuestions
Paper: https://aclanthology.org/2024.findings-emnlp.156
BoolQuestions
BoolQuestions has been uploaded to Hugging Face Hub. You can download the… See the full description on the dataset page: https://huggingface.co/datasets/ustc-zhangzm/BoolQuestions.boolq-mk
BoolQ MK version
This dataset is a Macedonian adaptation of the BoolQ dataset, originally curated (English -> Serbian) by Aleksa Gordić. It was translated from Serbian to Macedonian using the Google Translate API.
You can find this dataset as part of the macedonian-llm-eval GitHub and HuggingFace.
The dataset can be used to evaluate the models described in the paper Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language.
Why Translate from… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/boolq-mk.
