CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QCRI /AraDiCE-BoolQ AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the BoolQ split of the data. Evaluation We have used lm-harness eval framework to for the… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE-BoolQ.textquestion-answering1K<n<10K0 likes158 downloads2y agoHugging Face02sarvamai /boolq-indic Indic BoolQ Dataset A multilingual version of the BoolQ (Boolean Questions) dataset, translated from English into 10 Indian languages. It is a question-answering dataset for yes/no questions containing ~12k naturally occurring questions. Languages Covered The dataset includes translations in the following languages: Bengali (bn) Gujarati (gu) Hindi (hi) Kannada (kn) Marathi (mr) Malayalam (ml) Oriya (or) Punjabi (pa) Tamil (ta) Telugu (te) Dataset Format Each… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/boolq-indic.textquestion-answering100K<n<1M0 likes102 downloads2y agoHugging Face03ustc-zhangzm /BoolQuestions BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language? Official repository for BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language? GitHub Repository: https://github.com/zmzhang2000/boolean-dense-retrieval HuggingFace Hub: https://huggingface.co/datasets/ustc-zhangzm/BoolQuestions Paper: https://aclanthology.org/2024.findings-emnlp.156 BoolQuestions BoolQuestions has been uploaded to Hugging Face Hub. You can download the… See the full description on the dataset page: https://huggingface.co/datasets/ustc-zhangzm/BoolQuestions.texttext-retrieval10M<n<100M3 likes87 downloads2y agoHugging Face04hishab /boolq_bn Dataset Summary BoolQ Bangla (BN) is a question-answering dataset for yes/no questions, generated using GPT-4. The dataset contains 15,942 examples, with each entry consisting of a triplet: (question, passage, answer). The questions are naturally occurring, generated from unprompted and unconstrained settings. Input passages were sourced from Bangla Wikipedia, Banglapedia, and News Articles, and GPT-4 was used to generate corresponding yes/no questions with answers. The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/hishab/boolq_bn.textquestion-answering1K<n<10K1 likes57 downloads1y agoHugging Face05ibm-research /BoolQ_robustness Dataset Card for "BoolQ-robustness" Dataset Summary BoolQ-robustness is an expanded version of the BoolQ dataset (https://arxiv.org/abs/1905.10044) but with perturbations of the original input questions and passages. It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations. Data Instances boolq_robustness Size of downloaded dataset file: 21.8 MB Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BoolQ_robustness.tabularquestion-answering10K<n<100K0 likes38 downloads2y agoHugging Face06LVSTCK /boolq-mk BoolQ MK version This dataset is a Macedonian adaptation of the BoolQ dataset, originally curated (English -> Serbian) by Aleksa Gordić. It was translated from Serbian to Macedonian using the Google Translate API. You can find this dataset as part of the macedonian-llm-eval GitHub and HuggingFace. The dataset can be used to evaluate the models described in the paper Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language. Why Translate from… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/boolq-mk.tabularquestion-answering1K<n<10K0 likes38 downloads1y agoHugging Face07jensjepsen /esperanto-boolq-questions esperanto-boolq-questions BoolQ questions (train + validation, 12,697 rows) translated from English to Esperanto by jensjepsen/eo-mt-v13-large-bidir, with round-trip quality metadata for filtering. Row schema field description orig_idx original BoolQ row index (train first, then validation) split source split (train / validation) en_orig raw BoolQ question (lowercase, no ?, as in google/boolq) en_preproc preprocessed input fed to MT: spaCy… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-boolq-questions.tabulartranslation10K<n<100K0 likes28 downloads2mo agoHugging Face08mygitphase /boolq-indic Indic BoolQ Dataset A multilingual version of the BoolQ (Boolean Questions) dataset, translated from English into 10 Indian languages. It is a question-answering dataset for yes/no questions containing ~12k naturally occurring questions. Languages Covered The dataset includes translations in the following languages: Bengali (bn) Gujarati (gu) Hindi (hi) Kannada (kn) Marathi (mr) Malayalam (ml) Oriya (or) Punjabi (pa) Tamil (ta) Telugu (te) Dataset Format Each… See the full description on the dataset page: https://huggingface.co/datasets/mygitphase/boolq-indic.textquestion-answering100K<n<1M0 likes24 downloads5mo agoHugging Face09mygitphase /boolq_indic Indic BoolQ Dataset A multilingual version of the BoolQ (Boolean Questions) dataset, translated from English into 10 Indian languages. It is a question-answering dataset for yes/no questions containing ~12k naturally occurring questions. Languages Covered The dataset includes translations in the following languages: Bengali (bn) Gujarati (gu) Hindi (hi) Kannada (kn) Marathi (mr) Malayalam (ml) Oriya (or) Punjabi (pa) Tamil (ta) Telugu (te) Dataset Format Each… See the full description on the dataset page: https://huggingface.co/datasets/mygitphase/boolq_indic.textquestion-answering100K<n<1M0 likes21 downloads5mo agoHugging Face10tbilisi-ai-lab /boolq-ka boolq-ka Georgian translation of the BoolQ (Boolean Questions) benchmark. Dataset Summary Property Value Examples 3,270 Splits validation Languages Georgian, English Task Yes/No Question Answering Data Fields question: Question (English) answer: Boolean answer passage: Context passage (English) question_ka: Question (Georgian) answer_ka: Answer (Georgian) passage_ka: Context passage (Georgian) Translation Methodology… See the full description on the dataset page: https://huggingface.co/datasets/tbilisi-ai-lab/boolq-ka.textquestion-answering1K<n<10K0 likes12 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.