datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
boolq
Dataset Card for Boolq
Dataset Summary
BoolQ is a question answering dataset for yes/no questions containing 15942 examples. These questions are naturally
occurring ---they are generated in unprompted and unconstrained settings.
Each example is a triplet of (question, passage, answer), with the title of the page as optional additional context.
The text-pair classification setup is similar to existing natural language inference tasks.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/google/boolq.boolq_helmllama2_7b_chat-boolq-results
Dataset Card for "llama2_7b_chat-boolq-results"
More Information needed
boolq_n_shotboolq_nqboolq-audio
Dataset Card for Dataset Name
This is a derivative of https://huggingface.co/datasets/google/boolq, but with an audio version of the questions as an additional feature. The audio was generated by running the existing question values through the Azure TTS generator with a 16KHz sample rate.
Dataset Details
Dataset Description
Curated by: Fixie.ai
Language(s) (NLP): English
License: Creative Commons Share-Alike 3.0 license.
Uses
Training and… See the full description on the dataset page: https://huggingface.co/datasets/fixie-ai/boolq-audio.TextToText_boolqboolq-qwen3-vl-32bboolq-translatedfrench_boolq
Dataset Card for "test_fboolq"
More Information needed
task380_boolq_yes_no_question
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task380_boolq_yes_no_question
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task380_boolq_yes_no_question.boolq-indic
Indic BoolQ Dataset
A multilingual version of the BoolQ (Boolean Questions) dataset, translated from English into 10 Indian languages.
It is a question-answering dataset for yes/no questions containing ~12k naturally occurring questions.
Languages Covered
The dataset includes translations in the following languages:
Bengali (bn)
Gujarati (gu)
Hindi (hi)
Kannada (kn)
Marathi (mr)
Malayalam (ml)
Oriya (or)
Punjabi (pa)
Tamil (ta)
Telugu (te)
Dataset Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/boolq-indic.boolq-cot-opus5kor_boolq
Dataset Card for "kor_boolq"
More Information needed
Source Data Citation Information
@inproceedings{clark2019boolq,
title = {BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions},
author = {Clark, Christopher and Lee, Kenton and Chang, Ming-Wei, and Kwiatkowski, Tom and Collins, Michael, and Toutanova, Kristina},
booktitle = {NAACL},
year = {2019},
}
super_glue_boolq_promptsourceboolq_msmarcoboolq
Dataset Card for "boolq"
More Information needed
snli_boolq
Dataset Card for "snli_boolq"
More Information needed
boolq_kg
BoolQ (Kyrgyz)
This dataset is the Kyrgyz-translated version of the BoolQ benchmark, a reading comprehension task requiring a yes/no answer.
🏔️ Part of the KyrgyzLLM-Bench
This dataset is a component of the KyrgyzLLM-Bench, a comprehensive suite for evaluating LLMs in Kyrgyz.
Main Paper: Bridging the Gap in Less-Resourced Languages: Building a Benchmark for Kyrgyz Language Models
Hugging Face Hub: https://huggingface.co/TTimur
GitHub Project:… See the full description on the dataset page: https://huggingface.co/datasets/TTimur/boolq_kg.boolq_bn
Dataset Summary
BoolQ Bangla (BN) is a question-answering dataset for yes/no questions, generated using GPT-4. The dataset contains 15,942 examples, with each entry consisting of a triplet: (question, passage, answer). The questions are naturally occurring, generated from unprompted and unconstrained settings. Input passages were sourced from Bangla Wikipedia, Banglapedia, and News Articles, and GPT-4 was used to generate corresponding yes/no questions with answers.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/hishab/boolq_bn.boolq_logsnli_boolq_train
Dataset Card for "snli_boolq_train"
This dataset contains validation and training data from both boolq and snli
snli_boolq_train is a mixed dataset containing both boolq (https://huggingface.co/datasets/boolq) and the "entail" and "contradict" samples from snli (https://huggingface.co/datasets/snli). The selection of this data was to increase the finetuning dataset amount for large language models with uninitialized weights. neutral statements were removed to… See the full description on the dataset page: https://huggingface.co/datasets/wojemann/snli_boolq_train.boolq_with_dev_hpotars_boolq
Dataset Card for "tars_boolq"
More Information needed
boolq_ragsoluted_qwentask381_boolq_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task381_boolq_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task381_boolq_question_generation.phi-boolq-results
Dataset Card for "phi-boolq-results"
More Information needed
llama2_7b_chat-boolq
Dataset Card for "llama2_7b_chat-boolq"
More Information needed
boolq_ar
Dataset Card for "boolq_ar"
More Information needed
boolq_reformatted
