datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
natural_questions
Dataset Card for Natural Questions
Dataset Summary
The NQ corpus contains questions from real users, and it requires QA systems to
read and comprehend an entire Wikipedia article that may or may not contain the
answer to the question. The inclusion of real user questions, and the
requirement that solutions should read an entire page to find the answer, cause
NQ to be a more realistic and challenging task than prior QA datasets.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.researchy_questions_modified
Dataset Summary
This dataset is derived from corbyrosset/researchy_questions, a collection of ~100k non-factoid, multi-perspective "Researchy Questions" mined from real Bing search engine logs. After a labor-intensive filtering funnel from billions of queries, these "needles in the haystack" are questions that probably require a lot of sub-questions and research to answer adequately, and are shown to be harder than other open-domain QA datasets like Natural Questions.… See the full description on the dataset page: https://huggingface.co/datasets/minorproject-research/researchy_questions_modified.ml-research-questions-broad-llm-judgeResearch-BloomTaxonomy-LLM-Persian-QuestionsThis repository contains the code and datasets created and used for our paper:
"Evaluating LLM-Generated Persian Questions for Teaching Conditional Programming Using Bloom’s Taxonomy"
DOI: 10.1109/IST64061.2024.10843532
https://ieeexplore.ieee.org/document/10843532
Contributors:
Marzieh Alidadi (@marzieh-alidadi)
Narges Shahhoseini (@nshahhoseini)
ml-research-questions-llm-judgeresearch_questions
