datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ko-widesearch
Ko-WideSearch
A Korean breadth-search benchmark: each task asks a web agent to exhaustively
enumerate a closed set and fill every attribute cell of a table (e.g. "list every
award category at the 59th Grand Bell Awards and give each winner"). 228 tasks across
three difficulty tiers.
[!IMPORTANT]
The question and answer fields are encrypted. To keep this a fair,
leakage-aware test of web agents, the gold is not published as plain text — it is
canary-XOR obfuscated (same scheme… See the full description on the dataset page: https://huggingface.co/datasets/Minbyul/Ko-widesearch.Mistral_Trivia-QA_Dataset
Mistral Trivia QA Dataset
The Mistral Trivia QA Dataset is a collection of trivia questions and answers designed to evaluate and train question-answering models. It covers a wide range of topics and is particularly useful for assessing a model's ability to handle general knowledge and reasoning tasks.The documents are derived from WikiText-2, providing diverse and well-structured textual content suitable for extractive QA generation.
Model outputs for this dataset were generated… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Mistral_Trivia-QA_Dataset.Cloze_QA_Dataset_Wikitext2
Cloze QA Dataset (WikiText-2)
Dataset Description
The Cloze QA Dataset is automatically generated from the WikiText-2 corpus. It contains fill-in-the-blank (cloze) style questions derived directly from sentences in Wikipedia articles. This dataset is particularly useful for evaluating local recall, reading comprehension, and contextual understanding.
Each document produces exactly three unique QA pairs, preserving document structure and sentence alignment while… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Cloze_QA_Dataset_Wikitext2.
