datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nq_open_gold
Natural Questions Open Dataset with Gold Documents
This dataset is a curated version of the Natural Questions open dataset,
with the inclusion of the gold documents from the original Natural Questions (NQ) dataset.
The main difference with the NQ-open dataset is that some entries were excluded, as their respective gold documents exceeded 512 tokens in length.
This is due to the pre-processing of the gold documents, as detailed in this related dataset.
The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/nq_open_gold.kilt_corpus_wiki_dump2019
KILT Corpus in Chunks: Wikipedia Corpus Chunked for RAG
A preprocessed version of the KILT Wikipedia corpus where each article has been split into coherent, section-aware text chunks suitable for Retrieval Augmented Generation (RAG) pipelines.
Dataset Summary
This dataset provides a chunked view of the KILT Wikipedia snapshot (August 2019). Starting from full Wikipedia articles, each article is first segmented by its section structure, then each section is independently… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/kilt_corpus_wiki_dump2019.
