trec-ragtime/ragtime2
RAGTIME2 Collection This dataset contains the documents for TREC RAGTIME Track 2026. Please refer to the website for the details of the task. RAGTIME is a multilingual RAG task, which expects the participating system to retrieve relevant documents from all four languages and synthesize a response with citation to the report request. For convenience, we separate the documents by their languages into four .jsonl files. However, they are intended to be used as a whole set. The… See the full description on the dataset page: https://huggingface.co/datasets/trec-ragtime/ragtime2.
RAGTIME2 Collection
This dataset contains the documents for TREC RAGTIME Track 2026. Please refer to the website for the details of the task.
RAGTIME is a multilingual RAG task, which expects the participating system to retrieve relevant documents from all four languages and synthesize a response with citation to the report request. For convenience, we separate the documents by their languages into four .jsonl files. However, they are intended to be used as a whole set.
The documents are extracted from Common Crawl News and sampled between August 1, 2021, and July 31, 2024, with an even number of documents every day. Each language has 1,000,095 documents. Machine translation using a Sockeye model trained by HLTCOE are also released.
