datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cosmopedia-wikihow-chunked
Overview
This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face.
Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs.
Dataset Structure
Each record in the dataset represents a chunk of a larger article, and contains the following fields:
doc_id: A unique identifier for the parent article
chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.vietnamese-evidence-corpus-chunked-e5-v3
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
Chunked with multilingual-E5 token budget
Prefix-aware chunking using `passage: {title}
`
Sentence-aware overlap to preserve local context
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.vietnamese-evidence-corpus-chunked
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.vietnamese-evidence-corpus-chunked-e5-v2
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.
