CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MongoDB /cosmopedia-wikihow-chunked Overview This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face. Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs. Dataset Structure Each record in the dataset represents a chunk of a larger article, and contains the following fields: doc_id: A unique identifier for the parent article chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.tabularquestion-answering1M<n<10M9 likes208 downloads3y agoHugging Face02Loctran123 /vietnamese-evidence-corpus-chunked-e5-v3 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics Chunked with multilingual-E5 token budget Prefix-aware chunking using `passage: {title} ` Sentence-aware overlap to preserve local context Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.tabulartext-retrieval10K<n<100K0 likes39 downloads1mo agoHugging Face03Loctran123 /vietnamese-evidence-corpus-chunked Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.tabulartext-retrieval10K<n<100K0 likes23 downloads1mo agoHugging Face04Loctran123 /vietnamese-evidence-corpus-chunked-e5-v2 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.tabulartext-retrieval10K<n<100K0 likes20 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.