CoolFace
Datasetpublic

Loctran123/vietnamese-evidence-corpus-chunked-e5-v2

Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes6downloads
settings

This repository belongs to Loctran123 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namevietnamese-evidence-corpus-chunked-e5-v2
visibilitypublic
licenceother
gatedno
ownerLoctran123
Account settings
Loctran123/vietnamese-evidence-corpus-chunked-e5-v2 · CoolFace