datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vietnamese-evidence-retrieval-indexes-v2-r1
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v2-r1 at revision 2a18d35b6ea2e078db95c1aacdc2a28947268b4e.
Rows: 63,699
Source embedding shards: 13
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Dense input: title + text
Dense rows: deduplicated by content hash
Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes-v2-r1.vietnamese-evidence-retrieval-indexes
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large at revision e928944361ca7d4c80f80d19bec52ebad55a4f7f.
Rows: 52,605
Source embedding shards: 11
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Vietnamese tokenization: Unicode word tokens, no stemming and no stopword removal
row_id in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes.Vietnamese-Openorca-Multiplechoice-gg-translatedvietnamese-evidence-corpus-chunked-e5-v3
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
Chunked with multilingual-E5 token budget
Prefix-aware chunking using `passage: {title}
`
Sentence-aware overlap to preserve local context
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.vietnamese-evidence-corpus-chunked
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.vietnamese-evidence-corpus-embeddings-e5-large-v2vietnamese-evidence-corpus-chunked-e5-v2
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.vietnamese_ultrafeedback_binarizedAIDetection_Vietnamese_HumanDatavietnamese-evidence-retrieval-indexes-v3-1
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
aiMy144/vietnamese-evidence-corpus-embeddings-e5-large-v3-1 at revision ea0826c0a9d4273eec267eb66b6dbbfe64d5fb11.
Rows: 53,747
Source embedding shards: 11
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Dense input: title + text
Dense rows: deduplicated by content hash
Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/aiMy144/vietnamese-evidence-retrieval-indexes-v3-1.
