Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3
Vietnamese Evidence Corpus Embeddings Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v3, generated with intfloat/multilingual-e5-large. Rows: 53,114 Embedding dimension: 1024 Embedding dtype: float32 Input: `passage: {title} {text}` Inputs truncated to 512 tokens before embedding Deduplicated before embedding by normalized content_hash Provenance retained for duplicate content L2 normalized: yes Parquet shards: 11 Use query: for… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3.
Vietnamese Evidence Corpus Embeddings
Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v3, generated with intfloat/multilingual-e5-large.
- Rows: 53,114
- Embedding dimension: 1024
- Embedding dtype: float32
- Input: `passage: {title}
{text}`
- Inputs truncated to 512 tokens before embedding
- Deduplicated before embedding by normalized
content_hash - Provenance retained for duplicate content
- L2 normalized: yes
- Parquet shards: 11
Use query: for claims/queries and normalize query vectors before cosine or inner-product retrieval. The source chunk boundaries and provenance columns are kept for auditing, context assembly, and source tracing.
