CoolFace
Datasetpublic

Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3

Vietnamese Evidence Corpus Embeddings Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v3, generated with intfloat/multilingual-e5-large. Rows: 53,114 Embedding dimension: 1024 Embedding dtype: float32 Input: `passage: {title} {text}` Inputs truncated to 512 tokens before embedding Deduplicated before embedding by normalized content_hash Provenance retained for duplicate content L2 normalized: yes Parquet shards: 11 Use query: for… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes24downloads
Dataset Card

Vietnamese Evidence Corpus Embeddings

Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v3, generated with intfloat/multilingual-e5-large.

  • —Rows: 53,114
  • —Embedding dimension: 1024
  • —Embedding dtype: float32
  • —Input: `passage: {title}

{text}`

  • —Inputs truncated to 512 tokens before embedding
  • —Deduplicated before embedding by normalized content_hash
  • —Provenance retained for duplicate content
  • —L2 normalized: yes
  • —Parquet shards: 11

Use query: for claims/queries and normalize query vectors before cosine or inner-product retrieval. The source chunk boundaries and provenance columns are kept for auditing, context assembly, and source tracing.