CoolFace
Datasetpublic

Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v2-r1

Vietnamese Evidence Corpus Embeddings Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v2, generated with intfloat/multilingual-e5-large. Rows: 63,873 Embedding dimension: 1024 Embedding dtype: float32 Prefix: passage: L2 normalized: yes Parquet shards: 13 Use query: for claims/queries and normalize query vectors before cosine or inner-product retrieval. The chunk_id column is the stable join key back to the source corpus.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes10downloads
Dataset Card

Vietnamese Evidence Corpus Embeddings

Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v2, generated with intfloat/multilingual-e5-large.

  • —Rows: 63,873
  • —Embedding dimension: 1024
  • —Embedding dtype: float32
  • —Prefix: passage:
  • —L2 normalized: yes
  • —Parquet shards: 13

Use query: for claims/queries and normalize query vectors before cosine or inner-product retrieval. The chunk_id column is the stable join key back to the source corpus.