Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3
Vietnamese Evidence Corpus Embeddings Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v3, generated with intfloat/multilingual-e5-large. Rows: 53,114 Embedding dimension: 1024 Embedding dtype: float32 Input: `passage: {title} {text}` Inputs truncated to 512 tokens before embedding Deduplicated before embedding by normalized content_hash Provenance retained for duplicate content L2 normalized: yes Parquet shards: 11 Use query: for… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face