Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3
Vietnamese Evidence Corpus Embeddings Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v3, generated with intfloat/multilingual-e5-large. Rows: 53,114 Embedding dimension: 1024 Embedding dtype: float32 Input: `passage: {title} {text}` Inputs truncated to 512 tokens before embedding Deduplicated before embedding by normalized content_hash Provenance retained for duplicate content L2 normalized: yes Parquet shards: 11 Use query: for… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3.
Finalize E5 embeddings: manifest.json
Finalize E5 embeddings: README.md
Add embedding shard train-00010
Add embedding shard train-00009
Add embedding shard train-00008
Add embedding shard train-00007
Add embedding shard train-00006
Add embedding shard train-00005
Add embedding shard train-00004
Add embedding shard train-00003
Add embedding shard train-00002
Add embedding shard train-00001
Add embedding shard train-00000
Add E5 embedding configuration
initial commit
