CoolFace
Datasetpublic

Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3

Vietnamese Evidence Corpus Embeddings Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v3, generated with intfloat/multilingual-e5-large. Rows: 53,114 Embedding dimension: 1024 Embedding dtype: float32 Input: `passage: {title} {text}` Inputs truncated to 512 tokens before embedding Deduplicated before embedding by normalized content_hash Provenance retained for duplicate content L2 normalized: yes Parquet shards: 11 Use query: for… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes24downloads
15 commits on main
6ddedd91mo ago

Finalize E5 embeddings: manifest.json

Loctran123
4dc37171mo ago

Finalize E5 embeddings: README.md

Loctran123
5dd1c4a1mo ago

Add embedding shard train-00010

Loctran123
5deaebd1mo ago

Add embedding shard train-00009

Loctran123
8a884d21mo ago

Add embedding shard train-00008

Loctran123
947071c1mo ago

Add embedding shard train-00007

Loctran123
be88cd81mo ago

Add embedding shard train-00006

Loctran123
54c27a91mo ago

Add embedding shard train-00005

Loctran123
79b35041mo ago

Add embedding shard train-00004

Loctran123
678ff4a1mo ago

Add embedding shard train-00003

Loctran123
56d4aa51mo ago

Add embedding shard train-00002

Loctran123
ba3d7b11mo ago

Add embedding shard train-00001

Loctran123
36443011mo ago

Add embedding shard train-00000

Loctran123
08bc5df1mo ago

Add E5 embedding configuration

Loctran123
56b58711mo ago

initial commit

Loctran123