CoolFace
Datasetpublic

KShivendu/miriad-mlateon-colbert-smoke

MIRIAD 200, encoded with mLateOn-medical Multi-vector (ColBERT-style) embeddings for tomaarsen/miriad-benchmark-200k, produced with multi-vector-encoder/mLateOn-medical. passages 200 token vectors 176,014 mean vectors / passage 880.07 dim 128 stored dtype float16 embeddings size 0.05 GB raw text encoded 1 MB The embeddings are 49x larger than the text they came from, which is why late-interaction retrieval needs quantization or pooling.… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/miriad-mlateon-colbert-smoke.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes55downloads
8 commits on main
c2a1cef1mo ago

Upload meta.json with huggingface_hub

KShivendu
f4f2d861mo ago

Upload query_ids.json with huggingface_hub

KShivendu
e1a72661mo ago

Upload qrels.json with huggingface_hub

KShivendu
5d0a5e01mo ago

Upload queries_offs.npy with huggingface_hub

KShivendu
6e4fb041mo ago

Upload queries.npy with huggingface_hub

KShivendu
3bf3cc71mo ago

Upload README.md with huggingface_hub

KShivendu
e7686f31mo ago

Upload data/train-00000-of-00001.parquet with huggingface_hub

KShivendu
71ba5341mo ago

initial commit

KShivendu