KShivendu/miriad-mlateon-colbert-smoke
MIRIAD 200, encoded with mLateOn-medical Multi-vector (ColBERT-style) embeddings for tomaarsen/miriad-benchmark-200k, produced with multi-vector-encoder/mLateOn-medical. passages 200 token vectors 176,014 mean vectors / passage 880.07 dim 128 stored dtype float16 embeddings size 0.05 GB raw text encoded 1 MB The embeddings are 49x larger than the text they came from, which is why late-interaction retrieval needs quantization or pooling.… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/miriad-mlateon-colbert-smoke.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face