CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B5 likes8.3k downloads2mo agoHugging Face02lightonai /embeddings-fine-tuning-multilingual-unfiltered Overview This dataset provides multilingual and code retrieval data for fine-tuning text embedding models. It is composed of high quality data sources with mined documents annotated with bi-encoder scores. For each query, the 2048 closest documents are mined with snowflake-arctic-embed-l-v2.0 for MIRACL and MLDR and with gte-modernbert-base for CodeEditSearchTrain, and annotated with their bi-encoder similarity score. No false-negative filtering or cross-encoder annotation is… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered.4 likes1.3k downloads2mo agoHugging Face03GlobalCampus /openalex-multilingual-embeddings OpenAlex Multilingual Embeddings This dataset contains multilingual text embeddings of all records in OpenAlex with a title or an abstract from the snapshot of 2023-10-20. The dataset was created for the FORAS project to investigate the efficacy of different methods of searching in databases of academic publications. All scripts will be available in a GitHub repository. The project is supported by a grant from the Dutch Research Council (grant no. 406.22.GO.048)… See the full description on the dataset page: https://huggingface.co/datasets/GlobalCampus/openalex-multilingual-embeddings.text100M<n<1B0 likes1.2k downloads3y agoHugging Face04jadenhoch /corpus_embeddings_multilingual-e5-large-instruct0 likes128 downloads9mo agoHugging Face05ronak-wani /Wiki2019-multilingual-dense-embeddings Wiki2019-multilingual-dense-embeddings Qdrant vector database containing dense embeddings for the wiki_composite collection, built from 2019 Wikipedia dumps across 8 languages. Collection Details Property Value Embedding model nvidia/llama-embed-nemotron-8b Qdrant version 1.17.0 Vector size 4096 Distance metric Cosine Storage mode On-disk (payload + vectors) Total size 947 GB (882 GiB) Payload index Keyword index on "wiki" field Chunking… See the full description on the dataset page: https://huggingface.co/datasets/ronak-wani/Wiki2019-multilingual-dense-embeddings.1B<n<10B0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.