datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.embeddings-fine-tuning-multilingual-unfiltered
Overview
This dataset provides multilingual and code retrieval data for fine-tuning text embedding models. It is composed of high quality data sources with mined documents annotated with bi-encoder scores.
For each query, the 2048 closest documents are mined with snowflake-arctic-embed-l-v2.0 for MIRACL and MLDR and with gte-modernbert-base for CodeEditSearchTrain, and annotated with their bi-encoder similarity score. No false-negative filtering or cross-encoder annotation is… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered.openalex-multilingual-embeddings
OpenAlex Multilingual Embeddings
This dataset contains multilingual text embeddings of all records in OpenAlex with a title or an abstract from the snapshot of 2023-10-20.
The dataset was created for the FORAS project to investigate the efficacy of
different methods of searching in databases of academic publications. All scripts will be available in a GitHub repository.
The project is supported by a grant from the Dutch Research Council (grant no. 406.22.GO.048)… See the full description on the dataset page: https://huggingface.co/datasets/GlobalCampus/openalex-multilingual-embeddings.corpus_embeddings_multilingual-e5-large-instructWiki2019-multilingual-dense-embeddings
Wiki2019-multilingual-dense-embeddings
Qdrant vector database containing dense embeddings for the wiki_composite collection, built from 2019 Wikipedia dumps across 8 languages.
Collection Details
Property
Value
Embedding model
nvidia/llama-embed-nemotron-8b
Qdrant version
1.17.0
Vector size
4096
Distance metric
Cosine
Storage mode
On-disk (payload + vectors)
Total size
947 GB (882 GiB)
Payload index
Keyword index on "wiki" field
Chunking… See the full description on the dataset page: https://huggingface.co/datasets/ronak-wani/Wiki2019-multilingual-dense-embeddings.
