datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2
Dataset Card for "wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2"
More Information needed
msmarco-msmarco-MiniLM-L6-v3
MS MARCO with hard negatives from msmarco-MiniLM-L6-v3
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-MiniLM-L6-v3.msmarco-scores-ms-marco-MiniLM-L6-v2
MS MARCO query-passage scores using cross-encoder/ms-marco-MiniLM-L6-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
This dataset contains 160 million CrossEncoder scores on the MS MARCO dataset, using the cross-encoder/ms-marco-MiniLM-L6-v2 model.
The scores are unprocessed logits, i.e. they don't range between 0...1, and they can be used for finetuning search models using distillation.
See… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-scores-ms-marco-MiniLM-L6-v2.malicious-prompts-minilm-embeddingsfineweb-multilingual-minilm-l12-v2-shard-27468-20000STEM-wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2msmarco-hard-negatives-cross-encoder-ms-marco-MiniLM-L-6-v2-scoresnyaya-ae-all-MiniLM-L6-v2-ftlegal-v2
Dataset Card for "nyaya-ae-all-MiniLM-L6-v2-ftlegal-v2"
More Information needed
yemeksepeti-MiniLM-L12-v2nyaya-ae-all-MiniLM-L6-v2
Dataset Card for "nyaya-ae-all-MiniLM-L6-v2"
More Information needed
msmarco-margin-mse-minilmfineweb-multilingual-minilm-l12-v2-shard-27468-27460philippine-budget-2025-embeddings-minilm
Philippine Budget 2025 - Vector Embeddings (all-MiniLM-L6-v2)
Dataset Description
This dataset contains vector embeddings of the 2025 People's Budget of the Philippines, a citizen-friendly overview of the PHP 6.326 trillion national budget published by the Department of Budget and Management (DBM).
Source Document
These embeddings are based on the 2025 People's Enacted Budget (English version, revised as of April 22, 2025).
Direct Download Link: 2025 People's… See the full description on the dataset page: https://huggingface.co/datasets/pageman/philippine-budget-2025-embeddings-minilm.s1K-1.1-minilm-split-kmeans-dim384-20251118nyaya-ae-all-MiniLM-L6-v2-ftlegal-v1
Dataset Card for "nyaya-ae-all-MiniLM-L6-v2-ftlegal-v1"
More Information needed
technetcorpus-all-MiniLM-L6-v2mimarchive-all-MiniLM-L6-v2TheValley_embeddings_all-MiniLM-L6-v2MiniLMReduced
