CoolFace
10 results

embedding-training

sentence-transformers /embedding-training-data Training Data for Text Embedding Models [!NOTE] This repository contains raw datasets, all of which have also been formatted for easy training in the Embedding Model Datasets collection. We recommend looking there first. This repository contains training files to train text embedding models, e.g. using sentence-transformers. Data Format All files are in a jsonl.gz format: Each line contains a JSON-object that represent one training example. The JSON objects can… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/embedding-training-data.feature-extraction144 likes2.1k downloads27d agoHugging FaceKaLM-Embedding /KaLM-reranker-training-data Lychee-KaLM-Reranker Training Data A large-scale, ready-to-use multilingual dataset for fine-tuning reranking models. This repository contains 3,885,265 training samples collected from 54 datasets, covering English, Chinese, and multilingual retrieval tasks. Each sample includes task instructions, positive passages, at least 16 hard negatives, and teacher scores annotated by Qwen3-Reranker-8B. When expanded into point-wise query–passage pairs, the dataset provides at least 66… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-reranker-training-data.texttext-ranking1M<n<10M3 likes1k downloads18d agoHugging Faceselmanbaysan /turkish_embedding_model_training_datatextsentence-similarity100M<n<1B5 likes1k downloads1y agoHugging Faceselmanbaysan /cleaned_turkish_embedding_model_training_data_colabtext10M<n<100M1 likes481 downloads1y agoHugging Facetrmteb /cleaned_turkish_embedding_model_training_data_colab Citation If you use this dataset in your research, please cite the following paper: @inproceedings{baysan-gungor-2025-tr, title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations", author = "Baysan, Mehmet Selman and Gungor, Tunga", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025", month = nov, year = "2025", address = "Suzhou, China", publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.text10M<n<100M3 likes413 downloads10mo agoHugging FaceNorskHelsenett /eti-embedding-training-data-2048-v3 Version note (v3): third generation of the ETI dense training data. Reuses the questions of NorskHelsenett/eti-embedding-training-data-v2 (which trained eti-embeddinggemma-v2), but re-maps every question to its article URL, re-cuts positives at a 2048-token window, re-mines hard negatives with granite + reranker validation (pos_score/neg_score), adds the keyword half, and filters junk anchors. Trained NorskHelsenett/eti-embeddinggemma-v3. eti-embedding-training-data-2048… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-v3.tabularsentence-similarity100K<n<1M0 likes218 downloads25d agoHugging Face