datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embedding-training-data
Training Data for Text Embedding Models
[!NOTE]
This repository contains raw datasets, all of which have also been formatted for easy training in the Embedding Model Datasets collection. We recommend looking there first.
This repository contains training files to train text embedding models, e.g. using sentence-transformers.
Data Format
All files are in a jsonl.gz format: Each line contains a JSON-object that represent one training example.
The JSON objects can… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/embedding-training-data.KaLM-reranker-training-data
Lychee-KaLM-Reranker Training Data
A large-scale, ready-to-use multilingual dataset for fine-tuning reranking models.
This repository contains 3,885,265 training samples collected from 54 datasets, covering English, Chinese, and multilingual retrieval tasks. Each sample includes task instructions, positive passages, at least 16 hard negatives, and teacher scores annotated by Qwen3-Reranker-8B.
When expanded into point-wise query–passage pairs, the dataset provides at least 66… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-reranker-training-data.turkish_embedding_model_training_datacleaned_turkish_embedding_model_training_data_colabcleaned_turkish_embedding_model_training_data_colab
Citation
If you use this dataset in your research, please cite the following paper:
@inproceedings{baysan-gungor-2025-tr,
title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations",
author = "Baysan, Mehmet Selman and
Gungor, Tunga",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025",
month = nov,
year = "2025",
address = "Suzhou, China",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.eti-embedding-training-data-2048-v3
Version note (v3): third generation of the ETI dense training data. Reuses the questions of NorskHelsenett/eti-embedding-training-data-v2 (which trained eti-embeddinggemma-v2), but re-maps every question to its article URL, re-cuts positives at a 2048-token window, re-mines hard negatives with granite + reranker validation (pos_score/neg_score), adds the keyword half, and filters junk anchors. Trained NorskHelsenett/eti-embeddinggemma-v3.
eti-embedding-training-data-2048… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-v3.nordic-embedding-training-data
Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset
The purpose of this dataset is to pre- or post-train embedding models for Danish on text similarity tasks.
The dataset is structured for training using InfoNCE loss (also known as SimCSE loss, Cross-Entropy Loss with in-batch negatives, or simply in-batch negatives loss), with hard-negative samples for the tasks of retrieval and unit-triplet. Beware that if fine-tuning the unit-triplets for… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/nordic-embedding-training-data.scandinavian-embedding-training-dataturkish_embedding_model_training_dataeti-embedding-training-data-2048-triplets
ETI Embedding Training Data — Triplets with Hard Negatives
This dataset contains 330,120 (anchor, positive, negative) triplets for training and fine-tuning Norwegian-language embedding models, particularly for health-related retrieval and RAG applications.
How this dataset was created
Source data
The triplets were mined from the source dataset NorskHelsenett/eti-embedding-training-data-2048, which contains 78,888 anchor-positive pairs of Norwegian health… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-triplets.Vietnamese-Embedding-Training-Data
Vietnamese Embedding Training Data
A curated Vietnamese dataset for training semantic embedding and retrieval models.
This dataset combines multiple Vietnamese sources and normalizes them into a unified retrieval-oriented format with positive passages and hard negatives.
Format
Each sample contains:
{
"query": "...",
"positive": "...",
"negatives": ["...", "...", "..."],
"source": "...",
"task": "retrieval",
"domain": "...",
"group_id": "..."… See the full description on the dataset page: https://huggingface.co/datasets/PaxiAI/Vietnamese-Embedding-Training-Data.eti-embedding-training-data-v2
ETI Embedding Training Data v2
576,708 Norwegian-language (anchor, positive, negative) triplets for contrastive fine-tuning of embedding, sparse-retrieval, and reranker models in the Norwegian health and welfare domain.
This dataset is a weighted, deduplicated merge of four upstream triplet datasets, designed to give a single high-quality training signal for two-stage RAG pipelines that retrieve over Norwegian public health information.
The model trained on it —… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-v2.embedding_training_data_v1
Similarity scores from multiple embedding models will be added soon.
eti-embedding-training-data-2048-triplets-merged-v2powershell_embedding_model_training_dataEmbedding-training-sample13eti-embedding-training-data-2048-triplets
ETI Embedding Training Data — Triplets with Hard Negatives
This dataset contains 330,120 (anchor, positive, negative) triplets for training and fine-tuning Norwegian-language embedding models, particularly for health-related retrieval and RAG applications.
How this dataset was created
Source data
The triplets were mined from the source dataset thivy/eti-embedding-training-data-2048, which contains 78,888 anchor-positive pairs of Norwegian health content. That… See the full description on the dataset page: https://huggingface.co/datasets/thivy/eti-embedding-training-data-2048-triplets.eti-embedding-training-data-4096ko-legal-embedding-training-v1
Korean Public Legal Embedding Training v1
실제 embedding 학습 queue가 소비하는 provenance-preserving 파생 JSONL이다.
source-native query/positive 관계와 deterministic bootstrap negative로 컴파일했다. 최종 학습 전 current-student hard-negative mining을 수행해야 한다.
rows: 250,000
release eligible: true
visibility: public
use: public redistribution and model training
exact benchmark query/evaluation-text matches: 0
exact retrieval-corpus matches: 0 unique hashes
Sionic 9 및 MTEB task-family train source가 포함될 수… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models2/ko-legal-embedding-training-v1.eti-embedding-training-data-2048-triplets-v4
ETI Embedding Training Data — v4 (5-style, LLM-judged triplets)
Norwegian (nb) retrieval triplets built from the
NorskHelsenett/LOS_Document_classification_ETI
corpus of public-service / welfare / health documents.
Each row is a triplet (anchor, positive, negative) plus three metadata
columns (style, category, doc_url) you can use to filter, weight, or
build curriculum stages.
76,408 triplets · 2,530 source documents · 38,629 distinct anchors.
Schema
field… See the full description on the dataset page: https://huggingface.co/datasets/thivy/eti-embedding-training-data-2048-triplets-v4.eti-embedding-training-data-2048-triplets-merged-v3eti-embedding-training-data-2048-v2-triplets-v21-cleanedeti-embedding-training-data-2048-triplets-mergedeti-embedding-training-data-summaries-merged-v3-cleanEmbedding-training-sample12eti-embedding-training-data-2048-v2-tripletsEmbedding-training-sample10MNLP_M3_rag_embedding_trainingeti-embedding-training-data-2048
ETI Embedding Training Data (2048 tokens)
This dataset contains 78,888 anchor-positive pairs for training Norwegian-language embedding models focused on health-related content. Each pair consists of a question (anchor) and its corresponding relevant passage (positive).
Dataset format
Column
Description
Example
anchor
A question in Norwegian
"Hva er noen tips for å gjøre leken mer lystbetont for barnet mitt?"
positive
The correct/relevant passage
A passage… See the full description on the dataset page: https://huggingface.co/datasets/thivy/eti-embedding-training-data-2048.eti-embedding-training-data-2048-triplets-v3-cleaned
