datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
contrastive-pretraining
Contrastive Pretraining
Per-language query/document pairs produced by the retrieval-common-crawl pipeline.
Each config corresponds to a single language or source with identical LightOn-style schema.
Config overview
Configs available are: fw-edu, fw2-arb_Arab, fw2-ces_Latn, fw2-cmn_Hani, fw2-dan_Latn, fw2-deu_Latn, fw2-ell_Grek, fw2-fas_Arab, fw2-fra_Latn, fw2-hun_Latn, fw2-ind_Latn, fw2-ita_Latn, fw2-jpn_Jpan, fw2-nld_Latn, fw2-pol_Latn, fw2-por_Latn, fw2-rus_Cyrl… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/contrastive-pretraining.C4-contrastive-watermark
Dataset Card for "C4-contrastive-watermark"
More Information needed
opengloss-v1.3-contrastive-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Contrastive Examples v1.3
Dataset Summary
OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations
designed for contrastive learning and… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-contrastive-examples.gepa-rlm-exp-contrastive-20260219-191545
gepa-rlm-exp-contrastive-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: contrastive | Last updated: 2026-02-19 21:46 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
contrastive
48.89%
33.33%
826,936
$0.0000
6170s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-contrastive-20260219-191545.gepa-exp-contrastive-vanilla-20260220-083526
gepa-exp-contrastive-vanilla-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: contrastive | Last updated: 2026-02-20 18:21 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_vanilla_k3
vanilla
3
contrastive
31.11%
44.00%
51,600
$0.1236
32317s
Learning Curves
Experiment Config
{
"script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-contrastive-vanilla-20260220-083526.ms-contrastive-100kcontext-conditioned-record-transfer-v10.3-tdc-bbb-martins-contrastive-intern
BBB V10.3-TDC dual-view contrastive data
This release converts each frozen V10.3-TDC pair into two independent LM inputs.
The reference view exposes the known outcome or measurement; the query view hides it.
The original comparison question, choices, and completion are not included.
Source: jiosephlee/context-conditioned-molecule-transfer-v10.3-tdc-bbb-martins-mixed-continuous-intern at 49021c7e4609e9eb902200cb5a60a0c534a294fe
Train rows: 117,408
Direct/auxiliary ratio: 1:1… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/context-conditioned-record-transfer-v10.3-tdc-bbb-martins-contrastive-intern.gepa-exp-contrastive-rlm-20260220-083526
gepa-exp-contrastive-rlm-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: contrastive | Last updated: 2026-02-20 19:26 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
contrastive
35.56%
40.67%
365,790
$0.0000
33195s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-contrastive-rlm-20260220-083526.liars-bench-contrastiveturkish_weakly_supervised_contrastive_learning_dataset_filteredcontext-conditioned-record-transfer-v10.3-tdc-v2-hp-bioavailability-ma-contrastive-intern
Oral V10.3-TDC heldout-parent with Starling direct dual-view contrastive data
Each frozen pair is converted into independent reference and query LM inputs. The
reference exposes its known outcome or measurement; the query hides it. Pairwise
questions, choices, and completions are excluded.
Source: jiosephlee/context-conditioned-molecule-transfer-v10.3-tdc-v2-hp-bioavailability-ma-mixed-continuous-intern at 9db99d12536f5dbb3d906ab90e5acff1b036405c
Train rows: 98,448… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/context-conditioned-record-transfer-v10.3-tdc-v2-hp-bioavailability-ma-contrastive-intern.C4-contrastive-watermark
Dataset Card for "C4-contrastive-watermark"
More Information needed
opengloss-v1.1-contrastive-examples
OpenGloss Contrastive Examples v1.1
Dataset Summary
OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations
designed for contrastive learning and semantic similarity training. Each example contains
a source sentence and a 5-point semantic gradient showing how meaning shifts from
antonym to synonym poles.
This dataset is derived from the OpenGloss
encyclopedic dictionary, using example sentences and their lexical context to generate… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-contrastive-examples.context-conditioned-record-transfer-v9.0.2-tdc-v2-skin-reaction-contrastive-intern
Skin Reaction V9.0.2-TDC with Starling direct dual-view contrastive data
Each frozen pair is converted into independent reference and query LM inputs. The
reference exposes its known outcome or measurement; the query hides it. Pairwise
questions, choices, and completions are excluded.
Source: jiosephlee/context-conditioned-molecule-transfer-v9.0.2-tdc-v2-skin-reaction-mixed-continuous-intern at 4d97c745c735aa8d6dc07b81305e93f4e3bea496
Train rows: 197,714
Direct/auxiliary rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/context-conditioned-record-transfer-v9.0.2-tdc-v2-skin-reaction-contrastive-intern.context-conditioned-record-transfer-v10.3-tdc-hp-bioavailability-ma-contrastive-intern
Oral V10.3-TDC heldout-parent indirect-only dual-view contrastive data
Each frozen pair is converted into independent reference and query LM inputs. The
reference exposes its known outcome or measurement; the query hides it. Pairwise
questions, choices, and completions are excluded.
Source: jiosephlee/context-conditioned-molecule-transfer-v10.3-tdc-hp-bioavailability-ma-mixed-continuous-intern at d7e7eae329eeebe7860ea33fefbb47f25fa724ce
Train rows: 21,504
Direct/auxiliary rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/context-conditioned-record-transfer-v10.3-tdc-hp-bioavailability-ma-contrastive-intern.context-conditioned-record-transfer-v9.0.2-tdc-skin-reaction-contrastive-intern
Skin Reaction V9.0.2-TDC indirect-only dual-view contrastive data
Each frozen pair is converted into independent reference and query LM inputs. The
reference exposes its known outcome or measurement; the query hides it. Pairwise
questions, choices, and completions are excluded.
Source: jiosephlee/context-conditioned-molecule-transfer-v9.0.2-tdc-skin-reaction-mixed-continuous-intern at 9e4f3edf6796d754bac576fa645a1bf84f292ff4
Train rows: 27,072
Direct/auxiliary rows: 13,536/13… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/context-conditioned-record-transfer-v9.0.2-tdc-skin-reaction-contrastive-intern.dsprites-gold-wrong-contrastive-observational-over_specifiedNemotron-RL-Agentic-Contrastive-Eval-swe-v1-prefix-contrastattributionBench_contrastive1neg_mismatchogbert-v1-contrastive
OGBert Contrastive Combined Dataset
This dataset combines contrastive learning signals from two sources:
Contrastive Examples: Gradient-based semantic similarity pairs
Definitions: Word-level semantic relationships (synonyms, antonyms, definitions, etc.)
Dataset Statistics
Total pairs: 9,358,022
Training pairs: 8,890,120
Evaluation pairs: 467,902
Breakdown by Source
Contrastive Dataset (500,000 pairs):
Gradient signals: All C(5,2) = 10 pairwise… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/ogbert-v1-contrastive.opengloss-v1.2-contrastive-examples
OpenGloss Contrastive Examples v1.2
Dataset Summary
OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations
designed for contrastive learning and semantic similarity training. Each example contains
a source sentence and a 5-point semantic gradient showing how meaning shifts from
antonym to synonym poles.
This dataset is derived from the OpenGloss
encyclopedic dictionary, using example sentences and their lexical context to generate… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-contrastive-examples.contrastivedistillprefdataexopengloss-embedding-contrastiveLFQA_eval_dataset_unit_tests_contrastiveNemotron-RL-Agentic-Contrastive-Eval-contrast-v1Nemotron-RL-Agentic-Contrastive-Eval-swe-v1-contrastturkish_weakly_supervised_contrastive_learning_dataset_filtereddsprites-gold-wrong-contrastive-observational-under_specifiedasrs-data-contrastive-non-aircraftalbums_with_moods_contrastive
Dataset Details (Card Organization by HuggingFace)
Dataset Description
This dataset contains synthetic natural-language music queries paired with albums, designed for fine-tuning contrastive and embedding-based retrieval models. Queries are generated from album reviews using curated adjective and music-descriptor vocabularies, enabling semantic alignment between descriptive text and musical works.
The dataset emphasizes mood, texture, and stylistic descriptors rather than… See the full description on the dataset page: https://huggingface.co/datasets/mmarkusmalone/albums_with_moods_contrastive.
