datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-distilbert-margin-mse-mean-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.opacity-marginalization
Amortized opacity marginalization: data
Training sets, trained models, priors and held-out evaluation sets for Amortized Opacity Marginalization Improves C/O Interval Calibration for Brown-Dwarf Retrievals (Heraty 2026, arXiv:2609.01665). Code and paper source: github.com/kaileh57/opacity-marginalization.
Contents
spectra/v7_opmarg: one million simulated NIRSpec G395H spectra generated with randomized molecular opacities (the marginalized training set)… See the full description on the dataset page: https://huggingface.co/datasets/Kaileh57/opacity-marginalization.msmarco-co-condenser-margin-mse-sym-mnrl-mean-v1
MS MARCO with hard negatives from co-condenser-margin-mse-sym-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-co-condenser-margin-mse-sym-mnrl-mean-v1.msmarco-distilbert-margin-mse-sym-mnrl-mean-v2
MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.msmarco-mpnet-margin-mse-mean-v1
MS MARCO with hard negatives from mpnet-margin-mse-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-mpnet-margin-mse-mean-v1.msmarco-distilbert-margin-mse-cls-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.msmarco-distilbert-margin-mse-sym-mnrl-mean-v1
MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.msmarco-distilbert-margin-mse-mnrl-mean-v1
MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.msmarco-co-condenser-margin-mse-cls-v1
MS MARCO with hard negatives from co-condenser-margin-mse-cls-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-co-condenser-margin-mse-cls-v1.msmarco-distilbert-margin-mse-cls-dot-v2
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.hh-harmless-base-qwen3-8b-margin-dpo-margin-logsfixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
ultrafeedback-qwen3-8b-margin-dpo-margin-logsmarginal-fidelity-survey-eval
Synthetic Survey Evaluation: Marginal Fidelity and Response Contracts
Matching survey averages does not establish that an AI persona simulates an individual. This small, reproducible evaluation release accompanies Alexander Doudkin's arXiv:2609.07305v1, Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure.
Read the practical walkthrough: Do AI Personas Simulate People—or Just… See the full description on the dataset page: https://huggingface.co/datasets/getminds/marginal-fidelity-survey-eval.fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts
fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
MARGIN
Overview
Dataset of paper and implementation of MARGIN, Margin-Aware Regularized Geometry for Imbalance Vulnerability DetectioN
Reference
@misc{zhang2026MARGIN,
title={MARGIN: Margin-Aware Regularized Geometry for Imbalanced Vulnerability Detection},
author={Yuteng Zhang and Huifang Ma and Jiahui Wei and Qingqing Li and Yafei Yang},
year={2026},
eprint={2605.10240},
archivePrefix={arXiv},
primaryClass={cs.SE}… See the full description on the dataset page: https://huggingface.co/datasets/codemetic/MARGIN.pickapic-5k-high-margin-sortedfixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
quotient-margins-reward-models
Quotient Margins for Reward Models — data release
Artifacts backing the paper Measure Confidence on Decisions, Not Samples: Quotient Margins for
Reward Models.
The short version of the paper. Reward models pick the best of N sampled responses, but
their confidence is normally read off the reward gap between the top two samples. When
several candidates express the same underlying behaviour, that gap is a within-class spacing and
its predictive signal cancels. Measuring the margin… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/quotient-margins-reward-models.moshi-on-policy-dpo-margin3douvras-quote-margin-reasoning
Douvras Quote and Margin Reasoning v0.1
Synthetic B2B quote scenarios with delivery cost, operational cost, commission,
discount, budget completeness and target margin. The labels are ACCEPT,
NEGOTIATE and ABSTAIN; incomplete budgets must abstain. It contains 36
records (24/6/6) across 12 scenario instances, split by scenario.
This is a calculation protocol, not financial advice. Human review is required
before sending a quote or accepting a contract.
fixed-n-rb-offset-cost-aware-marginrl-qwen3-1.7b-base-math12k-offset2048-token-mean-rollouts
fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset2048_token_mean rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42-rollouts
er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
pick_tblock_mp_safe_margin_maxvelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "DualPanda",
"total_episodes": 1000,
"total_frames": 243000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/younghyopark/pick_tblock_mp_safe_margin_maxvel.eh-margin-evidence-responsiveness-worldknown
margin-evidence-responsiveness-worldknown -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-margin-evidence-responsiveness-worldknown
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-margin-evidence-responsiveness-worldknown.msmarco_marginmse_qwen3_wordglove
msmarco_marginmse_qwen3_wordglove
Self-contained MS MARCO MarginMSE dataset. Per unique text: wikigiga tokens (student input) + Qwen3-Embedding-8B teacher vector (MRL[1024], L2-normalized). row_map.parquet maps each triplet to (query_idx, positive_idx, neg_idxs).
Train margin: cos(t_q,t_pos)-cos(t_q,t_neg) computed on the fly from query_emb.npy / passage_emb.npy.
pick_tblock_mp_safe_margin_maxvel
pick_tblock_mp_safe_margin_maxvel TsFile Conversion
This dataset is a TsFile conversion of younghyopark/pick_tblock_mp_safe_margin_maxvel, a LeRobot v2.1 DualPanda robot dataset.
Modalities: Time-series. The source metadata declares total_videos=0 and video_path=null.
Source Dataset Facts
From the original dataset card and meta/info.json:
Source dataset: younghyopark/pick_tblock_mp_safe_margin_maxvel
License: apache-2.0
Codebase version: v2.1
Robot type:… See the full description on the dataset page: https://huggingface.co/datasets/THULab/pick_tblock_mp_safe_margin_maxvel.marginal-medqa_marginal_9b_v2eh-margin-mapping
margin-mapping -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-margin-mapping
Provenance
Experiment: experiments/margin-mapping
Amendment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-margin-mapping.bookmaker-margin-panel
Bookmaker Margin Panel — daily measured betting margins across 22 leagues and 3 market types
A daily panel of measured bookmaker margins (overround/vigorish) across head-to-head, spread
and totals markets: one row per day × league × region × market type, aggregated from over
4 million individually recorded odds observations (2.7 million cleaned markets).
To our knowledge this is the only openly published bookmaker margin panel. It quantifies the
price of sports betting — the… See the full description on the dataset page: https://huggingface.co/datasets/edushinka/bookmaker-margin-panel.
