CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.4k downloads2y agoHugging Face02sentence-transformers /msmarco-co-condenser-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from co-condenser-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-co-condenser-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M12 likes2.6k downloads2y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.3k downloads2y agoHugging Face04sentence-transformers /msmarco-mpnet-margin-mse-mean-v1 MS MARCO with hard negatives from mpnet-margin-mse-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-mpnet-margin-mse-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.7k downloads2y agoHugging Face05sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.7k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.5k downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes853 downloads2y agoHugging Face08sentence-transformers /msmarco-co-condenser-margin-mse-cls-v1 MS MARCO with hard negatives from co-condenser-margin-mse-cls-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-co-condenser-margin-mse-cls-v1.tabularfeature-extraction10M<n<100M1 likes775 downloads2y agoHugging Face09sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes239 downloads2y agoHugging Face10W-61 /hh-harmless-base-qwen3-8b-margin-dpo-margin-logstabular1K<n<10K0 likes222 downloads7mo agoHugging Face11zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes192 downloads1mo agoHugging Face12W-61 /ultrafeedback-qwen3-8b-margin-dpo-margin-logstabular1K<n<10K0 likes150 downloads7mo agoHugging Face13getminds /marginal-fidelity-survey-eval Synthetic Survey Evaluation: Marginal Fidelity and Response Contracts Matching survey averages does not establish that an AI persona simulates an individual. This small, reproducible evaluation release accompanies Alexander Doudkin's arXiv:2609.07305v1, Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure. Read the practical walkthrough: Do AI Personas Simulate People—or Just… See the full description on the dataset page: https://huggingface.co/datasets/getminds/marginal-fidelity-survey-eval.tabularn<1K0 likes120 downloads12d agoHugging Face14zjhhhh /fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes91 downloads20d agoHugging Face15codemetic /MARGIN Overview Dataset of paper and implementation of MARGIN, Margin-Aware Regularized Geometry for Imbalance Vulnerability DetectioN Reference @misc{zhang2026MARGIN, title={MARGIN: Margin-Aware Regularized Geometry for Imbalanced Vulnerability Detection}, author={Yuteng Zhang and Huifang Ma and Jiahui Wei and Qingqing Li and Yafei Yang}, year={2026}, eprint={2605.10240}, archivePrefix={arXiv}, primaryClass={cs.SE}… See the full description on the dataset page: https://huggingface.co/datasets/codemetic/MARGIN.text100K<n<1M0 likes89 downloads3mo agoHugging Face16drproduck /pickapic-5k-high-margin-sortedimage1K<n<10K0 likes85 downloads11mo agoHugging Face17zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes85 downloads20d agoHugging Face18matCercola18 /quotient-margins-reward-models Quotient Margins for Reward Models — data release Artifacts backing the paper Measure Confidence on Decisions, Not Samples: Quotient Margins for Reward Models. The short version of the paper. Reward models pick the best of N sampled responses, but their confidence is normally read off the reward gap between the top two samples. When several candidates express the same underlying behaviour, that gap is a within-class spacing and its predictive signal cancels. Measuring the margin… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/quotient-margins-reward-models.texttext-generation0 likes82 downloads3d agoHugging Face19kalbin /moshi-on-policy-dpo-margin3audio1K<n<10K0 likes72 downloads5mo agoHugging Face20dougdotcon /douvras-quote-margin-reasoning Douvras Quote and Margin Reasoning v0.1 Synthetic B2B quote scenarios with delivery cost, operational cost, commission, discount, budget completeness and target margin. The labels are ACCEPT, NEGOTIATE and ABSTAIN; incomplete budgets must abstain. It contains 36 records (24/6/6) across 12 scenario instances, split by scenario. This is a calculation protocol, not financial advice. Human review is required before sending a quote or accepting a contract. tabularn<1K0 likes53 downloads13d agoHugging Face21hi-todayis-jh /fixed-n-rb-offset-cost-aware-marginrl-qwen3-1.7b-base-math12k-offset2048-token-mean-rollouts fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset2048_token_mean rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes51 downloads7d agoHugging Face22zjhhhh /er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42-rollouts er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular10K<n<100K0 likes49 downloads14d agoHugging Face23professorsynapse /eh-margin-evidence-responsiveness-worldknown margin-evidence-responsiveness-worldknown -- aggregate exhaust Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those. HF repo: professorsynapse/eh-margin-evidence-responsiveness-worldknown Provenance Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-margin-evidence-responsiveness-worldknown.texttext-classificationn<1K0 likes44 downloads29d agoHugging Face24zjhhhh /fixed-n-rb-offset-marginrl-qwen3-4b-base-polaris53k-offset512-token-mean-1epoch-rollouts fixed_n_rb_offset_marginrl_Qwen3-4B-Base_polaris53k_offset512_token_mean_1epoch rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes38 downloads10d agoHugging Face25andersonbcdefg /combined_triples_with_marginstabular1M<n<10M0 likes35 downloads3y agoHugging Face26andersonbcdefg /synthetic_nli_with_marginstabular10K<n<100K0 likes34 downloads3y agoHugging Face27gupta-tanish /QwQ-Long-CoT-30k-subset-Llama3.1-8B-dynamic-perturbation-regex-generation-max-margintabular100K<n<1M0 likes34 downloads1y agoHugging Face28BigCatc /ultrafeedback_small_margin_high_chstabular10K<n<100K0 likes33 downloads2y agoHugging Face29deu05232 /repro_msmarco-w-instructions_seed42-multipos-margintext100K<n<1M1 likes32 downloads4mo agoHugging Face30Asap7772 /persona_gpt4_paired_margin1_allsplittabular100K<n<1M0 likes31 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.