CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.4k downloads2y agoHugging Face02sentence-transformers /msmarco-co-condenser-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from co-condenser-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-co-condenser-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M12 likes2.6k downloads2y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.3k downloads2y agoHugging Face04sentence-transformers /msmarco-mpnet-margin-mse-mean-v1 MS MARCO with hard negatives from mpnet-margin-mse-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-mpnet-margin-mse-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.7k downloads2y agoHugging Face05sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.7k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.5k downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes853 downloads2y agoHugging Face08sentence-transformers /msmarco-co-condenser-margin-mse-cls-v1 MS MARCO with hard negatives from co-condenser-margin-mse-cls-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-co-condenser-margin-mse-cls-v1.tabularfeature-extraction10M<n<100M1 likes775 downloads2y agoHugging Face09sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes239 downloads2y agoHugging Face10codemetic /MARGIN Overview Dataset of paper and implementation of MARGIN, Margin-Aware Regularized Geometry for Imbalance Vulnerability DetectioN Reference @misc{zhang2026MARGIN, title={MARGIN: Margin-Aware Regularized Geometry for Imbalanced Vulnerability Detection}, author={Yuteng Zhang and Huifang Ma and Jiahui Wei and Qingqing Li and Yafei Yang}, year={2026}, eprint={2605.10240}, archivePrefix={arXiv}, primaryClass={cs.SE}… See the full description on the dataset page: https://huggingface.co/datasets/codemetic/MARGIN.text100K<n<1M0 likes89 downloads3mo agoHugging Face11drproduck /pickapic-5k-high-margin-sortedimage1K<n<10K0 likes85 downloads11mo agoHugging Face12kalbin /moshi-on-policy-dpo-margin3audio1K<n<10K0 likes72 downloads5mo agoHugging Face13younghyopark /pick_tblock_mp_safe_margin_maxvelThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "DualPanda", "total_episodes": 1000, "total_frames": 243000, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:1000" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null, "features": {… See the full description on the dataset page: https://huggingface.co/datasets/younghyopark/pick_tblock_mp_safe_margin_maxvel.tabularrobotics100K<n<1M0 likes44 downloads1y agoHugging Face14jsanzolac /msmarco_marginmse_qwen3_wordglove msmarco_marginmse_qwen3_wordglove Self-contained MS MARCO MarginMSE dataset. Per unique text: wikigiga tokens (student input) + Qwen3-Embedding-8B teacher vector (MRL[1024], L2-normalized). row_map.parquet maps each triplet to (query_idx, positive_idx, neg_idxs). Train margin: cos(t_q,t_pos)-cos(t_q,t_neg) computed on the fly from query_emb.npy / passage_emb.npy. 1M<n<10M0 likes43 downloads4mo agoHugging Face15andersonbcdefg /combined_triples_with_marginstabular1M<n<10M0 likes35 downloads3y agoHugging Face16andersonbcdefg /synthetic_nli_with_marginstabular10K<n<100K0 likes34 downloads3y agoHugging Face17gupta-tanish /QwQ-Long-CoT-30k-subset-Llama3.1-8B-dynamic-perturbation-regex-generation-max-margintabular100K<n<1M0 likes34 downloads1y agoHugging Face18BigCatc /ultrafeedback_small_margin_high_chstabular10K<n<100K0 likes33 downloads2y agoHugging Face19deu05232 /repro_msmarco-w-instructions_seed42-multipos-margintext100K<n<1M1 likes32 downloads4mo agoHugging Face20Asap7772 /persona_gpt4_paired_margin1_allsplittabular100K<n<1M0 likes31 downloads2y agoHugging Face21mnoukhov /openai_summarize_comparisons_tldrprompt_relabel1b_margintabular10K<n<100K0 likes29 downloads3y agoHugging Face22Margin2003 /koch_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 2, "total_frames": 861, "total_tasks": 1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Margin2003/koch_test.tabularrobotics10K<n<100K0 likes29 downloads2mo agoHugging Face23Asap7772 /persona_gpt4_paired_margin1_tuplesplit_filteredtabular100K<n<1M0 likes28 downloads2y agoHugging Face24drproduck /fifa_100k_high_margin_sortedimage100K<n<1M1 likes28 downloads10mo agoHugging Face25Asap7772 /persona_gpt4_paired_margin5tabular100K<n<1M0 likes27 downloads3y agoHugging Face26Asap7772 /persona_gpt4_paired_margin10tabular100K<n<1M0 likes27 downloads3y agoHugging Face27younghyopark /pick_tblock_mp_safe_margin_maxvel_more_conservativeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "DualPanda", "total_episodes": 1000, "total_frames": 243000, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:1000" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null, "features": {… See the full description on the dataset page: https://huggingface.co/datasets/younghyopark/pick_tblock_mp_safe_margin_maxvel_more_conservative.tabularrobotics100K<n<1M0 likes27 downloads1y agoHugging Face28mnoukhov /openai_summarize_generated_20k_relabel_1b_margintabular10K<n<100K0 likes26 downloads3y agoHugging Face29andersonbcdefg /filtered_triples_with_marginstabular1M<n<10M0 likes25 downloads3y agoHugging Face30mnoukhov /openai_summarize_generated_20k_relabel_pythia410m-dpo1_margintabular10K<n<100K0 likes25 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.