CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.7k downloads3y agoHugging Face02yguooo /newyorker_caption_ranking New Yorker Caption Ranking Dataset Dataset Descriptions Homepage: https://nextml.github.io/caption-contest-data/ Repository: https://github.com/yguooo/cartoon-caption-generation Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning Point of Contact: yguo@cs.wisc.edu Dataset Summary We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.imagetext-generation1M<n<10M6 likes1.4k downloads2y agoHugging Face03ust-archive /ust-rankings UST Rankings Daily course and instructor rating marts for UST Rankings, built from the ust-archive datasets. File Contents courses.parquet Current Course metadata by Course Code. course-ratings.parquet Longitudinal course ratings by term and criterion. instructor-ratings.parquet Longitudinal instructor ratings by term and criterion. course-rankings.parquet Latest-term course ratings. instructor-rankings.parquet Latest-term instructor ratings.… See the full description on the dataset page: https://huggingface.co/datasets/ust-archive/ust-rankings.tabular10M<n<100M0 likes976 downloads4d agoHugging Face04arpesenti /tts-ranking-data0 likes805 downloads10mo agoHugging Face05suz22 /RoboTwin_blocks_ranking_rgb_randomizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aloha", "total_episodes": 500, "total_frames": 216980, "total_tasks": 136, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:500" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_blocks_ranking_rgb_randomized.imagerobotics100K<n<1M0 likes704 downloads6mo agoHugging Face06suz22 /RoboTwin_blocks_ranking_size_randomizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aloha", "total_episodes": 500, "total_frames": 217762, "total_tasks": 89, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:500" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_blocks_ranking_size_randomized.imagerobotics100K<n<1M0 likes651 downloads6mo agoHugging Face07jacklin /msmarco_passage_ranking_corpusThis is the preprocessed data from msmarco passage(v1) ranking corpus. MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,. text1M<n<10M0 likes399 downloads4y agoHugging Face08SakikoTogawa /robotwin-blocks_ranking_rgb-500-A8000 likes339 downloads5mo agoHugging Face09SakikoTogawa /robotwin-blocks_ranking_size-500-A8000 likes274 downloads5mo agoHugging Face10Shiki42 /robotwin_blocks_ranking_rgb_hybrid_200_dynFcam robotwin_blocks_ranking_rgb_hybrid_200_dynFcam A validated LeRobot v2.1 release of two native RoboTwin 2.0 expert schedules for blocks_ranking_rgb. The variants use independent accepted seeds and are concatenated into one training split; matching episode offsets are not paired scenes. dynFcam observation derivative This repository reuses exactly the same validated native episodes as Shiki42/robotwin_blocks_ranking_rgb_hybrid_200 at revision… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_blocks_ranking_rgb_hybrid_200_dynFcam.tabularrobotics10K<n<100K1 likes271 downloads2mo agoHugging Face11jason1966 /alitaqi000_world-university-rankings-2023 World University Rankings 2023 World University Rankings 2023 include 1,799 universities across 104 countries. Dataset Info Source: Kaggle Original Size: 0.07 MB Kaggle Downloads: 8,164 Files: 1 Files World University Rankings 2023.csv Mirrored from Kaggle tabular1K<n<10K0 likes266 downloads6mo agoHugging Face12zyznull /dureader-retrieval-ranking dureader 数据来自DuReader-Retreval数据集,这里是原始地址。 本数据集只用作学术研究使用。如果本仓库涉及侵权行为,会立即删除。 3 likes210 downloads4y agoHugging Face13YanAdjeNole /Document_ranking_testtext100K<n<1M0 likes171 downloads2y agoHugging Face14edwardgiamphy /Noisy-MSMARCO-Passage-RankingThis link gathers 72 noisy versions of the MS-Marco-Passage Ranking dataset consisting of three noise types (insertion, deletion, substitution), two different distributions of errors in the text (Batch 1 where errors are distributed in few words in the text and Batch2 where errors are more evenly spread out between words) and 12 different intensities of noise (CER varying from 3% to 36% with intervals of 3%). The exact dataset that has been used is the MS-Marco-passagetest2020-top1000. The… See the full description on the dataset page: https://huggingface.co/datasets/edwardgiamphy/Noisy-MSMARCO-Passage-Ranking.0 likes131 downloads3y agoHugging Face15axgroup /Ranking_TVR1 likes122 downloads2y agoHugging Face16carpenter-singh-lab /motive-v2-prediction-rankings MOTIVE v2 Gene-Compound Prediction Rankings This Hugging Face repository is the browsable Data Studio mirror of the two Parquet artifacts in Zenodo record 22105202. Zenodo is the canonical source for citation, provenance, methods, versioning, file integrity, and detailed interpretation. Exact version DOI: 10.5281/zenodo.22105202 Scientific context and limitations: MOTIVE Issue 12 Paper: MOTIVE: A Drug-Target Interaction Graph For Inductive Link Prediction Browse… See the full description on the dataset page: https://huggingface.co/datasets/carpenter-singh-lab/motive-v2-prediction-rankings.tabular10M<n<100M0 likes118 downloads24d agoHugging Face17DorayakiLin /blocks_ranking_rgb_20_10_31_v2.1videon<1K0 likes109 downloads11mo agoHugging Face18DorayakiLin /blocks_ranking_size_24_11_01_v2.1videon<1K0 likes108 downloads11mo agoHugging Face19double7 /TowerBlocks-MT-Ranking Dataset Card for TowerBlocks-MT-Ranking (GQM Ranking Annotations) Summary TowerBlocks-MT-Ranking is a group-wise machine translation ranking dataset annotated under the Group Quality Metric (GQM) paradigm.Each example contains a source sentence and a group of 2–4 candidate translations, which are jointly evaluated to produce a relative quality ranking (and associated group-relative scores/labels). The annotations are produced by Gemini-2.5-Pro using GQM-style… See the full description on the dataset page: https://huggingface.co/datasets/double7/TowerBlocks-MT-Ranking.text10K<n<100K0 likes97 downloads7mo agoHugging Face20VincentNi /robotwin-blocks-ranking-rgb-rollouts RoboTwin blocks_ranking_rgb — Wan2.2 TI2V Rollouts 160 text+image-to-video rollouts (10 initial conditions × 16 random seeds) for the blocks_ranking_rgb task from RoboTwin, generated with the Wan2.2 TI2V (5B) diffusion model fine-tuned with a merged Vidar LoRA adapter, and scored with the blocks_ranking_v2 reward (SAM3 object tracking + IDM inverse-dynamics + FK gripper ↔ block position matching). Companion to the EmbodiedVideoRL / DanceGRPO reward-model work. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/VincentNi/robotwin-blocks-ranking-rgb-rollouts.textn<1K0 likes91 downloads5mo agoHugging Face21SakikoTogawa /tmp_fastwam-data-blocks_ranking_rgb2videon<1K0 likes83 downloads4mo agoHugging Face22apararti /repro-beyond-model-ranking-predictability-aligned-evaluation-for-time-series-forecasting-results Beyond Model Ranking reproduction results This dataset repository contains scripts, tests, raw tables, figures, and intermediate predictions for an independent reproduction of Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting. The reproduction follows the algorithms in the authors official repository and uses the four public ETT datasets from the official ETT repository. The original ETT CSVs are not duplicated here; rerun commands fetch them… See the full description on the dataset page: https://huggingface.co/datasets/apararti/repro-beyond-model-ranking-predictability-aligned-evaluation-for-time-series-forecasting-results.1 likes81 downloads2mo agoHugging Face23zyznull /msmarco-passage-ranking1 likes80 downloads4y agoHugging Face24jacklin /msmarco_passage_ranking_official_trainThis is the preprocessed training data from msmarco passage(v1) ranking corpus. MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,. text100K<n<1M0 likes78 downloads4y agoHugging Face25brightkey /uni-rankings-2026 BrightKey Independent University Rankings Dataset (2026) 299 universities × 55 countries × 6 dimensions, evaluated independently. No payments from institutions accepted. Public data only. This is the open release of the BrightKey university rankings — an independent alternative to QS, THE, and Shanghai rankings. Released under CC BY 4.0. Live site: https://brightkey.co/en/rankings/methodology GitHub repo: https://github.com/arthurb2l/brightkey-university-dataset Zenodo DOI:… See the full description on the dataset page: https://huggingface.co/datasets/brightkey/uni-rankings-2026.tabulartabular-classificationn<1K0 likes76 downloads3mo agoHugging Face26ulab-ai /Ranking-benchtext10K<n<100K0 likes74 downloads1y agoHugging Face27Shiki42 /robotwin_blocks_ranking_rgb_hybrid_200 robotwin_blocks_ranking_rgb_hybrid_200 A validated LeRobot v2.1 release of two native RoboTwin 2.0 expert schedules for blocks_ranking_rgb. The variants use independent accepted seeds and are concatenated into one training split; matching episode offsets are not paired scenes. Composition LeRobot episodes Schedule variant Episodes Aligned samples Native config 0-99 non_overlapping_reference 100 50,954 parallelvla_non_overlapping_reference_verified_v1… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_blocks_ranking_rgb_hybrid_200.tabularrobotics10K<n<100K0 likes73 downloads2mo agoHugging Face28aliangdw /usc_xarm_policy_rankingtextn<1K0 likes72 downloads10mo agoHugging Face29yuvalkirstain /PickaPic-rankings Dataset Card for "PickaPic-rankings" More Information needed tabular10K<n<100K2 likes71 downloads4y agoHugging Face30BloomBerry /acm-icaif-2025_chunk_rankingtext10K<n<100K0 likes62 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.