datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
argument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.newyorker_caption_ranking
New Yorker Caption Ranking Dataset
Dataset Descriptions
Homepage: https://nextml.github.io/caption-contest-data/
Repository: https://github.com/yguooo/cartoon-caption-generation
Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
Point of Contact: yguo@cs.wisc.edu
Dataset Summary
We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.ust-rankings
UST Rankings
Daily course and instructor rating marts for UST Rankings, built from the
ust-archive datasets.
File
Contents
courses.parquet
Current Course metadata by Course Code.
course-ratings.parquet
Longitudinal course ratings by term and criterion.
instructor-ratings.parquet
Longitudinal instructor ratings by term and criterion.
course-rankings.parquet
Latest-term course ratings.
instructor-rankings.parquet
Latest-term instructor ratings.… See the full description on the dataset page: https://huggingface.co/datasets/ust-archive/ust-rankings.tts-ranking-dataRoboTwin_blocks_ranking_rgb_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 216980,
"total_tasks": 136,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_blocks_ranking_rgb_randomized.RoboTwin_blocks_ranking_size_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 217762,
"total_tasks": 89,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_blocks_ranking_size_randomized.msmarco_passage_ranking_corpusThis is the preprocessed data from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
robotwin-blocks_ranking_rgb-500-A800robotwin-blocks_ranking_size-500-A800robotwin_blocks_ranking_rgb_hybrid_200_dynFcam
robotwin_blocks_ranking_rgb_hybrid_200_dynFcam
A validated LeRobot v2.1 release of two native RoboTwin 2.0 expert schedules for blocks_ranking_rgb.
The variants use independent accepted seeds and are concatenated into one training split; matching episode offsets are not paired scenes.
dynFcam observation derivative
This repository reuses exactly the same validated native episodes as Shiki42/robotwin_blocks_ranking_rgb_hybrid_200 at revision… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_blocks_ranking_rgb_hybrid_200_dynFcam.alitaqi000_world-university-rankings-2023
World University Rankings 2023
World University Rankings 2023 include 1,799 universities across 104 countries.
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 8,164
Files: 1
Files
World University Rankings 2023.csv
Mirrored from Kaggle
dureader-retrieval-ranking
dureader
数据来自DuReader-Retreval数据集,这里是原始地址。
本数据集只用作学术研究使用。如果本仓库涉及侵权行为,会立即删除。
Document_ranking_testNoisy-MSMARCO-Passage-RankingThis link gathers 72 noisy versions of the MS-Marco-Passage Ranking dataset consisting of three noise types (insertion, deletion, substitution), two different distributions of errors in the text (Batch 1 where errors are distributed in few words in the text and Batch2 where errors are more evenly spread out between words) and 12 different intensities of noise (CER varying from 3% to 36% with intervals of 3%).
The exact dataset that has been used is the MS-Marco-passagetest2020-top1000. The… See the full description on the dataset page: https://huggingface.co/datasets/edwardgiamphy/Noisy-MSMARCO-Passage-Ranking.Ranking_TVRmotive-v2-prediction-rankings
MOTIVE v2 Gene-Compound Prediction Rankings
This Hugging Face repository is the browsable Data Studio mirror of the two Parquet artifacts in Zenodo record 22105202.
Zenodo is the canonical source for citation, provenance, methods, versioning, file integrity, and detailed interpretation.
Exact version DOI: 10.5281/zenodo.22105202
Scientific context and limitations: MOTIVE Issue 12
Paper: MOTIVE: A Drug-Target Interaction Graph For Inductive Link Prediction
Browse… See the full description on the dataset page: https://huggingface.co/datasets/carpenter-singh-lab/motive-v2-prediction-rankings.blocks_ranking_rgb_20_10_31_v2.1blocks_ranking_size_24_11_01_v2.1TowerBlocks-MT-Ranking
Dataset Card for TowerBlocks-MT-Ranking (GQM Ranking Annotations)
Summary
TowerBlocks-MT-Ranking is a group-wise machine translation ranking dataset annotated under the Group Quality Metric (GQM) paradigm.Each example contains a source sentence and a group of 2–4 candidate translations, which are jointly evaluated to produce a relative quality ranking (and associated group-relative scores/labels). The annotations are produced by Gemini-2.5-Pro using GQM-style… See the full description on the dataset page: https://huggingface.co/datasets/double7/TowerBlocks-MT-Ranking.robotwin-blocks-ranking-rgb-rollouts
RoboTwin blocks_ranking_rgb — Wan2.2 TI2V Rollouts
160 text+image-to-video rollouts (10 initial conditions × 16 random seeds) for the
blocks_ranking_rgb task from RoboTwin, generated with the Wan2.2 TI2V (5B)
diffusion model fine-tuned with a merged Vidar LoRA adapter, and scored with the
blocks_ranking_v2 reward (SAM3 object tracking + IDM inverse-dynamics + FK
gripper ↔ block position matching).
Companion to the EmbodiedVideoRL / DanceGRPO
reward-model work.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/VincentNi/robotwin-blocks-ranking-rgb-rollouts.tmp_fastwam-data-blocks_ranking_rgb2repro-beyond-model-ranking-predictability-aligned-evaluation-for-time-series-forecasting-results
Beyond Model Ranking reproduction results
This dataset repository contains scripts, tests, raw tables, figures, and intermediate predictions for an independent reproduction of Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting.
The reproduction follows the algorithms in the authors official repository and uses the four public ETT datasets from the official ETT repository. The original ETT CSVs are not duplicated here; rerun commands fetch them… See the full description on the dataset page: https://huggingface.co/datasets/apararti/repro-beyond-model-ranking-predictability-aligned-evaluation-for-time-series-forecasting-results.msmarco-passage-rankingmsmarco_passage_ranking_official_trainThis is the preprocessed training data from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
uni-rankings-2026
BrightKey Independent University Rankings Dataset (2026)
299 universities × 55 countries × 6 dimensions, evaluated independently. No payments from institutions accepted. Public data only.
This is the open release of the BrightKey university rankings — an independent alternative to QS, THE, and Shanghai rankings. Released under CC BY 4.0.
Live site: https://brightkey.co/en/rankings/methodology
GitHub repo: https://github.com/arthurb2l/brightkey-university-dataset
Zenodo DOI:… See the full description on the dataset page: https://huggingface.co/datasets/brightkey/uni-rankings-2026.Ranking-benchrobotwin_blocks_ranking_rgb_hybrid_200
robotwin_blocks_ranking_rgb_hybrid_200
A validated LeRobot v2.1 release of two native RoboTwin 2.0 expert schedules for blocks_ranking_rgb.
The variants use independent accepted seeds and are concatenated into one training split; matching episode offsets are not paired scenes.
Composition
LeRobot episodes
Schedule variant
Episodes
Aligned samples
Native config
0-99
non_overlapping_reference
100
50,954
parallelvla_non_overlapping_reference_verified_v1… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_blocks_ranking_rgb_hybrid_200.usc_xarm_policy_rankingPickaPic-rankings
Dataset Card for "PickaPic-rankings"
More Information needed
acm-icaif-2025_chunk_ranking
