datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
argument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.newyorker_caption_ranking
New Yorker Caption Ranking Dataset
Dataset Descriptions
Homepage: https://nextml.github.io/caption-contest-data/
Repository: https://github.com/yguooo/cartoon-caption-generation
Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
Point of Contact: yguo@cs.wisc.edu
Dataset Summary
We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.ust-rankings
UST Rankings
Daily course and instructor rating marts for UST Rankings, built from the
ust-archive datasets.
File
Contents
courses.parquet
Current Course metadata by Course Code.
course-ratings.parquet
Longitudinal course ratings by term and criterion.
instructor-ratings.parquet
Longitudinal instructor ratings by term and criterion.
course-rankings.parquet
Latest-term course ratings.
instructor-rankings.parquet
Latest-term instructor ratings.… See the full description on the dataset page: https://huggingface.co/datasets/ust-archive/ust-rankings.msmarco_passage_ranking_corpusThis is the preprocessed data from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
alitaqi000_world-university-rankings-2023
World University Rankings 2023
World University Rankings 2023 include 1,799 universities across 104 countries.
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 8,164
Files: 1
Files
World University Rankings 2023.csv
Mirrored from Kaggle
Document_ranking_testmotive-v2-prediction-rankings
MOTIVE v2 Gene-Compound Prediction Rankings
This Hugging Face repository is the browsable Data Studio mirror of the two Parquet artifacts in Zenodo record 22105202.
Zenodo is the canonical source for citation, provenance, methods, versioning, file integrity, and detailed interpretation.
Exact version DOI: 10.5281/zenodo.22105202
Scientific context and limitations: MOTIVE Issue 12
Paper: MOTIVE: A Drug-Target Interaction Graph For Inductive Link Prediction
Browse… See the full description on the dataset page: https://huggingface.co/datasets/carpenter-singh-lab/motive-v2-prediction-rankings.TowerBlocks-MT-Ranking
Dataset Card for TowerBlocks-MT-Ranking (GQM Ranking Annotations)
Summary
TowerBlocks-MT-Ranking is a group-wise machine translation ranking dataset annotated under the Group Quality Metric (GQM) paradigm.Each example contains a source sentence and a group of 2–4 candidate translations, which are jointly evaluated to produce a relative quality ranking (and associated group-relative scores/labels). The annotations are produced by Gemini-2.5-Pro using GQM-style… See the full description on the dataset page: https://huggingface.co/datasets/double7/TowerBlocks-MT-Ranking.robotwin-blocks-ranking-rgb-rollouts
RoboTwin blocks_ranking_rgb — Wan2.2 TI2V Rollouts
160 text+image-to-video rollouts (10 initial conditions × 16 random seeds) for the
blocks_ranking_rgb task from RoboTwin, generated with the Wan2.2 TI2V (5B)
diffusion model fine-tuned with a merged Vidar LoRA adapter, and scored with the
blocks_ranking_v2 reward (SAM3 object tracking + IDM inverse-dynamics + FK
gripper ↔ block position matching).
Companion to the EmbodiedVideoRL / DanceGRPO
reward-model work.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/VincentNi/robotwin-blocks-ranking-rgb-rollouts.uni-rankings-2026
BrightKey Independent University Rankings Dataset (2026)
299 universities × 55 countries × 6 dimensions, evaluated independently. No payments from institutions accepted. Public data only.
This is the open release of the BrightKey university rankings — an independent alternative to QS, THE, and Shanghai rankings. Released under CC BY 4.0.
Live site: https://brightkey.co/en/rankings/methodology
GitHub repo: https://github.com/arthurb2l/brightkey-university-dataset
Zenodo DOI:… See the full description on the dataset page: https://huggingface.co/datasets/brightkey/uni-rankings-2026.msmarco_passage_ranking_official_trainThis is the preprocessed training data from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
Ranking-benchPickaPic-rankings
Dataset Card for "PickaPic-rankings"
More Information needed
usc_xarm_policy_rankingE2Rank_ranking_datasetsacm-icaif-2025_chunk_rankingusc_koch_p_ranking_rfmusc_franka_policy_rankingranking-fas-results
Ranking FAS Results
Ranking FAS Results v1 is a compact, project-generated metrics dataset (1,874 rows across 6 tables) that compares algorithms for ranking from pairwise comparisons — the problem of turning a set of noisy, possibly-cyclic pairwise preference judgments (A beats B, B beats C, C beats A, ...) into a single consistent ranking. It reports how well feedback-arc-set-style methods (Soroush Vahidi's own OURS_MFAS family), classical ranking baselines (e.g. SpringRank… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/ranking-fas-results.utd_so101_clean_policy_ranking_topmsmarco_passage_ranking
MS MARCO Passage Ranking
Dataset description
MS MARCO (MicroSoft MAchine Reading COmprehension) is a large-scale collection built for machine reading comprehension and information retrieval research. The original release introduced more than one million real user questions sampled from Bing search logs, paired with passages drawn from web documents, and human-authored answers where applicable.
The passage ranking track uses a fixed corpus of short text passages and asks… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/msmarco_passage_ranking.Document_ranking6_testmsmarco_passage_ranking_queriesThis is the preprocessed queries from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
tooldrift-model-rankings
ToolDrift: OpenRouter model usage rankings, captured daily
One row per model per ranking window per capture: its rank, the tokens and requests behind that rank, and its share of the window. The series shows which models the market actually routes work to, day by day.
Rows in this cut
44,369
One row is
one model in one ranking window on one capture day
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-model-rankings.tooldrift-app-rankings
ToolDrift: OpenRouter app usage rankings, captured daily
One row per app per ranking window per capture, with the tool it maps to where ToolDrift tracks one. It is the same series as the model rankings, read from the consumer side.
Rows in this cut
641
One row is
one app in one ranking window on one capture day
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method
https://toolproof.thecompound.tech/methodology
Licence… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-app-rankings.PickaPic-rankings-6-3-2023
Dataset Card for "PickaPic-rankings-6-3-2023"
More Information needed
esci_ranking_liteambrosia-ranking-binarised-pcracm-icaif-2025_document_rankingBASIR_Budget_Assisted_Sectoral_Impact_RankingSector Classification Dataset: sector_classification/sectorwise_budget_text_extracted.xlsx
Description of columns:
date - Date of Budget
year - Year of Budget
budget_full_text_filename - Name of file having full budget transcript
text_segment - Text excerpts related to the identified sector
sector - Identified sector (TARGET)
Sector Ranking Dataset: sector_ranking/sectorwise_budgetdaywise_performance_with_text_ranked_full.xlsx
Description of columns:
date - Date… See the full description on the dataset page: https://huggingface.co/datasets/sohomghosh/BASIR_Budget_Assisted_Sectoral_Impact_Ranking.
