datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ArenaBench
ArenaBench
Dataset Description
ArenaBench is a large-scale benchmark designed for comprehensive evaluation of deepfake detection methods. It contains approximately 45K testing samples and evaluates detectors across five levels, including in-domain, cross-model, cross-manipulation, commercial models, and real-world scenarios.
Overview
Overview of the ArenaBench benchmark.
Benchmark Design
ArenaBench evaluates deepfake… See the full description on the dataset page: https://huggingface.co/datasets/ceil2/ArenaBench.vibe-landing-page-arena
Vibe Landing Page Arena
A large-scale human preference dataset for evaluating AI-generated landing page design quality. 36,000 pairwise judgments from 3,492 annotators comparing landing pages generated by Claude Code, Cursor, Lovable, and Replit across 100 prompts and 4 design dimensions.
Overview
Metric
Value
Total judgments
36,000
Unique annotators
3,492
Prompts
100
Business categories
97
Design tones
82
Tools compared
4 (Claude Code, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/vibe-landing-page-arena.vibe-design-arena
Vibe Design Arena
The first public pairwise human preference dataset for real-world vibe-coded web applications.
60 apps from the Vibe Coding Showcase were compared pairwise by human annotators who judged which app has better visual design based on screenshots. Every possible pair was evaluated (C(60,2) = 1,770 comparisons), with 30 human votes per pair.
Dataset Summary
Stat
Value
Apps
60
Pairwise comparisons
1,770
Human votes per comparison
30
Total… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/vibe-design-arena.
