datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EASI-Leaderboard-Data
EASI Leaderboard Data
A consolidated dataset for the EASI Leaderboard, containing the evaluation data (inputs/prompts) actually used on the leaderboard across spatial reasoning benchmarks for VLMs.
Looking for the Spatial Intelligence leaderboard?https://huggingface.co/spaces/lmms-lab-si/EASI-Leaderboard
🔎 Dataset Summary
Question types: MCQ (multiple choice) and NA (numeric answer).
File format: TSV only.
Usage: These TSVs are directly consumable by the EASI… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-si/EASI-Leaderboard-Data.tts_leaderboard_screenshotsidp-leaderboard-resultsALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.leaderboard-dataocr-leaderboard
OmniAI OCR Leaderboard
A comprehensive leaderboard comparing OCR and data extraction performance across traditional OCR providers and multimodal LLMs, such as gpt-4o and gemini-2.0. The dataset includes full results from testing 9 providers on 1,000 pages each.
Benchmark Results (Feb 2025) | Source Code
quantum-leaderboard-assetsperturb-leaderboard-history
Perturb Leaderboard History
Append-oriented archive derived from the public Perturb leaderboard API.
challenges: 500 tasks, one clean image per task.
miner_outputs: 9294 archived top-miner responses with images and metrics.
dashboard_history: 377254 task/validator/UID snapshots with 50-point score history.
Output images are selected from the union of top 15 rolling-score miners and
top 15 current-score valid miners. Dashboard history contains every reported
UID. Presigned URL… See the full description on the dataset page: https://huggingface.co/datasets/defqon-1/perturb-leaderboard-history.leaderboard-artifactsidp-leaderboard-results
