datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VOST-TAS
[NeurIPS 2025] Tracking and Understanding Object Transformations
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Dataset Visualizations: GitHub
Paper: arXiv:2511.04678
Project Page: tubelet-graph.github.io
Project Repository: GitHub
Point of Contact: Yihong Sun
📊 Dataset Overview
VOST-TAS (TrackAnyState) is an extended version of the VOST validation set with explicit transformation annotations for tracking and… See the full description on the dataset page: https://huggingface.co/datasets/yihongs/VOST-TAS.regx-benchmark
RegX
Cross-Domain Multi-View Point Cloud Registration Benchmark
RegX evaluates multi-view point cloud registration across scales spanning nine orders
of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors
never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and
solid-state LiDAR, terrestrial and airborne laser scanners.
Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question
instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.ps4mas-final-test-rollouts-0813
PS4MAS Final Test Rollouts (0813)
Source split: ps4mas-0521-splits final_test_scenarios.jsonl
Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json.
Files
Model
Rows
Path
best_rl_gigpo_debate_step40
800
traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.officeqa-checkpoint-eval-data
Checkpoint evaluation plot data
Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures.
No model execution, grading, publication, or source-result changes were performed to make this export.
Contents
checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds.
pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.androidlife-530
AndroidLife-530 — Android agent benchmark (real phone, real LLM)
AndroidLife runs Android agent tasks against a real phone (via ADB/MobileRun)
and a real LLM, and grades the agent on reaching a verifiable device end-state.
This repo ships the 530-task corpus plus everything needed to reproduce runs.
Benchmark, or template — your call. The 530 tasks are an extended version
of the benchmark, usable as a larger evaluation set for further benchmarking of
models beyond the 60-task… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/androidlife-530.geometry-dash-levels
Geometry Dash Level Dataset
Subsets
2024_300k
Dump of ~300k levels from the Geometry Dash servers, sorted by the number of likes. 66 JSONL shards (~660MB each, ~43GB total).
Files: 2024_300k/levels-v1-00000.jsonl through 2024_300k/levels-v1-00065.jsonl
2026_50k_rated
~50k rated/featured levels scraped from the Geometry Dash servers in February 2026. 39 JSONL shards (~500MB each, ~19GB total).
Files: 2026_50k_rated/levels-v2-00000.jsonl through… See the full description on the dataset page: https://huggingface.co/datasets/yusp48/geometry-dash-levels.hhtools_parc_ms
hhtools PARC MS — terrain-aware humanoid motion clips
中文说明
Motion clips in PARC MS layout for human-humanoid-tools (hhtools): per-clip folders with a PARC MSFileData pickle and a static terrain mesh. Suitable for meshmimic / interaction-mesh retargeting (parkour, climbing, box traversal, etc.).
Clips
26,396
Total size
~10.8 GB
Skeleton
15-bone PARC humanoid (humanoid.xml topology)
Terrain
Heightfield + Wavefront OBJ per clip
Source
Converted from PARC… See the full description on the dataset page: https://huggingface.co/datasets/YaojieShen/hhtools_parc_ms.youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.arxiv-metadata-2020-2026
arXiv Metadata, enriched (2020–2026)
Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09,
enriched with abstracts, citation counts, and Semantic Scholar identifiers, and
organized as a two-level hierarchy: field of study → year.
Unlike a bare title index, every record here carries the abstract, the
full author list, citation counts, and the Semantic Scholar corpusId,
so you can do retrieval, classification, citation analysis, and corpus building
directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.YCB-V-DSWe provide our stereo recordings for the YCB-V object dataset, as well as some of our physically-based stereo renderings.
Contents
stereo+depth recordings/ — raw stereo video recordings (.mkv)
train_pbr_left*.z*, train_pbr_right*.z* — physically-based rendered training data (split zip archives), left/right stereo
test_pbr_left.zip — physically-based rendered test data
test_bboxes/ — BOP-format annotation JSONs:
test_targets_pbr_left.json — flat list of {im_id, inst_count… See the full description on the dataset page: https://huggingface.co/datasets/tpoellabauer/YCB-V-DS.FOLIOastra-robodojo-rollouts
Astra RoboDojo Evaluation Records
Rollout records from the evaluations in GPT 6 Astra as an Embodied Policy,
by Jiayi Su, Yixin Zheng, Mi Yan, Li Yi, Zhizheng Zhang, and He Wang. This archive
includes action proposals, executed actions, observations, robot states,
model-provided explanations and reasoning summaries, and metadata for reproducing
the evaluation settings, together with a reader and documentation.
Report
Public controller source
Data schema and alignment… See the full description on the dataset page: https://huggingface.co/datasets/YuMoool/astra-robodojo-rollouts.MMMC
MMMC: Massive Multi-discipline Multimodal Coding Benchmark for Educational Video Generation
Dataset Summary
The MMMC (Massive Multi-discipline Multimodal Coding) benchmark is a curated dataset for Code2Video research, focusing on the automatic generation of professional, discipline-specific educational videos. Unlike pixel-only video datasets, MMMC provides structured metadata that links lecture content with executable code, visual references, and topic-level annotations… See the full description on the dataset page: https://huggingface.co/datasets/YanzheChen/MMMC.final-best-raw-episodes-2026-09-14-v2
Final and canonical-best evaluation episodes: frozen preparation
Full local packaging is now running. See materialization status and instructions. This preparation folder is not the full payload; the separate full export remains incomplete until its verified COMPLETED marker is written.
Prepared inventory: 72 evaluations / 119,608 expected episodes. There are
36 canonical-best and 44 final memberships, with 8 evaluations tagged both.
One incomplete Q38-teacher OfficeQA v2 final… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/final-best-raw-episodes-2026-09-14-v2.OntoKGFollowBench
FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
We introduce FollowBench, a Multi-level Fine-grained Constraints Following Benchmark for systemically and precisely evaluate the instruction-following capability of LLMs.
FollowBench comprehensively includes five different types (i.e., Content, Situation, Style, Format, and Example) of fine-grained constraints.
To enable a precise constraint following estimation on diverse… See the full description on the dataset page: https://huggingface.co/datasets/YuxinJiang/FollowBench.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.netryx-new-york-5km
New York 5km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.712800, -74.006000
Radius: 5.0 km
Panoramas: 196,824
Index entries: 787,296
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("new-york-5km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in the Netryx GUI.
Details… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-5km.netryx-new-york-city-13km
Nyc-Core-Usethis 13km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.713200, -74.002500
Radius: 13.0 km
Panoramas: 663,084
Index entries: 2,652,336
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.netryx-new-york-0kmxiangqi-dataset
Xiangqi (Chinese Chess) Gameplay Trajectories & Visualizations Dataset
This dataset contains 1,000 high-quality Chinese Chess (Xiangqi) matches extracted and processed from the open-source training pipeline of Pikafish (the leading neural-network-backed Xiangqi engine). The original source training trajectories are credited to the px0data dataset on Kaggle.
For each match, this dataset provides both structured, step-by-step action sequences (JSONL format) suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/ysong18/xiangqi-dataset.HealthChat-11K
HealthChat-11K
This repository contains HealthChat-11K, a curated dataset of approximately 11,000 real-world conversations, composed of 25,000 user messages, where users seek healthcare information from Large Language Models (LLMs). The goal of this work is to provide a high-quality resource for systematically studying and improving health conversations involving humans and AI (e.g., LLMs).
The dataset was presented in the paper: "What's Up, Doc?": Analyzing How Users Seek Health… See the full description on the dataset page: https://huggingface.co/datasets/yahskapar/HealthChat-11K.lm-eval-results-yleo-EmertonMonarch-7B-slerp-private
Dataset Card for Evaluation run of yleo/EmertonMonarch-7B-slerp
Dataset automatically created during the evaluation run of model yleo/EmertonMonarch-7B-slerp
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yleo-EmertonMonarch-7B-slerp-private.icongenai-svg-captions
IconGenAI SVG Captions
Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models.
Part of the IconGenAI research project.
Files
Two files are provided at different stages of the processing pipeline:
File
Records
Purpose
icons_captioned_merged.jsonl
275,912
Full license-filtered corpus with VLM-generated captions and collection metadata
icons_training_captioned.jsonl227,821
Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.I-SHEEP-Self-Datayoutube-highlights-full
YouTube Highlights 完整媒体与标注
本仓库面向数据集协作交付,提供一个可断点续传的完整 tar 文件。解压后即可得到视频、
官方标签转换结果、字段说明和本地可视化检查页。
数据概况
6 个类别:dog、gymnastics、parkour、skating、skiing、surfing
417 个通过 ffprobe 完整性检查的 MP4
315 个 human_mturk 视频:具有 MTurk 人工软投票分数
102 个 weak_match 视频:只有官方自动匹配弱标签
官方清单中另有 1 个当前不可下载的视频,未进入训练标注
19 个已下载视频存在媒体帧数与官方标注帧号差异,保留在数据集中并单独列入复核清单
Linux 下载与解压
BASE_URL="https://huggingface.co/datasets/jhanglee/youtube-highlights-full/resolve/main"
wget -c… See the full description on the dataset page: https://huggingface.co/datasets/jhanglee/youtube-highlights-full.agentic_vbench_rollouts
agentic-vbench calibration rollout archive
Complete, immutable copies of the calibration trajectories referenced from the
agentic-vbench task PRs. Unlike the copies committed in each task's
calibration/rollouts/, nothing here is elided: every sampled-frame payload the agent
saw is present.
Redaction policy (mechanical, applied identically to every file):
Absolute host filesystem paths from the runner machine are replaced with /workspace.
Capture-hardware and source-collection… See the full description on the dataset page: https://huggingface.co/datasets/yalesunxiatao/agentic_vbench_rollouts.Ivy-Fake
IVY-FAKE: Unified Explainable Benchmark and Detector for AIGC Content
This repository provides the official implementation of IVY-FAKE and IVY-xDETECTOR, a unified explainable framework and benchmark for detecting AI-generated content (AIGC) across both images and videos.
🔍 Overview
IVY-FAKE is the first large-scale dataset designed for multimodal explainable AIGC detection. It contains:
150K+ training samples (images + videos)
18.7K evaluation samples
Fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/yishiliu/Ivy-Fake.minimax-m3-deepsearchqa-skill-eval
MiniMax M3 DeepSearchQA Skill Eval
Evaluates minimax/minimax-m3 on google/deepsearchqa using a Pi agent, You.com MCP tools, and a research skill optimized for this harness, model, and tool surface.
MiniMax M3 Medium Reasoning with the You.com research skill reached 74.85% adjusted F1 on DeepSearchQA, above the paper's GPT-5 High Reasoning F1 result. Public artifacts are available for inspection and reproduction.
Links
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/youdotcom/minimax-m3-deepsearchqa-skill-eval.lm-eval-results-yunconglong-DARE_TIES_13B-private
Dataset Card for Evaluation run of yunconglong/DARE_TIES_13B
Dataset automatically created during the evaluation run of model yunconglong/DARE_TIES_13B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yunconglong-DARE_TIES_13B-private.
