datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CineBoard3D-plus
🎬 CineBoard3D++: Dynamic 3D Story World Dataset
📊 Dataset Summary
CineBoard3D++ is a collection of editable, movie-inspired 3D story worlds built with StoryBlender for narrative-grounded camera planning and world visual attention. It brings together story scripts, animated characters, scene geometry, and shot-level configurations in native Blender projects.
The benchmark covers 50 stories, 457 scenes, 1,585 shots, and 3,197 3D assets (836 plot-related and 2,361… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringAI-LAB/CineBoard3D-plus.floorplans-cityscapes
Dataset Summary
This is a curated collection of floorplan images sourced from across the internet. It is intended for research in architectural AI, layout generation, and urban scene understanding.
Data format: Image files with associated integer labels.
Sources: Publicly available images from various web sources (This dataset is one unified collections).
Purpose: Educational and research use.
Dataset Structure
The dataset follows the standard Hugging Face Image… See the full description on the dataset page: https://huggingface.co/datasets/wheres-my-python/floorplans-cityscapes.CityCube-Benchcinematic-world-stills
Council of AI — cinematic stills
Cinematic stills produced for Council of AI surfaces. metadata.jsonl gives the
Hub image viewer a caption per file. These are illustrations — they carry no measurement and back no slot.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that GET, never a second
engine. If the fetch fails the honest answer is UNCHECKABLE — never a fabricated 0.000.
Status… See the full description on the dataset page: https://huggingface.co/datasets/csoai/cinematic-world-stills.CineBenchSyn
CineBenchSyn
CineBenchSyn is the synthetic benchmark for CineOrchestra,
a unified model for cinematic video generation that jointly controls subjects, events, camera, and shot transitions.
It contains 512 hand-authored 10.2-second scenarios that target under-represented, edge-case
cinematic compositions (large casts, dense events, frequent shot transitions). Each scenario is
expressed with the same entity-centric primitive used by CineOrchestra: every cinematic element —
a… See the full description on the dataset page: https://huggingface.co/datasets/sharathgirish/CineBenchSyn.cities
Geomelon — World Cities, Regions & Countries
Multilingual geographic reference data — cities, regions, and countries — sourced from
Wikidata and maintained by Geomelon. Every
city carries population, coordinates, elevation, area, postal/dialing codes, settlement type, and
name translations into 50+ languages. Released under CC0 1.0 — public domain, no attribution
required, free for commercial use.
One config per country (see the dropdown above) — pick a country to avoid… See the full description on the dataset page: https://huggingface.co/datasets/geomelon/cities.CIGEval_sft_data
Dataset Card for CIGEval_sft_data
CIGEval_sft_data is the dataset used for fine-tuning LMMs in the paper CIGEval. It contains data on both tool selection and image evaluation, which can be combined into 2.3k complete evaluation trajectories. The dataset was constructed through the following steps:
Using GPT-4o + CIGEval to evaluate the full ImageHub dataset, generating 4,903 evaluation trajectories.
Randomly selecting 60% of these and filtering out the ones where the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/HIT-TMG/CIGEval_sft_data.satellite-civilian-conflict-disruption-reporter-v1
Satellite Civilian Conflict Disruption Reporter v1
Dataset ID: ChrisRPL/satellite-civilian-conflict-disruption-reporter-v1
Status
This is a valid diagnostic reporter-schema dataset, not the current Blackline Atlas canonical model gate. The canonical compact calibration/gold dataset remains ChrisRPL/satellite-disruption-triage-aux-v2-2.
Use this dataset for future schema-simplification experiments only after respecting the mixed source licenses. Do not treat the associated… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-civilian-conflict-disruption-reporter-v1.FiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/CiaranCw/FiVE-Fine-Grained-Video-Editing-Benchmark.circuitvqadesc2circocigeval-sft-datacircuitvqacircuitvqadescCIS_4190_Final_Project
