datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FilmBench
FilmBench — Video Generation Benchmark Dataset
📢 Update (2026-08-02)
Added English prompt files: filmbench_prompts_en.csv is now available with English prompts.
filmbench_prompts_en.csv (1,169 rows): Prompt-level table with English prompts. Columns: uid, task, movie_type (English), en_prompt, reference_url.
📢 Update (2026-07-31)
Fixed a batch of misaligned prompts in filmbench_videos.csv: the zh_prompt column has been recalibrated against the… See the full description on the dataset page: https://huggingface.co/datasets/skylenage/FilmBench.SNU_FILMIndustryCorpus2_film_entertainment
IndustryCorpus2: Film & Entertainment
This repository contains the IndustryCorpus2: Film & Entertainment domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_film_entertainment.real01b-md2-r5-iql3-film150-taub08f-tiidk4c200-n32-heval-s2026070704This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-md2-r5-iql3-film150-taub08f-tiidk4c200-n32-heval-s2026070704.real01b-md2-r5-repeat-base-dp-filmtiidk4-c200k-n32-s2026070704This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-md2-r5-repeat-base-dp-filmtiidk4-c200k-n32-s2026070704.FilmEval
FilmEval
FilmEval is a benchmark dataset for evaluating novel-to-film generation. It contains source novels grouped by difficulty and generated film/video outputs from multiple inference systems. The dataset is designed for comparing how well different models transform a written story into a complete short film.
Dataset Summary
FilmEval contains:
15 source novels
3 difficulty levels: easy, medium, and hard
5 source novels per difficulty level
6 model inference… See the full description on the dataset page: https://huggingface.co/datasets/ZuoHaotong/FilmEval.th-en-zh-tts-200k-enhanced
TH-EN-ZH Multi-speaker TTS Dataset (200K, RE-USE Enhanced)
Speech-enhanced variant of FILM6912/th-en-zh-tts-200k.
Every clip has been processed through NVIDIA RE-USE (universal speech enhancement, SEMamba) at its native sample rate, then re-encoded losslessly as FLAC (PCM_16).
Same schema, same row order, same 200,000 rows (th 100k / en 50k / zh 50k):
Column
Type
Description
text
string
Transcript (identical to the original dataset)
audio
Audio
Enhanced audio, FLAC… See the full description on the dataset page: https://huggingface.co/datasets/FILM6912/th-en-zh-tts-200k-enhanced.real01b-square-d2-r5-redo-base-dp-filmtiidk4-c200k-n32-s2026081801This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-square-d2-r5-redo-base-dp-filmtiidk4-c200k-n32-s2026081801.real01b-md2-r4-repeat-base-dp-filmtiidk4-c200k-n32-s2026070604This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-md2-r4-repeat-base-dp-filmtiidk4-c200k-n32-s2026070604.photon47-extended-film-highlights
Photon47 extended-film highlights
Seven at-least-60-second, silent H.264/MP4 real-gameplay reels derived from the public creator film
Photon47: Multiplayer Cute Monster Battle Game.
Each website mode card plays its matching reel first, followed by the existing
automated mouse/keyboard gameplay capture.
Contents
moba.addendum.mp4 — team clashes, hero swarms, beams, ice, and vortex effects
arena.addendum.mp4 — card deployment, counterpush, and finale… See the full description on the dataset page: https://huggingface.co/datasets/ryan-superman/photon47-extended-film-highlights.real01b-square-d2-r5-cand3-taub09-c250k-film-c200k-h6bs1024-c200k-heval-sobolseed2026070901This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-square-d2-r5-cand3-taub09-c250k-film-c200k-h6bs1024-c200k-heval-sobolseed2026070901.sonic-forage-h3-film-lab
Sonic Forage — H3 Film Lab
8 short films, each built from six chained MiniMax H3 shots with native generated audio, rendered on a single 33 GB GPU with 8 sampling steps per shot.
Everything here is AI-generated video — label it as such when you publish it.
Download everything
hf download TheMindExpansionNetwork/sonic-forage-h3-film-lab --repo-type dataset --local-dir sonic-forage-h3-films
One file at a time:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/sonic-forage-h3-film-lab.types-of-film-shots
What a Shot!
2,919 film frames labeled with shot scale (8 classes), from film-grab.com.
annotator
count
how labeled
human
863
hand-labeled (gold)
ai
2,056
DINOv2 classifier; low-confidence frames re-judged by Claude Opus (active learning)
Columns
image, label — ambiguous, closeUp, detail, extremeLongShot, fullShot, longShot, mediumCloseUp, mediumShot
annotator — human or ai
source — human / v2_highconf / opus_review / opus_resolved_amb… See the full description on the dataset page: https://huggingface.co/datasets/szymonrucinski/types-of-film-shots.LiCo-Film
LiCo-Film: Standardized Film-Based LiDAR Cover Contamination Dataset
The release includes the approved BSD-3-Clause license documentation in
LICENCE.txt.
LiCo-Film is a dataset and benchmark for LiDAR cover contamination under
dynamic driving conditions. It uses standardized transparent contamination
films to create reproducible dry and wet proxy-contamination states and
separate winter-road recordings to evaluate transfer to naturally accumulated
dry real contamination.
The… See the full description on the dataset page: https://huggingface.co/datasets/JannisGrimminger/LiCo-Film.hitchcock-psycho-1960-film-dataset-transformed
Psycho → AI-Model Dataset (Transformed)
A thematic re-skin of the Psycho (1960) Q&A dataset into an original AI-model setting where the world is transformed into an AI/data-center environment.
Character names, actor names, objects, locations, production references, dates, and thematic elements are remapped to AI/ML concepts and modern technology.
File: psycho_dataset_transformed.jsonl
Format: JSONL — one JSON object per line
Schema: each line has prompt and completion string… See the full description on the dataset page: https://huggingface.co/datasets/antfr99/hitchcock-psycho-1960-film-dataset-transformed.filmsupply_028magnetic-pendulum-films
The magnetic pendulum in 2D — basin-of-attraction films
Four 1080p60 films, ~30 s each, of the basin-of-attraction map of the magnetic
pendulum: release the bob from rest at every point of the plane, integrate until
it settles, and colour that point by the magnet that captured it.
This is a real, named, well-studied object, not a repurposed Newton fractal. Its
signature is the Wada property: every point on the boundary between two basins
is also on the boundary of the third, at… See the full description on the dataset page: https://huggingface.co/datasets/damiavicens/magnetic-pendulum-films.IndustryCorpus_film[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_film.filmsupply_026korean_film_window_30sth-en-zh-tts-200k
TH-EN-ZH Multi-speaker TTS Dataset (200K)
A unified multi-speaker text-to-speech dataset combining three high-quality speech corpora, normalized to a single schema:
Column
Type
Description
text
string
Transcript (whitespace-normalized; LibriTTS uses normalized text)
audio
Audio
Embedded audio bytes (original format: MP3 for th / WAV for en, zh)
speaker_id
string
Namespaced speaker ID (cv_th_*, libritts_*, aishell3_*)
lang
string
th, en, or zh… See the full description on the dataset page: https://huggingface.co/datasets/FILM6912/th-en-zh-tts-200k.filmsupply_016filmsupply_000200_Filmes_de_Maior_Faturamento_-_Outubro_2022Artigos da Wikipédia para os 200 filmes de maior faturamento em outubro de 2022
license: mit
filmsupply_024film-metricsfilmsupply_029filmseek-dataset
SentenceTransformer based on sentence-transformers/all-MiniLM-L6-v2
This is a sentence-transformers model finetuned from sentence-transformers/all-MiniLM-L6-v2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for retrieval.
Model Details
Model Description
Model Type: Sentence Transformer
Base model: sentence-transformers/all-MiniLM-L6-v2
Maximum Sequence Length: 256 tokens
Output Dimensionality: 384… See the full description on the dataset page: https://huggingface.co/datasets/celaphee/filmseek-dataset.filmfrench-film-reviews
