datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Movie101
Movie101
[!NOTE]
Please carefully read the Movie101 license before using the data.Current dataset version: Movie101v2
Audio Description (AD) describes movie content in real time to help visually impaired individuals enjoy movies, where a narration speech briefly summarizes the ongoing plots during pauses in character dialogue, help its audience keep up with the movie.
The AD creation involves extensive work by human experts, which is costly and difficult to cover the vast array… See the full description on the dataset page: https://huggingface.co/datasets/yuezih/Movie101.MovieChat-1K_trainmovie_posters-100k
Dataset Card for "movie_posters-100k"
More Information needed
movielens-25m-thumb
🍿 Popcorn Thumbnails Embeddings
This dataset contains deep visual features obtained from +65000 movie thumbnails.
It contains extracted visual features using modern VLMs.
To simply load it, Popcorn framework has been developed that can be used in movie recommendation, information retrieval, classification, etc tasks.
📚 Citation
@article{popcorn,
title={Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation},
author={Tourani… See the full description on the dataset page: https://huggingface.co/datasets/alitourani/movielens-25m-thumb.movie-scenes-captionedmovies-dataset
+9000 Movie Dataset
Overview
This dataset is sourced from Kaggle and has been granted CC0 1.0 Universal (CC0 1.0) Public Domain Dedication by the original author. This means you can copy, modify, distribute, and perform the work, even for commercial purposes, all without asking permission.
I would like to express our gratitude to the original author for their contribution to the data community.
License
This dataset is released under the CC0 1.0 Universal… See the full description on the dataset page: https://huggingface.co/datasets/Pablinho/movies-dataset.movielens_ratingsblender-open-movies-medialetterboxd-all-movie-data
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/letterboxd-all-movie-data.movie-scenesembedded_movies
sample_mflix.embedded_movies
This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast.
In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature.
Overview
This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.MovieTection
Dataset Description 🎬
The MovieTection dataset is a benchmark designed for detecting pretraining data in Large Vision-Language Models (VLMs). It serves as a resource for analyzing model exposure to Copyrighted Visual Content ©️.
Paper: DIS-CO: Discovering Copyrighted Content in VLMs Training Data
Direct Use 🖥️
The dataset is designed for image/caption-based question-answering, where models predict the movie title given a frame or its corresponding textual… See the full description on the dataset page: https://huggingface.co/datasets/DIS-CO/MovieTection.movie_stills_captioned_dataset_local
Dataset Card for "movie_stills_captioned_dataset_local"
More Information needed
movie-posters
Dataset Card for "movie_posters"
More Information needed
MovieStills_Captioned_SmolVLM
From the Frontier Research Team at Takara.ai we present MovieStills_Captioned_SmolVLM, a dataset of 75,000 movie stills with high-quality synthetic captions generated using SmolVLM.
Dataset Description
This dataset contains 75,000 movie stills, each paired with a high-quality synthetic caption. It was generated using the HuggingFaceTB/SmolVLM-256M-Instruct model, designed for instruction-tuned multimodal tasks. The dataset aims to support image captioning tasks, particularly for… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/MovieStills_Captioned_SmolVLM.Movie-Poster-WebURL-Dataset-1874-2025
Movie Poster WebURL Dataset 1874–2025
A TMDB-derived metadata index of movie poster WebURLs covering 1874–2025.
The dataset contains metadata and external TMDB poster URLs. Poster image binaries are not redistributed in this repository.
Data
Split: train
Rows: 804,304
Format: Parquet
Columns: 15
The publication artifact was produced from a larger local TMDB harvest and passed a conservative metadata-based content filtering and post-filter verification process… See the full description on the dataset page: https://huggingface.co/datasets/ROSCOSMOS/Movie-Poster-WebURL-Dataset-1874-2025.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
Movieseance-movie
SÉANCE — projected presence
A 1 m 46 s explainer film for SÉANCE, the studio's projected-presence concept: turning one
phone video into an apparition you can project into a real room.
Watch it here: seance-explainer-1080p.mp4 (1920×1080, 30 fps,
H.264 + AAC, −14.4 LUFS)
The companion concept page, with a playable demo of the Crossover effect, is live at
https://sonicforage.com/seance/
What the film says
Hook — every projected trick stops at the wall; this one… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/seance-movie.MovieRatingDBAll data acquired from - https://www.omdbapi.com/
MovieTection_Mini
Dataset Description 🎬
The MovieTection_Mini dataset is a benchmark designed for detecting pretraining data in Large Vision-Language Models (VLMs). It serves as a resource for analyzing model exposure to Copyrighted Visual Content ©️.
This dataset is a compact subset of the full MovieTection dataset, containing only 4 movies instead of 100. It is designed for users who want to experiment with the benchmark without the need to download the entire dataset, making it a more… See the full description on the dataset page: https://huggingface.co/datasets/DIS-CO/MovieTection_Mini.Movie_Profitability_Analysis
Movie Profitability Analysis - EDA Summary
Dataset Overview
This project explores the “Movies Metrics, Features and Statistics” dataset from Kaggle.The dataset contains 6,569 movies and 32 features, including:
Production Budget
Worldwide & Domestic Gross
Running Time
Genre
Creative Type
Production Method
Ratings and Release Date
The goal is to understand which pre-release factors influence a movie’s ability to generate positive profit.… See the full description on the dataset page: https://huggingface.co/datasets/Leelu1002/Movie_Profitability_Analysis.movie-posters-genres-80k
Dataset Card for "movie-posters-genres-80k"
More Information needed
movie-postersmovie_posters_100k_controlnetDataset Name: 10k Movie Poster Images with Layouts and Captions
Description:
This dataset contains 10,000 movie poster images, along with their extracted layout information and captions. The captions are generated by concatenating the movie title and genre(s). The layout annotations were extracted using PaddleOCR, providing precise structural details of the posters.
Source:
The dataset is a curated set of the movie-posters-100k dataset.
Key Features:
Images: 10,000 high-resolution movie… See the full description on the dataset page: https://huggingface.co/datasets/stzhao/movie_posters_100k_controlnet.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
EDA_on_IMDB_Movies_Datasetmovie_stills_dataset
Dataset Card for "movie_stills_dataset"
More Information needed
movies_CLIP_ViT-L14
🎬 Movie Frame & Caption Dataset
📖 Introduction
This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted.
This dataset can be used for tasks such as:
Video understanding
Multimodal learning (image + text)
Image captioning
Vision-language retrieval
📂 Data… See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.movies-cinemagoer
Movies Cinemagoer
A 7.5k-sized 34-column movie dataset extracted from Cinemagoer.
