datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rotten_tomatoes
Dataset Card for "rotten_tomatoes"
Dataset Summary
Movie Review Dataset.
This is a dataset of containing 5,331 positive and 5,331 negative processed
sentences from Rotten Tomatoes movie reviews. This data was first used in Bo
Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for
sentiment categorization with respect to rating scales.'', Proceedings of the
ACL, 2005.
Supported Tasks and Leaderboards
More Information Needed
Languages… See the full description on the dataset page: https://huggingface.co/datasets/cornell-movie-review-data/rotten_tomatoes.movie_binaries_0012movie_binaries_0010MovieNet
Dataset
This is an unofficial host for the MovieNet dataset.Refer to the official project repo for more details (data format, license, etc.).
Usage
Combine the split files into a large zip: cat frames_part_* > frames.zip.
P.S. I host this repo for easier public access to the original large MovieNet dataset, since the official website has problems for years.If the authors find it an infringement of rights, plz contact me for deletion.
training-movies-test-stuffMovie101
Movie101
[!NOTE]
Please carefully read the Movie101 license before using the data.Current dataset version: Movie101v2
Audio Description (AD) describes movie content in real time to help visually impaired individuals enjoy movies, where a narration speech briefly summarizes the ongoing plots during pauses in character dialogue, help its audience keep up with the movie.
The AD creation involves extensive work by human experts, which is costly and difficult to cover the vast array… See the full description on the dataset page: https://huggingface.co/datasets/yuezih/Movie101.movie_binaries_0013Japanese_NicoNico_Douga_Movie_Comment_Data_2018Japanese_NicoNico_Douga_Movie_Comment_Data_2016Japanese_NicoNico_Douga_Movie_Meta_Data_2016Japanese_NicoNico_Douga_Movie_Meta_Data_2013movie_binaries_0014llm-movielens
LLM-MovieLens
A portable pipeline that turns a catalogue's structured metadata into LLM-synthesized
item features, with the evaluation harness needed to find out whether they help.
Instantiated on two catalogues that share no metadata source: 10,381 films in
MovieLens 20M, and 9,289 books.
Accompanies the ECIR 2027 Resource-track submission by Tan Nghia Duong and Manh Hoang Tran
(School of Electrical and Electronic Engineering, Hanoi University of Science and
Technology).… See the full description on the dataset page: https://huggingface.co/datasets/minastik-ai/llm-movielens.cornell-movie-dialog
Dataset Card for "cornell-movie-dialog"
This is a reduced version of the Cornell Movie Dialog Corpus by Cristian Danescu-Niculescu-Mizil.
The original dataset contains 220,579 conversational exchanges between 10,292 pairs of movie characters, involving 9,035 characters from 617 movies for a total 304,713 utterances.
This reduced version of the dataset contains only the character tags and utterances from the movie_lines.txt file, with one utterance per line, suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/mylesmharrison/cornell-movie-dialog.imdb-movie-genres
Dataset Card for "imdb-movie-genres"
MDb (an acronym for Internet Movie Database) is an online database of information related to films, television programs, home videos, video games, and streaming content online – including cast, production crew and personal biographies, plot summaries, trivia, ratings, and fan and critical reviews. An additional fan feature, message boards, was abandoned in February 2017. Originally a fan-operated website, the database is now owned and operated by… See the full description on the dataset page: https://huggingface.co/datasets/adrienheymans/imdb-movie-genres.movies800time
movies800time
Audio validation set organized as validation/audio/* plus validation/metadata.jsonl. The metadata contains only file_name and transcription; transcriptions include timestamp and speaker markers.
MovieChat-1K_trainmovielens-100kmovie_posters-100k
Dataset Card for "movie_posters-100k"
More Information needed
wiki_moviesThe WikiMovies dataset consists of roughly 100k (templated) questions over 75k entities based on questions with answers in the open movie database (OMDb).MovieChat-1K_trainMovieChat-1K-testmovieswiki-movie-plots-with-summaries
Dataset Card for Wikipedia Movie Plots with AI Plot Summaries
Dataset Summary
Context
Wikipedia Movies Plots dataset by JustinR ( https://www.kaggle.com/jrobischon/wikipedia-movie-plots )
Content
Everything is the same as in https://www.kaggle.com/jrobischon/wikipedia-movie-plots
Acknowledgements
Please, go upvote https://www.kaggle.com/jrobischon/wikipedia-movie-plots dataset, since this is 100% based on that.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/vishnupriyavr/wiki-movie-plots-with-summaries.moviesmovie_gen_video_bench
Dataset Card for the Movie Gen Benchmark
Movie Gen is a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio.
Here, we introduce our evaluation benchmark "Movie Gen Bench Video Bench", as detailed in the Movie Gen technical report (Section 3.5.2).
To enable fair and easy comparison to Movie Gen for future works on these evaluation benchmarks, we additionally release the non cherry-picked generated videos from… See the full description on the dataset page: https://huggingface.co/datasets/meta-ai-for-media-research/movie_gen_video_bench.movie_rationalesThe movie rationale dataset contains human annotated rationales for movie
reviews.movie_mp4movielens-25m-thumb
🍿 Popcorn Thumbnails Embeddings
This dataset contains deep visual features obtained from +65000 movie thumbnails.
It contains extracted visual features using modern VLMs.
To simply load it, Popcorn framework has been developed that can be used in movie recommendation, information retrieval, classification, etc tasks.
📚 Citation
@article{popcorn,
title={Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation},
author={Tourani… See the full description on the dataset page: https://huggingface.co/datasets/alitourani/movielens-25m-thumb.full_tmdb_movies_dataset
