datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full_tmdb_movies_datasetdouban_movie_reviewTMDB-IMDB-Movies-Datasetmovie_reviews_with_context_drift
Dataset Card for reviews_with_drift
Dataset Description
Dataset Summary
This dataset was crafted to be used in our tutorial [Link to the tutorial when ready]. It consists on a large Movie Review Dataset mixed with some reviews from a Hotel Review Dataset. The training/validation set are purely obtained from the Movie Review Dataset while the production set is mixed. Some other features have been added (age, gender, context) as well as a made up timestamp… See the full description on the dataset page: https://huggingface.co/datasets/arize-ai/movie_reviews_with_context_drift.reddit_movie_large_v1
Dataset Card for Reddit-Movie-large-V1
Dataset Summary
This dataset contains the recommendation-related conversations in movie domain, only for research use in e.g., conversational recommendation, long-query retrieval tasks.
This dataset is ranging from Jan. 2012 to Dec. 2022. Another smaller version dataset (from Jan. 2022 to Dec. 2022) can be found here.
Dataset Processing
We dump Reddit conversations from pushshift.io, converted them into raw text on Reddit… See the full description on the dataset page: https://huggingface.co/datasets/ZhankuiHe/reddit_movie_large_v1.moviesreddit_movie_small_v1
Dataset Card for Reddit-Movie-small-V1
Dataset Summary
This dataset contains the recommendation-related conversations in movie domain, only for research use in e.g., conversational recommendation, long-query retrieval tasks.
This dataset is ranging from Jan. 2022 to Dec. 2022. Another larger version dataset (from Jan. 2012 to Dec. 2022) can be found here.
Dataset Processing
We dump Reddit conversations from pushshift.io, converted them into raw text on Reddit… See the full description on the dataset page: https://huggingface.co/datasets/ZhankuiHe/reddit_movie_small_v1.pixar_movies
Pixar Movies Dataset
A comprehensive dataset of Pixar movies, including details on their release dates, directors, cast, box office performance, and ratings. This dataset is gathered from official sources, including Pixar, Rotten Tomatoes, and IMDb. For more information, visit Pixar.
How the Data is Compiled
All information in this dataset has been collected from public sources, including official information from Pixar, Rotten Tomatoes, and IMDb. Cells are each… See the full description on the dataset page: https://huggingface.co/datasets/RummageLabs/pixar_movies.TMDB-all-moviesdouban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
ml-interview-examples-movielens-1mtmdb_5000_movies.csvTMDB 5000 Movie Dataset
Original source: https://www.kaggle.com/datasets/tmdb/tmdb-movie-metadata
IMDB_MoviesMovielensLatest_x1
MovielensLatest_x1
The MovieLens dataset consists of users' tagging records on movies. The task is formulated as personalized tag recommendation with each tagging record (user_id, item_id, tag_id) as an data instance. The target value denotes whether the user has assigned a particular tag to the movie. We provide the reusable, processed dataset released by the BARS benchmark, which are randomly split into 7:2:1 as the training set, validation set, and test set, respectively.… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/MovielensLatest_x1.spam-douban-movie-review
Description
The Spam Douban Movie Reviews Dataset is a collection of movie reviews scraped from Douban, a popular Chinese social networking platform for movie enthusiasts. This dataset consists of reviews that have been manually classified as either spam or genuine by human reviewers. It contains a total of 1,600 data.
This dataset is created for our project Spam Movie Reviews Detection through Supervised Learning.
douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
EDA_on_IMDB_Movies_Datasetletterboxd-movies
Letterboxd Movies Dataset
Dataset Description
A comprehensive dataset of movies scraped from Letterboxd, including genres, ratings, runtime, countries, and detailed movie characteristics.
This dataset contains 16246 movies with 28 features each, scraped from Letterboxd. It's perfect for:
🎬 Movie recommendation systems
📊 Film industry analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Movie discovery algorithms
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/letterboxd-movies.dirty-movie-a53fa0
dirty-movie-a53fa0
Synthetic sensors test data: 35 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/shenmin91/dirty-movie-a53fa0.moviesIMDb_Top_250_MoviesTMDB_movie_dataset_reducedBLUE-Steer-MovieLens100KRec-Gaze-Click-Cursor-Eye-Tracking-Movie-Recommendation-Dataset-for-Carousel-Interfaces
RecGaze Dataset
This is the HuggingFace RecGaze dataset from the paper: 'RecGaze: The First Eye Tracking and User Interaction Dataset for Carousel Interfaces'.
Link to open-acess paper: SIGIR 2025
Dataset Description
The RecGaze dataset is the first comprehensive feedback dataset on carousels that includes eye tracking results, clicks, cursor movements, and selection explanations. The dataset comprises of interactions from 3 movie selection tasks with 40… See the full description on the dataset page: https://huggingface.co/datasets/santideleon/Rec-Gaze-Click-Cursor-Eye-Tracking-Movie-Recommendation-Dataset-for-Carousel-Interfaces.movie-metadataanime-dataset-2025
Anime Dataset 2025
Description
This dataset contains anime metadata used for machine learning and recommendation systems.
Splits
train
test
Columns
Includes anime title, genres, score, members and other metadata.
Use cases
Anime recommendation systems
NLP tasks
Machine learning projects
License
CC-BY-4.0
movie-ratings
Movie Ratings Database
154,965 movies with ratings, vote counts, release dates, languages, genres, and runtime.
Source
DropThe.org — Data platform tracking 209K+ movies.
Analysis
209K Movies Ratings Analysis
191K Movies Feelgood Score
Links
DropThe.org
Movie Statistics
Methodology
Movie_Datasetmoviedbahsanaseer_top-rated-tmdb-movies-10k
TMDB Movies Dataset
Dataset of 10k top rated TMDB movies for text preprocessing (NLP)
Dataset Info
Source: Kaggle
Original Size: 1.43 MB
Kaggle Downloads: 8,035
Files: 1
Files
top10K-TMDB-movies.csv
Mirrored from Kaggle
