datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imdb-ciimdb-genres
Dataset Card for IMDb Movie Dataset: All Movies by Genre
Dataset Summary
This dataset is an adapted version of "IMDb Movie Dataset: All Movies by Genre" found at: https://www.kaggle.com/datasets/rajugc/imdb-movies-dataset-based-on-genre?select=history.csv.
Within the dataset, the movie title and year columns were combined, the genre was extracted from the seperate csv files, the pre-existing genre column was renamed to expanded-genres, any movies missing a description… See the full description on the dataset page: https://huggingface.co/datasets/jquigl/imdb-genres.imdbThis is the sentiment analysis dataset based on IMDB reviews initially released by Stanford University.
This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets.
We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well.
Raw text and already processed bag of words formats are provided. See the README file contained in the release for more… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/imdb.TMDB-IMDB-Movies-Datasetcounterfactually-augmented-imdb@article{kaushik2020learning,
title={Learning the Difference that Makes a Difference with Counterfactually Augmented Data},
author={Kaushik, Divyansh and Hovy, Eduard and Lipton, Zachary C},
journal={International Conference on Learning Representations (ICLR)},
year={2020}
}
IMDb_movie_reviews
Dataset Card for IMDb Movie Reviews
Dataset Summary
This is a custom train/test/validation split of the IMDb Large Movie Review Dataset available from http://ai.stanford.edu/~amaas/data/sentiment/.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
IMDb_movie_reviews
An example of 'train':
{
"text": "Beautifully photographed and ably acted, generally, but the… See the full description on the dataset page: https://huggingface.co/datasets/jahjinx/IMDb_movie_reviews.IMDb-Media
Dataset Card for "BrightData/IMDb-Media"
Dataset Summary
Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more.
For a complete list of data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/IMDb-Media.imdb-movie-reviewshcad_imdbIMDB_MoviesIMDB-Dataset-of-50K-Movie-Reviews-Backupimdb_trainimdb_ckb
IMDB Kurdish (imdb_ckb)
Central Kurdish (Sorani) movie-review sentiment: 49,595 reviews labelled positive
or negative, translated from the Stanford IMDB review set. Balanced labels, so a
0.50 accuracy baseline is meaningless — report F1.
At a glance
Rows
49,595 — train 24,903 / test 24,692
Columns
text (string), label (0 = negative, 1 = positive)
Files
train.csv (59.4 MB), test.csv (57.5 MB) — CSV, not parquet
Language
Central Kurdish / Sorani… See the full description on the dataset page: https://huggingface.co/datasets/razhan/imdb_ckb.imdb-tlmml-interview-examples-mm-imdbEDA_on_IMDB_Movies_Datasetasude-imdb-tr-sonIMDb_Top_250_MoviesIMDb-Movie-Reviews-Sentiment-DatasetIMDB-Dataset-of-50K-Movie-Reviews-Backupimdb-poisoned-50-bddr-word-deletion-defenseimdb-poisoned-25-bddr-word-deletion-defenseimdb_review_3000imdb_multiple_genresIMDB-SAMPLEDimdb_top_1000ImdbMovieDataSetThe IMDB dataset is a rich collection of movie-related information, including:
• Movie details: Titles (original and localized), release dates, and production status
• Ratings & popularity: User ratings and overall reception
• Genres & themes: Categorization of movies by type and style
• Summaries: Overviews describing movie plots
• Cast & crew: Information on actors, directors, writers, and other contributors
• Financial data: Budgets, revenues, and country of origin… See the full description on the dataset page: https://huggingface.co/datasets/ExecuteAutomation/ImdbMovieDataSet.synthetic-imdb-movie-reviews-parallelspanish_imdb_synopsis
Dataset Card for Spanish IMDb Synopsis
Dataset Description
4969 movie synopsis from IMDb in spanish.
Dataset Summary
[N/A]
Languages
All descriptions are in spanish, the other fields have some mix of spanish and english.
Dataset Structure
[N/A]
Data Fields
description: IMDb description for the movie (string), should be spanish
keywords: IMDb keywords for the movie (string), mix of spanish and english
genre: The genres of the… See the full description on the dataset page: https://huggingface.co/datasets/mathigatti/spanish_imdb_synopsis.imdb-sample-250
imdb-sample-250
This is a sample of imdb-sample-250.
