datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IMDb-Media
Dataset Card for "BrightData/IMDb-Media"
Dataset Summary
Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more.
For a complete list of data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/IMDb-Media.task284_imdb_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task284_imdb_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task284_imdb_classification.task285_imdb_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task285_imdb_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task285_imdb_answer_generation.spanish_imdb_synopsis
Dataset Card for Spanish IMDb Synopsis
Dataset Description
4969 movie synopsis from IMDb in spanish.
Dataset Summary
[N/A]
Languages
All descriptions are in spanish, the other fields have some mix of spanish and english.
Dataset Structure
[N/A]
Data Fields
description: IMDb description for the movie (string), should be spanish
keywords: IMDb keywords for the movie (string), mix of spanish and english
genre: The genres of the… See the full description on the dataset page: https://huggingface.co/datasets/mathigatti/spanish_imdb_synopsis.IMDB_Sentiment
Dataset Card for "imdb"
Dataset Summary
Large Movie Review Dataset.
This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well.
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset files: 84.13 MB
Size of the… See the full description on the dataset page: https://huggingface.co/datasets/Kwaai/IMDB_Sentiment.imdb
Dataset Card for "imdb"
Dataset Summary
Large Movie Review Dataset.
This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well.
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset files: 84.13 MB
Size of the… See the full description on the dataset page: https://huggingface.co/datasets/pt-sk/imdb.IMDb-Media
Dataset Card for "BrightData/IMDb-Media"
Dataset Summary
Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more.
For a complete list of… See the full description on the dataset page: https://huggingface.co/datasets/RyanHalliwell/IMDb-Media.imdb_rewardedThis is the imdb dataset, https://huggingface.co/datasets/imdb
We've used a reward / sentiment model, https://huggingface.co/lvwerra/distilbert-imdb to compute the rewards of the offline data.
This is so that we can use offline RL on the data.
