datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imdbimdb-movie-reviews
IMDB Movie Reviews
This is a dataset for binary sentiment classification containing substantially huge data. This dataset contains a set of 50,000 highly polar movie reviews for training models for text classification tasks.
The dataset is downloaded from
https://ai.stanford.edu/~amaas/data/sentiment/aclImdb_v1.tar.gz
This data is processed and splitted into training and test datasets (0.2% test split). Training dataset contains 40000 reviews and test dataset contains 10000… See the full description on the dataset page: https://huggingface.co/datasets/ajaykarthick/imdb-movie-reviews.imdbimdbimdb-movie-reviews
IMDB Movie Reviews
This is a dataset for binary sentiment classification containing substantially huge data. This dataset contains a set of 50,000 highly polar movie reviews for training models for text classification tasks.
The dataset is downloaded from
https://ai.stanford.edu/~amaas/data/sentiment/aclImdb_v1.tar.gz
This data is processed and splitted into training and test datasets (0.2% test split). Training dataset contains 40000 reviews and test dataset contains 10000… See the full description on the dataset page: https://huggingface.co/datasets/puneet44/imdb-movie-reviews.imdb_filteredhttps://archive.ics.uci.edu/dataset/331/sentiment+labelled+sentences
This dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015
Please cite the paper if you want to use it :)
It contains sentences labelled with positive or negative sentiment, extracted from reviews of products, movies, and restaurants
=======
Format:
sentence \t score \n
=======
Details:
Score is either 1 (for positive) or 0 (for negative)
The… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/imdb_filtered.imdb-dutch-instruct
Dataset Card for "imdb-dutch-instruct"
Dataset Description
The original IMBD dataset was translated to Dutch with yhavinga/ul2-large-en-nl.
Then, the dataset is converted to an instruct-style dataset with the following templates:
The instruction templates:
"Is deze recensie positief of negatief?",
"Wat is het sentiment van de recensie?",
"Wat voor toon heeft de volgende recensie?",
"Met wat voor sentiment zou je deze recensie beoordelen?"
The target templates:
"De… See the full description on the dataset page: https://huggingface.co/datasets/jjzha/imdb-dutch-instruct.imdb_rewardedThis is the imdb dataset, https://huggingface.co/datasets/imdb
We've used a reward / sentiment model, https://huggingface.co/lvwerra/distilbert-imdb to compute the rewards of the offline data.
This is so that we can use offline RL on the data.
imdb_contrastsetimdb_v0imdb-gpt-selftalk_500kimdbsimdb-sentiment-classifier-evalsimdb_reviewsimdb_comparison_feedbackimdb-preference-10kimdbIMDB_top_100imdb-2025-moreimdb-2025
