datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imdb
Dataset Card for "imdb"
Dataset Summary
Large Movie Review Dataset.
This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/imdb.imdb
ImdbClassification
An MTEB dataset
Massive Text Embedding Benchmark
Large Movie Review Dataset
Task category
t2c
Domains
Reviews, Written
Reference
http://www.aclweb.org/anthology/P11-1015
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["ImdbClassification"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more… See the full description on the dataset page: https://huggingface.co/datasets/mteb/imdb.imdb-wikiimdb-posters-and-description-512imdb-movie-genres
Dataset Card for "imdb-movie-genres"
MDb (an acronym for Internet Movie Database) is an online database of information related to films, television programs, home videos, video games, and streaming content online – including cast, production crew and personal biographies, plot summaries, trivia, ratings, and fan and critical reviews. An additional fan feature, message boards, was abandoned in February 2017. Originally a fan-operated website, the database is now owned and operated by… See the full description on the dataset page: https://huggingface.co/datasets/adrienheymans/imdb-movie-genres.imdb_faces_age_gender_name_256imdb_wiki_facesimdb_pt
Dataset Card for "imdb_pt"
More Information needed
MM-IMDbIMDBimdb_urdu_reviews
Dataset Card for ImDB Urdu Reviews
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
sentence: The movie review which was translated into Urdu.
sentiment: The sentiment exhibited in the review, either positive or negative.
Data Splits
[More… See the full description on the dataset page: https://huggingface.co/datasets/mirfan899/imdb_urdu_reviews.rl-lm-imdb-promptsimdb_clusteringimdb
Dataset Card for "imdb"
More Information needed
llm-eval-imdbimdb-truncated
Dataset Card for "imdb-truncated"
More Information needed
imdb_sentiment_finetune_dataset20ktask284_imdb_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task284_imdb_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task284_imdb_classification.IMDB-testImdbtask285_imdb_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task285_imdb_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task285_imdb_answer_generation.test_imdb_embedd2
Dataset Card for "test_imdb_embedd2"
More Information needed
imdb
Dataset Card for "imdb"
Dataset Summary
Large Movie Review Dataset.
This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/anondodawan/imdb.test_imdb_embedd
Dataset Card for "test_imdb_embedd"
More Information needed
imdb-posters-and-description-256imdb-ind-classification
IMDB_ind_Classification
Deduplicated copy of kornwtp/imdb-ind-classification.
Splits
split
rows
test
24,800
train
24,902
unsupervised
49,505
imdb
Dataset Card for "imdb"
Dataset Summary
Large Movie Review Dataset.
This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/hassanz123/imdb.imdb_dataset_offical_tripletIMDBSentiment25000_eval_data_imdb
