datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
filtered_yelp_restaurant_reviews
Dataset Card for "filtered_yelp_restaurant_reviews"
More Information needed
yelp-open-dataset-reviewsyelp_restaurant_reviews_5kyelp_review_sampledyelp-open-dataset-top-reviews-per-businessitalian-yelp-reviews-datasetYelp_Reviews_for_Sentiment_Analysis_fine_grained_5_classes
Dataset Card for Dataset Name
The Yelp reviews full star dataset is constructed by randomly taking 130,000 training samples and 10,000 testing samples for each review star from 1 to 5. In total there are 650,000 trainig samples and 50,000 testing samples.
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 2 columns in them, corresponding to class index (1 to 5) and review text. The review texts are… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yelp_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.tokenized-yelp-reviews-distillyelp_restaurant_reviewsThis dataset is a filtered version of https://huggingface.co/datasets/vincha77/filtered_yelp_restaurant_reviews
tokenized-yelp-reviews-distill-smalltokenized-yelp-reviewsYelpReviewsThis dataset contains the Fictional Yelp review.
Two jsonl files (df_neg.jsonl & df_pos.jsonl) can be directly used for OpenAI / OpenPipe finetuning.
(This is what you need for replication)
More detailed data at df_pos_generated and df_neg_generated file. The oos files are out-of-sample cues.
yelp_boba_reviewsYelp_Reviews_for_Binary_Senti_Analysis
Dataset Card for Dataset Name
The Yelp reviews polarity dataset is constructed by considering stars 1 and 2 negative, and 3 and 4 positive. For each polarity 280,000 training samples and 19,000 testing samples are take randomly. In total there are 560,000 trainig samples and 38,000 testing samples. Negative polarity is class 1, and positive class 2.
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 2… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yelp_Reviews_for_Binary_Senti_Analysis.yelp-open-dataset-top-reviewsyelp_reviews_encoded_hidden_outputs_truncatedyelp-open-dataset-reviewsyelp_reviews_2LDataset used in the paper:
A thorough benchmark of automatic text classification
From traditional approaches to large language models
https://github.com/waashk/atcBench
To guarantee the reproducibility of the obtained results, the dataset and its respective CV train-test partitions is available here.
Each dataset contains the following files:
data.parquet: pandas DataFrame with texts and associated encoded labels for each document.
split_<k>.pkl: pandas DataFrame with k-cross validation… See the full description on the dataset page: https://huggingface.co/datasets/waashk/yelp_reviews_2L.clcp_yelpreviewsyelp_reviews_reducedyelp-review-sentiment-subsetyelp-open-dataset-top-reviewsyelpreviewsyelp_reviews_encoded_hidden_outputsflan_combined_yelp_polarity_reviews_0_2_0yelp-reviews-propioyelp-reviewsyelp_ca_reviewsA dataset for the reviews of different restaurents in Santa Barbara using the Yelp Dataset
yelp_reviewsyelpreviews
