waashk/yelp_2013
Dataset used in the paper: A thorough benchmark of automatic text classification From traditional approaches to large language models https://github.com/waashk/atcBench To guarantee the reproducibility of the obtained results, the dataset and its respective CV train-test partitions is available here. Each dataset contains the following files: data.parquet: pandas DataFrame with texts and associated encoded labels for each document. split_<k>.pkl: pandas DataFrame with k-cross validation… See the full description on the dataset page: https://huggingface.co/datasets/waashk/yelp_2013.
077
data.parquetdownload
split_5.pkldownload
test_fold_0.parquetdownload
test_fold_1.parquetdownload
test_fold_2.parquetdownload
test_fold_3.parquetdownload
test_fold_4.parquetdownload
train_fold_0.parquetdownload
train_fold_1.parquetdownload
train_fold_2.parquetdownload
train_fold_3.parquetdownload
train_fold_4.parquetdownload
