CoolFace
Datasetpublic

waashk/yelp_2013

Dataset used in the paper: A thorough benchmark of automatic text classification From traditional approaches to large language models https://github.com/waashk/atcBench To guarantee the reproducibility of the obtained results, the dataset and its respective CV train-test partitions is available here. Each dataset contains the following files: data.parquet: pandas DataFrame with texts and associated encoded labels for each document. split_<k>.pkl: pandas DataFrame with k-cross validation… See the full description on the dataset page: https://huggingface.co/datasets/waashk/yelp_2013.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes77downloads
filedata.parquet138.1 MBdownload
filesplit_5.pkl7.4 MBdownload
filetest_fold_0.parquet27.4 MBdownload
filetest_fold_1.parquet27.6 MBdownload
filetest_fold_2.parquet27.6 MBdownload
filetest_fold_3.parquet27.5 MBdownload
filetest_fold_4.parquet27.4 MBdownload
filetrain_fold_0.parquet110.7 MBdownload
filetrain_fold_1.parquet110.4 MBdownload
filetrain_fold_2.parquet110.5 MBdownload
filetrain_fold_3.parquet110.6 MBdownload
filetrain_fold_4.parquet110.7 MBdownload

waashk/yelp_2013 · main · files are served by the source, never re-hosted here