datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-dataset-french-splitprocessed_sroie_donut_dataset_train_test_split
Dataset Card for "processed_sroie_donut_dataset_train_test_split"
More Information needed
tripadvisor-split-dataset-v2
TripAdvisor Review Rating Split Dataset
This dataset contains 80,000 TripAdvisor reviews with corresponding ratings. It is derived from the original TripAdvisor dataset available here and was created to train different models for a university project in the class of NLP.
Dataset Structure
Training Set: 30,400 examples
Validation Set: 1,600 examples
Test Set: 8,000 examples
Each set is balanced, ensuring equal representation of all sentiment labels.
Label
The… See the full description on the dataset page: https://huggingface.co/datasets/nhull/tripadvisor-split-dataset-v2.tripadvisor-split-dataset
New Version Available
A newer version of this dataset with improved annotations and additional examples is available here.
StreetView-Image-Dataset-10K-train-test-splitmorin-khuur-split-datasetfunction_dataset_final_splitShuffled-split-datasetfake_news_detection_dataset_cross_lingual_formatted_uncased_splitmoroccan-darija-asr-dataset-splitSplit_Port_Ship_Classification_Dataset_cococatalan-dataset-phonemized-splitneural-bridge_rag-dataset-12000-ShareGPT-splitncar-ocr-dataset5-splitncar-ocr-dataset10-splithuman-like-sft-dataset-splitbharatanatyam-mudra-dataset-splitcleaned-quora-dataset-train-test-splitThis is a cleaned version of the Quora dataset that's been configured with a train-test-val split.
Train : For training model
Test : For running experiments and comparing different OSS models and closed sourced models
Val : Only to be used at the end!
Colab Notebook to reproduce : https://colab.research.google.com/drive/1dGjGiqwPV1M7JOLfcPEsSh3SC37urItS?usp=sharing
FINGPT_QA_V23-split-datasetllm-japanese-dataset-split_10FINGPT_QA_V31-split-datasetslurp_clustered_split_dataset_fold1cti-dataset-split#these dictionary are useful for this dataset
pos_2_id = {'#': 0, '$': 1, "''": 2, '(': 3, ')': 4, '.': 5, ':': 6, 'CC': 7, 'CD': 8, 'DT': 9, 'EX': 10, 'FW': 11, 'IN': 12, 'JJ': 13, 'JJR': 14, 'JJS': 15, 'MD': 16, 'NN': 17, 'NNP': 18, 'NNPS': 19, 'NNS': 20, 'PDT': 21, 'POS': 22, 'PRP': 23, 'PRP$': 24, 'RB': 25, 'RBR': 26, 'RBS': 27, 'RP': 28, 'TO': 29, 'VB': 30, 'VBD': 31, 'VBG': 32, 'VBN': 33, 'VBP': 34, 'VBZ': 35, 'WDT': 36, 'WP': 37, 'WP$': 38, 'WRB': 39}
id_2_pos = {0: '#', 1: '$', 2: "''"… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/cti-dataset-split.punc_dataset_splithappychat-dataset-half-split
Dataset Card for "happychat-dataset-half-split"
More Information needed
ncar-ocr-dataset7-splitmy_split_datasetsplit_table-formula-dataset-augmentedpsychology-dataset-splitFINGPT_QA_V222-split-dataset
