split-dataset
wiki_splitOne million English sentences, each split into two sentences that together preserve the original meaning, extracted from Wikipedia
Google's WikiSplit dataset was constructed automatically from the publicly available Wikipedia revision history. Although
the dataset contains some inherent noise, it can serve as valuable training data for models that split or merge sentences.Emilia-dataset-french-splitprocessed_sroie_donut_dataset_train_test_split
Dataset Card for "processed_sroie_donut_dataset_train_test_split"
More Information needed
tripadvisor-split-dataset-v2
TripAdvisor Review Rating Split Dataset
This dataset contains 80,000 TripAdvisor reviews with corresponding ratings. It is derived from the original TripAdvisor dataset available here and was created to train different models for a university project in the class of NLP.
Dataset Structure
Training Set: 30,400 examples
Validation Set: 1,600 examples
Test Set: 8,000 examples
Each set is balanced, ensuring equal representation of all sentiment labels.
Label
The… See the full description on the dataset page: https://huggingface.co/datasets/nhull/tripadvisor-split-dataset-v2.tripadvisor-split-dataset
New Version Available
A newer version of this dataset with improved annotations and additional examples is available here.
robocasa365_datasets_split_target_source_human
