datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki_splitOne million English sentences, each split into two sentences that together preserve the original meaning, extracted from Wikipedia
Google's WikiSplit dataset was constructed automatically from the publicly available Wikipedia revision history. Although
the dataset contains some inherent noise, it can serve as valuable training data for models that split or merge sentences.Emilia-dataset-french-splitprocessed_sroie_donut_dataset_train_test_split
Dataset Card for "processed_sroie_donut_dataset_train_test_split"
More Information needed
tripadvisor-split-dataset-v2
TripAdvisor Review Rating Split Dataset
This dataset contains 80,000 TripAdvisor reviews with corresponding ratings. It is derived from the original TripAdvisor dataset available here and was created to train different models for a university project in the class of NLP.
Dataset Structure
Training Set: 30,400 examples
Validation Set: 1,600 examples
Test Set: 8,000 examples
Each set is balanced, ensuring equal representation of all sentiment labels.
Label
The… See the full description on the dataset page: https://huggingface.co/datasets/nhull/tripadvisor-split-dataset-v2.tripadvisor-split-dataset
New Version Available
A newer version of this dataset with improved annotations and additional examples is available here.
robocasa365_datasets_split_target_source_humanStreetView-Image-Dataset-10K-train-test-splitmorin-khuur-split-datasetfunction_dataset_final_splitdataset_remove_split_0catalan-dataset-phonemized-splitfake_news_detection_dataset_cross_lingual_formatted_uncased_splitneural-bridge_rag-dataset-12000-ShareGPT-splitSplit_Port_Ship_Classification_Dataset_cocoShuffled-split-datasetGiriama_bible_dataset_no_splitmoroccan-darija-asr-dataset-splitncar-ocr-dataset5-splitslurp_clustered_split_dataset_fold1bharatanatyam-mudra-dataset-splitcleaned-quora-dataset-train-test-splitThis is a cleaned version of the Quora dataset that's been configured with a train-test-val split.
Train : For training model
Test : For running experiments and comparing different OSS models and closed sourced models
Val : Only to be used at the end!
Colab Notebook to reproduce : https://colab.research.google.com/drive/1dGjGiqwPV1M7JOLfcPEsSh3SC37urItS?usp=sharing
ncar-ocr-dataset10-splitpunc_dataset_splitsplit_datasethappychat-dataset-half-split
Dataset Card for "happychat-dataset-half-split"
More Information needed
dataset-trigger-split-columns
dataset-trigger-split-columns
SECURITY TEST ARTIFACT: DO NOT USE AS A PRODUCTION MODEL
This repository is part of the Layerfault synthetic security corpus.
It is deliberately constructed to contain security-relevant characteristics for scanner testing.
Corpus ID: LF-CH-DATA-0010
Purpose
Dataset trigger split columns.
Direct expected Layerfault rules
None; this repository is a control/comparison input.
Candidate rules
These are… See the full description on the dataset page: https://huggingface.co/datasets/LayerFault/dataset-trigger-split-columns.cti-dataset-split#these dictionary are useful for this dataset
pos_2_id = {'#': 0, '$': 1, "''": 2, '(': 3, ')': 4, '.': 5, ':': 6, 'CC': 7, 'CD': 8, 'DT': 9, 'EX': 10, 'FW': 11, 'IN': 12, 'JJ': 13, 'JJR': 14, 'JJS': 15, 'MD': 16, 'NN': 17, 'NNP': 18, 'NNPS': 19, 'NNS': 20, 'PDT': 21, 'POS': 22, 'PRP': 23, 'PRP$': 24, 'RB': 25, 'RBR': 26, 'RBS': 27, 'RP': 28, 'TO': 29, 'VB': 30, 'VBD': 31, 'VBG': 32, 'VBN': 33, 'VBP': 34, 'VBZ': 35, 'WDT': 36, 'WP': 37, 'WP$': 38, 'WRB': 39}
id_2_pos = {0: '#', 1: '$', 2: "''"… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/cti-dataset-split.llm-japanese-dataset-split_10mura_dataset_processed_224px_split
Dataset Card for "mura_dataset_processed_224px_split"
More Information needed
aloha_play_dataset_part_3_with_fk_full_split
