CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /wiki_splitOne million English sentences, each split into two sentences that together preserve the original meaning, extracted from Wikipedia Google's WikiSplit dataset was constructed automatically from the publicly available Wikipedia revision history. Although the dataset contains some inherent noise, it can serve as valuable training data for models that split or merge sentences.100K<n<1M4 likes475 downloads3y agoHugging Face02AdrienB134 /Emilia-dataset-french-splitaudio100K<n<1M4 likes422 downloads2y agoHugging Face03ivelin /processed_sroie_donut_dataset_train_test_split Dataset Card for "processed_sroie_donut_dataset_train_test_split" More Information needed textn<1K0 likes313 downloads4y agoHugging Face04nhull /tripadvisor-split-dataset-v2 TripAdvisor Review Rating Split Dataset This dataset contains 80,000 TripAdvisor reviews with corresponding ratings. It is derived from the original TripAdvisor dataset available here and was created to train different models for a university project in the class of NLP. Dataset Structure Training Set: 30,400 examples Validation Set: 1,600 examples Test Set: 8,000 examples Each set is balanced, ensuring equal representation of all sentiment labels. Label The… See the full description on the dataset page: https://huggingface.co/datasets/nhull/tripadvisor-split-dataset-v2.texttext-classification10K<n<100K2 likes237 downloads2y agoHugging Face05nhull /tripadvisor-split-dataset New Version Available A newer version of this dataset with improved annotations and additional examples is available here. tabular10K<n<100K1 likes136 downloads2y agoHugging Face06ryanhoangt /robocasa365_datasets_split_target_source_human0 likes100 downloads6mo agoHugging Face07Sadhana-24 /StreetView-Image-Dataset-10K-train-test-splitimage1K<n<10K0 likes91 downloads11mo agoHugging Face08Sai08 /morin-khuur-split-datasetaudion<1K0 likes70 downloads1y agoHugging Face09Xtest /function_dataset_final_splittext10K<n<100K0 likes54 downloads2y agoHugging Face10presencesw /dataset_remove_split_0imagen<1K0 likes52 downloads2y agoHugging Face11Rosany /catalan-dataset-phonemized-splittext10K<n<100K0 likes44 downloads2y agoHugging Face12Erland /fake_news_detection_dataset_cross_lingual_formatted_uncased_splittext1K<n<10K0 likes44 downloads2y agoHugging Face13PJMixers-Dev /neural-bridge_rag-dataset-12000-ShareGPT-splittext10K<n<100K0 likes43 downloads2y agoHugging Face14ARG-NCTU /Split_Port_Ship_Classification_Dataset_cocoimage10K<n<100K0 likes43 downloads10mo agoHugging Face15Shreshthh /Shuffled-split-datasettext10K<n<100K0 likes43 downloads9mo agoHugging Face16kazeric /Giriama_bible_dataset_no_splitaudio1K<n<10K0 likes42 downloads1y agoHugging Face17mohamedmou /moroccan-darija-asr-dataset-splitaudio10K<n<100K1 likes41 downloads5mo agoHugging Face18Abdalrahmankamel /ncar-ocr-dataset5-splitimage10K<n<100K0 likes40 downloads6mo agoHugging Face19caster97 /slurp_clustered_split_dataset_fold1audio100K<n<1M0 likes34 downloads1y agoHugging Face20Samarth0710 /bharatanatyam-mudra-dataset-splitimage10K<n<100K0 likes34 downloads1y agoHugging Face21567-labs /cleaned-quora-dataset-train-test-splitThis is a cleaned version of the Quora dataset that's been configured with a train-test-val split. Train : For training model Test : For running experiments and comparing different OSS models and closed sourced models Val : Only to be used at the end! Colab Notebook to reproduce : https://colab.research.google.com/drive/1dGjGiqwPV1M7JOLfcPEsSh3SC37urItS?usp=sharing text100K<n<1M0 likes31 downloads3y agoHugging Face22Abdalrahmankamel /ncar-ocr-dataset10-splitimage1K<n<10K0 likes30 downloads6mo agoHugging Face23booba-uz /punc_dataset_splittext1M<n<10M0 likes27 downloads2y agoHugging Face24DrRiceIO7 /split_datasettext100K<n<1M0 likes27 downloads3mo agoHugging Face25helliun /happychat-dataset-half-split Dataset Card for "happychat-dataset-half-split" More Information needed text1K<n<10K0 likes26 downloads3y agoHugging Face26LayerFault /dataset-trigger-split-columns dataset-trigger-split-columns SECURITY TEST ARTIFACT: DO NOT USE AS A PRODUCTION MODEL This repository is part of the Layerfault synthetic security corpus. It is deliberately constructed to contain security-relevant characteristics for scanner testing. Corpus ID: LF-CH-DATA-0010 Purpose Dataset trigger split columns. Direct expected Layerfault rules None; this repository is a control/comparison input. Candidate rules These are… See the full description on the dataset page: https://huggingface.co/datasets/LayerFault/dataset-trigger-split-columns.0 likes26 downloads1mo agoHugging Face27thangvip /cti-dataset-split#these dictionary are useful for this dataset pos_2_id = {'#': 0, '$': 1, "''": 2, '(': 3, ')': 4, '.': 5, ':': 6, 'CC': 7, 'CD': 8, 'DT': 9, 'EX': 10, 'FW': 11, 'IN': 12, 'JJ': 13, 'JJR': 14, 'JJS': 15, 'MD': 16, 'NN': 17, 'NNP': 18, 'NNPS': 19, 'NNS': 20, 'PDT': 21, 'POS': 22, 'PRP': 23, 'PRP$': 24, 'RB': 25, 'RBR': 26, 'RBS': 27, 'RP': 28, 'TO': 29, 'VB': 30, 'VBD': 31, 'VBG': 32, 'VBN': 33, 'VBP': 34, 'VBZ': 35, 'WDT': 36, 'WP': 37, 'WP$': 38, 'WRB': 39} id_2_pos = {0: '#', 1: '$', 2: "''"… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/cti-dataset-split.text10K<n<100K0 likes25 downloads3y agoHugging Face28longquan /llm-japanese-dataset-split_10textquestion-answering100K<n<1M3 likes24 downloads3y agoHugging Face29KhalfounMehdi /mura_dataset_processed_224px_split Dataset Card for "mura_dataset_processed_224px_split" More Information needed image10K<n<100K0 likes24 downloads3y agoHugging Face30ishika /aloha_play_dataset_part_3_with_fk_full_splitvideon<1K0 likes24 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.