CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsgoncalves /roberta_pretrain Dataset Card for RoBERTa Pretrain Dataset Summary This is the concatenation of the datasets used to Pretrain RoBERTa. The dataset is not shuffled and contains raw text. It is packaged for convenicence. Essentially is the same as: from datasets import load_dataset, concatenate_datasets bookcorpus = load_dataset("bookcorpus", split="train") openweb = load_dataset("openwebtext", split="train") cc_news = load_dataset("cc_news", split="train") cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.textfill-mask10M<n<100M5 likes402 downloads3y agoHugging Face02closji /wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largetext1M<n<10M0 likes186 downloads4y agoHugging Face03closji /cc12m_princeton-nlp_sup-simcse-roberta-largeimage10M<n<100M0 likes110 downloads4y agoHugging Face04TannerGladson /chess-roberta-basetabular100M<n<1B0 likes109 downloads2y agoHugging Face05xorushi /roberta-pii-synth Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.texttoken-classification100K<n<1M0 likes84 downloads27d agoHugging Face06Aadithyak /roberta-largastabular100K<n<1M0 likes60 downloads2y agoHugging Face07Warawreh /PII-cleaned-roberta-classes-merged-ignoretext10K<n<100K0 likes57 downloads1y agoHugging Face08CabraVC /vector_dataset_roberta-fine-tunedtext1K<n<10K0 likes46 downloads3y agoHugging Face09oeg /CelebA_RoBERTa_Sp Corpus Summary This corpus contains 250000 entries made up of a pair of sentences in Spanish and their respective similarity value in the range 0 to 1. This corpus was used in the training of the sentence-transformer library to improve the efficiency of the RoBERTa-large-bne base model. Each of the pairs of sentences are textual descriptions of the faces of the CelebA dataset, which were previously translated into Spanish. The process followed to generate it was: First, a… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_RoBERTa_Sp.texttable-question-answering100K<n<1M1 likes42 downloads3y agoHugging Face10spoiled /ecqa_model_generate_robertatext10K<n<100K0 likes40 downloads4y agoHugging Face11closji /mscoco_2014_captions_princeton-nlp_sup-simcse-roberta-largetext100K<n<1M0 likes33 downloads4y agoHugging Face12sunhaozhepy /ag_news_roberta_keywords_embeddingstext100K<n<1M1 likes31 downloads3y agoHugging Face13Shweta-singh /test_data_roberta_base_6_racetabular10K<n<100K0 likes31 downloads2y agoHugging Face14Shweta-singh /test_data_roberta_basetabular10K<n<100K0 likes30 downloads2y agoHugging Face15nikchar /retrieval_verification_roberta Dataset Card for "retrieval_verification_roberta" More Information needed tabular10K<n<100K0 likes29 downloads3y agoHugging Face16insub /imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english Dataset Card for "imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english" 1. Purpose of creating the dataset For reproduction of DPO (direct preference optimization) thesis experiments(https://arxiv.org/abs/2305.18290) 2. How data is produced To reproduce the paper's experimental results, we need (x, chosen, rejected) data.However, imdb data only contains good or bad reviews, so the data must be readjusted. 2.1 prepare imdb… See the full description on the dataset page: https://huggingface.co/datasets/insub/imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english.text10K<n<100K2 likes28 downloads3y agoHugging Face17thibaudltn /twitter_ae_xlm_roberta_sentiment_stratifiedtabular10K<n<100K0 likes26 downloads7mo agoHugging Face18insub /imdb_prefix3_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english Dataset Card for "imdb_prefix3_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english" More Information needed text10K<n<100K1 likes24 downloads3y agoHugging Face19nikchar /retrieval_verification_bm25_roberta Dataset Card for "retrieval_verification_bm25_roberta" More Information needed tabular10K<n<100K0 likes23 downloads3y agoHugging Face20davidfant /rapidapi-example-responses-tokenized-xlm-roberta Dataset Card for "rapidapi-example-responses-tokenized-xlm-roberta" More Information needed text10K<n<100K0 likes23 downloads3y agoHugging Face21EgilKarlsen /CSIC_RoBERTa_FT Dataset Card for "CSIC_RoBERTa_FT" More Information needed tabular10K<n<100K0 likes22 downloads3y agoHugging Face22EgilKarlsen /Thunderbird_RoBERTa_FT Dataset Card for "Thunderbird_RoBERTa_FT" More Information needed tabular10K<n<100K0 likes21 downloads3y agoHugging Face23ejdis /roberta_datasettext100M<n<1B0 likes21 downloads2y agoHugging Face24theekshana /NER_medical_reports_tokenized_deid_roberta_i2b2 Dataset Card for "NER_medical_reports_tokenized_deid_roberta_i2b2" More Information needed text1K<n<10K1 likes21 downloads2y agoHugging Face25merve /parsed-dataset-xlm-robertatextn<1K0 likes20 downloads4y agoHugging Face26EgilKarlsen /PKDD_RoBERTa_FT Dataset Card for "PKDD_RoBERTa_FT" More Information needed tabular10K<n<100K0 likes20 downloads3y agoHugging Face27rajendrabaskota /hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512 Dataset Card for "hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512" More Information needed tabular100K<n<1M0 likes20 downloads3y agoHugging Face28artianand /bbq_roberta_large_race_custom_loss_lamda_07_predictionstabular10K<n<100K0 likes20 downloads1y agoHugging Face29orYx-models /roberta-leadership-dataset-finetunetextn<1K0 likes16 downloads2y agoHugging Face30gguichard /ontonotes_val-roberta-large-v2text10K<n<100K0 likes16 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.