CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eduagarcia-temp /roberta-pt-checkpoints0 likes2.3k downloads3y agoHugging Face02gsgoncalves /roberta_pretrain Dataset Card for RoBERTa Pretrain Dataset Summary This is the concatenation of the datasets used to Pretrain RoBERTa. The dataset is not shuffled and contains raw text. It is packaged for convenicence. Essentially is the same as: from datasets import load_dataset, concatenate_datasets bookcorpus = load_dataset("bookcorpus", split="train") openweb = load_dataset("openwebtext", split="train") cc_news = load_dataset("cc_news", split="train") cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.textfill-mask10M<n<100M5 likes405 downloads3y agoHugging Face03elricwan /roberta-data10M<n<100M0 likes261 downloads5y agoHugging Face04joegolk /roberta-tokenized-data1M<n<10M0 likes219 downloads8mo agoHugging Face05closji /wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largetext1M<n<10M0 likes186 downloads4y agoHugging Face06AnonymousSub /recipe_RL_data_roberta-base Dataset Description Structure Consists of 5 fields Each row corresponds to a policy - sequence of actions, given an initial <START> state, and corresponding rewards at each step. Fields steps, step_attn_masks, rewards, actions, dones Field descriptions steps (List of lists of Ints) - tokenized step tokens of all the steps in the policy sequence (here we use the roberta-base tokenizer, as roberta-base would be used to encode each step of a recipe)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSub/recipe_RL_data_roberta-base.1M<n<10M0 likes134 downloads4y agoHugging Face07tursunait /roberta-pii-synth Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tursunait/roberta-pii-synth.token-classification100K<n<1M0 likes134 downloads10mo agoHugging Face08AnonymousSub /recipe_RL_data_ONLY_CORRECT_roberta-base1M<n<10M0 likes121 downloads4y agoHugging Face09hriaz /wikitext-tags-roberta1M<n<10M0 likes118 downloads1y agoHugging Face10closji /cc12m_princeton-nlp_sup-simcse-roberta-largeimage10M<n<100M0 likes110 downloads4y agoHugging Face11TannerGladson /chess-roberta-basetabular100M<n<1B0 likes109 downloads2y agoHugging Face12emilylearning /cond_ft_none_on_reddit__prcnt_100__test_run_False__xlm-roberta-base1M<n<10M0 likes99 downloads4y agoHugging Face13AnonymousSub /recipe_RL_data_roberta-base_BASE_REWARD_32_STEPDIFF_reward1M<n<10M0 likes91 downloads4y agoHugging Face14emilylearning /cond_ft_subreddit_on_reddit__prcnt_100__test_run_False__xlm-roberta-base1M<n<10M0 likes85 downloads4y agoHugging Face15xorushi /roberta-pii-synth Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.texttoken-classification100K<n<1M0 likes84 downloads27d agoHugging Face16Aadithyak /roberta-largastabular100K<n<1M0 likes61 downloads2y agoHugging Face17emilylearning /cond_ft_none_on_reddit__prcnt_100__test_run_False__roberta-base1M<n<10M0 likes50 downloads4y agoHugging Face18mor40 /tokenized-chitanka-roberta100K<n<1M0 likes48 downloads2y agoHugging Face19peeper /vitmae-roberta-processed Dataset Card for "vitmae-roberta-processed" More Information needed 1K<n<10K0 likes46 downloads4y agoHugging Face20CabraVC /vector_dataset_roberta-fine-tunedtext1K<n<10K0 likes46 downloads3y agoHugging Face21emilylearning /cond_ft_subreddit_on_reddit__prcnt_100__test_run_False__roberta-base1M<n<10M0 likes42 downloads4y agoHugging Face22oeg /CelebA_RoBERTa_Sp Corpus Summary This corpus contains 250000 entries made up of a pair of sentences in Spanish and their respective similarity value in the range 0 to 1. This corpus was used in the training of the sentence-transformer library to improve the efficiency of the RoBERTa-large-bne base model. Each of the pairs of sentences are textual descriptions of the faces of the CelebA dataset, which were previously translated into Spanish. The process followed to generate it was: First, a… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_RoBERTa_Sp.texttable-question-answering100K<n<1M1 likes42 downloads3y agoHugging Face23Rinasution /roberta-retrieval-lite preprocess.py Dataset Summary A news media dataset with audio text modality, stored in arrow format. Preprocessing & Augmentation Preprocessing: domain specific Augmentation: mixup cutmix Splits & Sampling Split strategy: kfold 5 Sampling: random Quality & Labeling Quality filtering: lenient Labeling: pseudo label Files preprocess.py — main artifact of this repository License See… See the full description on the dataset page: https://huggingface.co/datasets/Rinasution/roberta-retrieval-lite.0 likes42 downloads27d agoHugging Face24emilylearning /cond_ft_none_on_reddit__prcnt_na__test_run_True__roberta-base0 likes41 downloads4y agoHugging Face25spoiled /ecqa_model_generate_robertatext10K<n<100K0 likes40 downloads4y agoHugging Face26lucasvandijk /roberta-ocr build_dataset.py Dataset Summary A wildlife dataset with video text modality, stored in hdf5 format. Preprocessing & Augmentation Preprocessing: adaptive Augmentation: mixup cutmix Splits & Sampling Split strategy: random 90 10 Sampling: curriculum Quality & Labeling Quality filtering: moderate Labeling: pseudo label Files build_dataset.py — main artifact of this repository License… See the full description on the dataset page: https://huggingface.co/datasets/lucasvandijk/roberta-ocr.0 likes40 downloads27d agoHugging Face27emilylearning /cond_ft_subreddit_on_reddit__prcnt_na__test_run_True__roberta-base0 likes39 downloads4y agoHugging Face28Danylolysenko /roberta-sentiment-lite preprocess.py Dataset Summary A wildlife dataset with audio video modality, stored in csv format. Preprocessing & Augmentation Preprocessing: progressive Augmentation: autoaugment Splits & Sampling Split strategy: temporal Sampling: random Quality & Labeling Quality filtering: moderate Labeling: weak supervision Files preprocess.py — main artifact of this repository License See the… See the full description on the dataset page: https://huggingface.co/datasets/Danylolysenko/roberta-sentiment-lite.0 likes37 downloads27d agoHugging Face29emilylearning /cond_ft_none_on_reddit__prcnt_na__test_run_True__xlm-roberta-base0 likes36 downloads4y agoHugging Face30Shweta-singh /test_data_roberta_base_6_racetabular10K<n<100K0 likes36 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.