datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
roberta_pretrain
Dataset Card for RoBERTa Pretrain
Dataset Summary
This is the concatenation of the datasets used to Pretrain RoBERTa.
The dataset is not shuffled and contains raw text. It is packaged for convenicence.
Essentially is the same as:
from datasets import load_dataset, concatenate_datasets
bookcorpus = load_dataset("bookcorpus", split="train")
openweb = load_dataset("openwebtext", split="train")
cc_news = load_dataset("cc_news", split="train")
cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largecc12m_princeton-nlp_sup-simcse-roberta-largechess-roberta-baseroberta-pii-synth
Synthetic PII Detection Dataset (RoBERTa-PII-Synth)
A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text.
This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere.
📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.roberta-largasPII-cleaned-roberta-classes-merged-ignorevector_dataset_roberta-fine-tunedCelebA_RoBERTa_Sp
Corpus Summary
This corpus contains 250000 entries made up of a pair of sentences in Spanish and their respective similarity value in the range 0 to 1. This corpus was used in the training of the
sentence-transformer library to improve the efficiency of the RoBERTa-large-bne base model.
Each of the pairs of sentences are textual descriptions of the faces of the CelebA dataset, which were previously translated into Spanish. The process followed to generate it was:
First, a… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_RoBERTa_Sp.ecqa_model_generate_robertamscoco_2014_captions_princeton-nlp_sup-simcse-roberta-largeag_news_roberta_keywords_embeddingstest_data_roberta_base_6_racetest_data_roberta_baseretrieval_verification_roberta
Dataset Card for "retrieval_verification_roberta"
More Information needed
imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english
Dataset Card for "imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english"
1. Purpose of creating the dataset
For reproduction of DPO (direct preference optimization) thesis experiments(https://arxiv.org/abs/2305.18290)
2. How data is produced
To reproduce the paper's experimental results, we need (x, chosen, rejected) data.However, imdb data only contains good or bad reviews, so the data must be readjusted.
2.1 prepare imdb… See the full description on the dataset page: https://huggingface.co/datasets/insub/imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english.twitter_ae_xlm_roberta_sentiment_stratifiedimdb_prefix3_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english
Dataset Card for "imdb_prefix3_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english"
More Information needed
retrieval_verification_bm25_roberta
Dataset Card for "retrieval_verification_bm25_roberta"
More Information needed
rapidapi-example-responses-tokenized-xlm-roberta
Dataset Card for "rapidapi-example-responses-tokenized-xlm-roberta"
More Information needed
CSIC_RoBERTa_FT
Dataset Card for "CSIC_RoBERTa_FT"
More Information needed
Thunderbird_RoBERTa_FT
Dataset Card for "Thunderbird_RoBERTa_FT"
More Information needed
roberta_datasetNER_medical_reports_tokenized_deid_roberta_i2b2
Dataset Card for "NER_medical_reports_tokenized_deid_roberta_i2b2"
More Information needed
parsed-dataset-xlm-robertaPKDD_RoBERTa_FT
Dataset Card for "PKDD_RoBERTa_FT"
More Information needed
hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512
Dataset Card for "hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512"
More Information needed
bbq_roberta_large_race_custom_loss_lamda_07_predictionsroberta-leadership-dataset-finetuneontonotes_val-roberta-large-v2
