deberta
Datasets
All datasets matching “deberta”aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
aac_c4_deberta_classified_0.90This dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
This is a subset of the dataset figmtu/aac_c4_deberta_classified.
It contains only the sentences that had a dialogue or forum probability of 0.90 or greater.
See our EMNLP 2025 paper for details.
wikitext-tags-deberta-basewikitext-tags-deberta-v3DeBERTa_multi-class_cb_datasetaac_subtitle_deberta_classifiedThis dataset contains sentences from the OpenSubtitles2016 movie subtitle corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
