datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aac_c4_deberta_classified_0.90This dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
This is a subset of the dataset figmtu/aac_c4_deberta_classified.
It contains only the sentences that had a dialogue or forum probability of 0.90 or greater.
See our EMNLP 2025 paper for details.
aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
wikitext-tags-deberta-basewikitext-tags-deberta-v3DeBERTa_multi-class_cb_datasetwsd_UFSAC_deberta_v3_largeaac_subtitle_deberta_classifiedThis dataset contains sentences from the OpenSubtitles2016 movie subtitle corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
aac_c4_deberta_classified_0.90_small_4mThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
This is a subset of the dataset figmtu/aac_c4_deberta_classified.
It contains only the sentences that had a dialogue or forum probability of 0.90 or greater.
This dataset is further limited to only 4M training examples for use in hyperparameter tuning.
See our EMNLP 2025 paper for details.
test_data_deberta_v3_large_npretest_data_deberta_v3_large_raceDeberta_results_racebbq_deberta_v3_large_custom_dataset_custom_headDeberta_results_race_new_input_format_2deberta-base-pii-300krs_deberta_faithful_summary_unannotatedbbq_deberta_v3_large_race_custom_loss_less_adapter_categories_predictionsc_corpus_br_finetuning_language_model_deberta
Dataset Card for "c_corpus_br_finetuning_language_model_deberta"
More Information needed
deberta_rmbbq_deberta_v3_large_race_custom_loss_less_data_predictionsDeberta_results_race_new_input_formatbbq_deberta_v3_large_race_custom_loss_race_format_predictionsbbq_deberta_v3_large_race_finetuned_predictionsdemo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2This is a demo constructed dataset for alignment/preference learning.
With paritially handcrafted questions (prompts), the answers are genreated by the phi-2 model with temperature 0.2 and the answers are scores select by the deberta-large-v2.
The dataset containing questions and the selected answers from highest to lowest, decoding with rejection sampling K=8.
Example loading:
import datasets
ds = datasets.load_dataset('yizhilll/demo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2')… See the full description on the dataset page: https://huggingface.co/datasets/yizhilll/demo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2.test_data_deberta_v3_large_raceclimbmix1k-deberta-v3-smallbbq_deberta_v3_large_race_custom_loss_custom_datasetbbq_deberta_v3_large_race_custom_loss_lamda_07_predictionsbbq_deberta_v3_large_race_custom_loss_predictionssquad_v2_with_answerable_with_debertav3_logitsdeberta_v3_large_race_custom_loss_our_dataset_predictions
