datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cycic_classificationhttps://storage.googleapis.com/ai2-mosaic/public/cycic/CycIC-train-dev.zip
https://colab.research.google.com/drive/16nyxZPS7-ZDFwp7tn_q72Jxyv0dzK1MP?usp=sharing
@article{Kejriwal2020DoFC,
title={Do Fine-tuned Commonsense Language Models Really Generalize?},
author={Mayank Kejriwal and Ke Shen},
journal={ArXiv},
year={2020},
volume={abs/2011.09159}
}
added for
@article{sileo2023tasksource,
title={tasksource: Structured Dataset Preprocessing Annotations for Frictionless Extreme… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/cycic_classification.Vietnamese-toxic-classificationPatent_classification_QAVietnamese-toxic-classificationText_classification_by_subject_area
🇰🇿 Kazakh Topic and Domain Identification Dataset
Dataset Summary
Kazakh Topic and Domain Identification Dataset is a Kazakh-language instruction-following dataset designed for topic recognition, domain classification, and text understanding tasks.
Each sample contains a short Kazakh prompt, a long Kazakh text passage, a target response, a domain label, and a unique sample identifier. The dataset is intended to help Large Language Models (LLMs) and NLP systems… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Text_classification_by_subject_area.
