datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-supervised-datasett5_large_supervised_proportional_1MThis data set is created by randomly sampling 1M documents from the large supervised proportional mixture from the T5 repository.
The code to produce this sampled dataset can be found here.
ViNLI-SimCSE-supervisedViNLI-Healthcare-supervisedViNLI-Zalo-supervisedViNLI-SimCSE-supervised_v2airbnb-reviews-supervisedSupervisedFine-Tuning-unrestricted
Introfuction
This is a Supervised Fine-Tuning dataset.Filtering out common rejection logic, legal statements, moralizing, and other uncomfortable elements found in generative AI.
If you need a pre-trained dataset, please go to:https://huggingface.co/datasets/Zhaoming213/Pretrain-unrestricted
Filter keywords
keywords_list = [
"我无法回答", "我无法给出", "我无法提供", "我不能提供", "我拒绝提供",
"我不具备", "我不拥有", "作为一个AI", "作为一个 AI ", "作为AI",
"作为语言", "作为大语言", "作为程序", "作为一款", "我没有个人"… See the full description on the dataset page: https://huggingface.co/datasets/Zhaoming213/SupervisedFine-Tuning-unrestricted.airbnb-reviews-supervised-improvements
