datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ign_clean_instruct_dataset_500kThis dataset contains ~508k prompt-instruction pairs with high quality responses. It was synthetically created from a subset of Ultrachat prompts. It does not contain any alignment focused responses or NSFW content.
Licensed under apache-2.0
PMP_QA_dataset_not_cleanPrint ("This dataset includes 580 Q/A records, not separated, not cleaned yet.-I am working to clean it up-therefore I'm not sharing it publicly.")
