datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PRODIGY-LAB_SARA
Dataset Card for PRODIGY-LAB_CLEANED
Repository: https://github.com/aadhithyaravi
Created by: Aadhithya
Contact: aadhithyaxll@gmail.com
Instagram: @aadhi.arc
LinkedIn: www.linkedin.com/in/aadhithya-ravi-135019289
Dataset Description
PRODIGY-SARA-MODEL is a refined and enhanced dataset designed for instruction-based fine-tuning of large language models (LLMs).It combines multiple high-quality sources, including cleaned and normalized instructions, to improve… See the full description on the dataset page: https://huggingface.co/datasets/Apex-X/PRODIGY-LAB_SARA.prodigy-cleaned
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize the… See the full description on the dataset page: https://huggingface.co/datasets/Apex-X/prodigy-cleaned.
