Mayank6255/translated_dolly_spec_decode
Aya Multilingual SFT Dataset This dataset is derived from the CohereForAI/aya_collection (translated_dolly subset) and formatted for supervised fine-tuning (SFT) with LLaMA-Factory. Dataset Subsets Subset Languages Train Test Description eng ENG ✓ ✓ English only subset hin HIN ✓ ✓ Hindi only subset deu DEU ✓ ✓ German only subset eng_hin_deu ENG, HIN, DEU ✓ ✓ Combined English, Hindi, and German subset eng_hin_deu_sampled ENG, HIN, DEU, SAMPLED ✓… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/translated_dolly_spec_decode.
028
