CoolFace
Datasetpublic

Mayank6255/translated_dolly_spec_decode

Aya Multilingual SFT Dataset This dataset is derived from the CohereForAI/aya_collection (translated_dolly subset) and formatted for supervised fine-tuning (SFT) with LLaMA-Factory. Dataset Subsets Subset Languages Train Test Description eng ENG ✓ ✓ English only subset hin HIN ✓ ✓ Hindi only subset deu DEU ✓ ✓ German only subset eng_hin_deu ENG, HIN, DEU ✓ ✓ Combined English, Hindi, and German subset eng_hin_deu_sampled ENG, HIN, DEU, SAMPLED ✓… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/translated_dolly_spec_decode.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes28downloads

Mayank6255/translated_dolly_spec_decode · main · files are served by the source, never re-hosted here