CoolFace
Datasetpublic

ayush-shunyalabs/translate-low-resource

Translation Dataset - Low Resource Indian Languages Parallel translation datasets for 50 Indian languages, generated using GPT-5-mini for NLLB-200 finetuning. Dataset Details Total configs: 239 Examples per config: ~4,000 Total examples: ~956,000 Languages: 50 Indian languages across Indo-Aryan, Dravidian, Austroasiatic, and Sino-Tibetan families Hub languages: English, Hindi, Bengali, Tamil, Odia, Assamese Usage Language Codes… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/translate-low-resource.

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes55downloads
settings

This repository belongs to ayush-shunyalabs on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nametranslate-low-resource
visibilitypublic
licencenot set
gatedno
ownerayush-shunyalabs
Account settings
ayush-shunyalabs/translate-low-resource · CoolFace