ayush-shunyalabs/translate-low-resource
Translation Dataset - Low Resource Indian Languages Parallel translation datasets for 50 Indian languages, generated using GPT-5-mini for NLLB-200 finetuning. Dataset Details Total configs: 239 Examples per config: ~4,000 Total examples: ~956,000 Languages: 50 Indian languages across Indo-Aryan, Dravidian, Austroasiatic, and Sino-Tibetan families Hub languages: English, Hindi, Bengali, Tamil, Odia, Assamese Usage Language Codes… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/translate-low-resource.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face