ayush-shunyalabs/translate-low-resource
Translation Dataset - Low Resource Indian Languages Parallel translation datasets for 50 Indian languages, generated using GPT-5-mini for NLLB-200 finetuning. Dataset Details Total configs: 239 Examples per config: ~4,000 Total examples: ~956,000 Languages: 50 Indian languages across Indo-Aryan, Dravidian, Austroasiatic, and Sino-Tibetan families Hub languages: English, Hindi, Bengali, Tamil, Odia, Assamese Usage Language Codes… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/translate-low-resource.
This repository belongs to ayush-shunyalabs on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
