CoolFace
Datasetpublic

pulipakav-1/translated-babylm-telugu

Translated BabyLM — Telugu (translated-babylm-telugu) Dataset Description This dataset is a Telugu translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Telugu, following the BabyLM challenge setup. Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B) Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-telugu.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes51downloads

pulipakav-1/translated-babylm-telugu · main · files are served by the source, never re-hosted here