aplycaebous/BanglaTLit
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla Dataset Overview BanglaTLit-PT: A pre-training corpus with 245727 transliterated or romanized Bangla samples for further pre-training language models. BanglaTLit: Subset of the BanglaTLit-PT dataset containing 42705 romanized Bangla and its corresponding Bangla back-transliteration pairs. Data Description Column Title Description id A unique identifier for… See the full description on the dataset page: https://huggingface.co/datasets/aplycaebous/BanglaTLit.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face