datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
romanized_hindi
Romanized Hindi Dataset
Dataset Description
The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation.
It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods.
The dataset is designed for training and evaluating Hindi↔Roman transliteration models.
Language(s): Hindi, Romanized Hindi
Size: ~1.82M rows
License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.telugu_alpaca_yahma_cleaned_filtered_romanizedtelugu_teknium_GPTeacher_general_instruct_filtered_romanizedSwabhasha_RomanizedSinhala_Dataset
Model Card for Model ID
This Repo is about Romanized Sinhala to Sinhala Transliteration using the Ngram and Rule Base Model.
Model Description
This dataset is capable of handling short-hand typing(Adhoc Transliteration).
eg
Input: khmda
Output : කොහොමද
If you are using this work:
Kindly cite :
T. G. D. K. Sumanathilaka, R. Weerasinghe and Y. H. P. P. Priyadarshana, "Swa-Bhasha: Romanized Sinhala to Sinhala Reverse Transliteration using a Hybrid Approach," 2023 3rd… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/Swabhasha_RomanizedSinhala_Dataset.Devnagari-Romanized-Pair
Dataset Overview
The dataset devanagari romanized pair contains, 959 rows, where each row has one English sentence and its corresponding Nepali translations both in the devanagari script and in romanized format, the size of the data set is less than 1,000 elements and it's designed for use in Translation, text generation and text to text generation tasks.
nepaliflow-romanized-nepali-to-devanagari-dataset
NepaliFlow Romanized Nepali to Devanagari Dataset
This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script.
Task
The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output.
Columns
prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari
completion: expected Nepali Devanagari output
Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.Legacy-Font-and-Romanized-Tamil-Corpusuonlp_culturaX_telugu_romanized_100kromanized_bangla
romanized_bangla — Dataset Card
Repository / id: sk-community/romanized_bangla
Derived from: wikimedia/wikipedia subset 20231101.bn (Bangla Wikipedia dump)
1. Short description
A romanized (Latin-script) version of Bangla Wikipedia text derived from the wikimedia/wikipedia dataset (subset 20231101.bn). The original dataset contained whole-article entries; you split article paragraphs into individual rows and transliterated Bangla text into a romanized representation… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_bangla.aya_romanized_sinhalatelugu_teknium_GPTeacher_general_instruct_filtered_romanizedtelugu_teknium_GPTeacher_general_instruct_filtered_romanized
