CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mugezhang /pair_tamil_malayalam_ipa_transcription_romanizedtext10M<n<100M0 likes277 downloads9mo agoHugging Face02mugezhang /massive_ipa_romanizedtext100K<n<1M0 likes173 downloads8mo agoHugging Face03kshitizgajurel /Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari Dataset Card for Dataset Name यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ। Dataset Prepared by: Manoj Kumar Baniya Aakash Kumar Thakur Manish Kathet Kshitiz Gajurel Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.text-generation10K<n<100K0 likes107 downloads2y agoHugging Face04sk-community /romanized_hindi Romanized Hindi Dataset Dataset Description The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation. It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods. The dataset is designed for training and evaluating Hindi↔Roman transliteration models. Language(s): Hindi, Romanized Hindi Size: ~1.82M rows License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.text1M<n<10M0 likes106 downloads1y agoHugging Face05mugezhang /indicxnli_ipa_romanizedtext100K<n<1M0 likes89 downloads8mo agoHugging Face06mugezhang /pair_russian_polish_ipa_romanizedtext1M<n<10M0 likes70 downloads9mo agoHugging Face07mugezhang /eng_spa_owt_phonemized_full_romanizedtext10M<n<100M0 likes54 downloads9mo agoHugging Face08Telugu-LLM-Labs /telugu_alpaca_yahma_cleaned_filtered_romanizedtext10K<n<100K19 likes53 downloads3y agoHugging Face09vrclc /Dakshina-romanized-mltext100K<n<1M0 likes38 downloads2y agoHugging Face10mugezhang /xlsum_unseen_phonemized_romanizedtext100K<n<1M0 likes34 downloads4mo agoHugging Face11ravithejads /yahma_alpaca_cleaned_telugu_filtered_and_romanizedtext10K<n<100K1 likes31 downloads3y agoHugging Face12tanziro /bangla-romanized-dhaka-shortchat-small-sfttext10K<n<100K0 likes31 downloads4mo agoHugging Face13Telugu-LLM-Labs /telugu_teknium_GPTeacher_general_instruct_filtered_romanizedtext10K<n<100K15 likes29 downloads3y agoHugging Face14mugezhang /xnli-en-ipa_ipa_romanizedtext100K<n<1M0 likes29 downloads6mo agoHugging Face15indiehackers /winogrande_debiased-telugu-romanized-nodicttext10K<n<100K0 likes28 downloads2y agoHugging Face16tanziro /bangla-romanized-dhaka-shortchat-1b-sfttext10K<n<100K0 likes28 downloads4mo agoHugging Face17deshanksuman /Swabhasha_RomanizedSinhala_Dataset Model Card for Model ID This Repo is about Romanized Sinhala to Sinhala Transliteration using the Ngram and Rule Base Model. Model Description This dataset is capable of handling short-hand typing(Adhoc Transliteration). eg Input: khmda Output : කොහොමද If you are using this work: Kindly cite : T. G. D. K. Sumanathilaka, R. Weerasinghe and Y. H. P. P. Priyadarshana, "Swa-Bhasha: Romanized Sinhala to Sinhala Reverse Transliteration using a Hybrid Approach," 2023 3rd… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/Swabhasha_RomanizedSinhala_Dataset.text1M<n<10M2 likes26 downloads1y agoHugging Face18jayasuryajsk /google-fleurs-te-romanizedaudio1K<n<10K0 likes25 downloads2y agoHugging Face19nirajandhakal /Devnagari-Romanized-Pair Dataset Overview The dataset devanagari romanized pair contains, 959 rows, where each row has one English sentence and its corresponding Nepali translations both in the devanagari script and in romanized format, the size of the data set is less than 1,000 elements and it's designed for use in Translation, text generation and text to text generation tasks. texttranslationn<1K3 likes25 downloads2y agoHugging Face20krishnAbadikelA /massive-unseen-ipa-romanizedtext10K<n<100K0 likes25 downloads5mo agoHugging Face21ai4bharat /IndicQA-romanizedtext1K<n<10K1 likes22 downloads2y agoHugging Face22mugezhang /pair_hindi_urdu_ipa_transcription_romanizedtext10M<n<100M0 likes22 downloads6mo agoHugging Face23tanziro /bangla-romanized-dhaka-shortchat-small-clean-v2-sfttext1K<n<10K1 likes21 downloads4mo agoHugging Face24dipeshch71 /nepaliflow-romanized-nepali-to-devanagari-dataset NepaliFlow Romanized Nepali to Devanagari Dataset This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script. Task The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output. Columns prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari completion: expected Nepali Devanagari output Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.text10K<n<100K1 likes19 downloads3mo agoHugging Face25sanujen /Legacy-Font-and-Romanized-Tamil-Corpustexttext-classification10K<n<100K0 likes18 downloads1y agoHugging Face26sabin1234 /Grounded_HPV_Cervical_Cancer_Romanized Grounded HPV & Cervical Cancer — Nepali (Romanized, Fixed) ShareGPT Dataset File: hpv_fixed.jsonl Format: JSON Lines (.jsonl), one JSON object per line Conversation schema: ShareGPT ("from": "human" / "from": "gpt") Language: Nepali (ne / ISO 639-3 npi), written in romanized script (Latin letters), not Devanagari License: CC-BY-4.0 Total records: 24,597 File size: ~47 MB 1. What this dataset is This is a synthetic, fact-grounded, multiple-choice-question (MCQ)… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Grounded_HPV_Cervical_Cancer_Romanized.text10K<n<100K0 likes17 downloads1d agoHugging Face27indiehackers /telugu_romanized_2000_mistraltext100K<n<1M0 likes15 downloads2y agoHugging Face28Telugu-LLM-Labs /uonlp_culturaX_telugu_romanized_100ktext100K<n<1M2 likes14 downloads3y agoHugging Face29sk-community /romanized_bangla romanized_bangla — Dataset Card Repository / id: sk-community/romanized_bangla Derived from: wikimedia/wikipedia subset 20231101.bn (Bangla Wikipedia dump) 1. Short description A romanized (Latin-script) version of Bangla Wikipedia text derived from the wikimedia/wikipedia dataset (subset 20231101.bn). The original dataset contained whole-article entries; you split article paragraphs into individual rows and transliterated Bangla text into a romanized representation… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_bangla.text100K<n<1M0 likes14 downloads1y agoHugging Face30indiehackers /telugu_romanizedtext10K<n<100K0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.