datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yoda_sentences
Yoda Speak
This small dataset was built using two resources:
Harvard Sentences, a list of 720 short sentences grouped into 72 sets of 10 sentences each
English to Yoda Translator, an online translator that converts normal English into Yoda's way of speaking.
Fun with this dataset I hope you have! Yes, hrrrm.
US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences
We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.10452_kurmanji-corrected-sentences
Cleaned Kurmanji Kurdish Sentences Dataset (Hawar Standard)
Dataset Description
This dataset contains over 10,000 highly curated and grammatically corrected Kurmanji Kurdish sentences. While the original raw sentences were sourced from the open-source Tatoeba project, they have undergone extensive and meticulous editorial correction to meet the strict standards of the Hawar orthography and authentic Kurdish grammar (Celadet Alî Bedirxan rules).
The… See the full description on the dataset page: https://huggingface.co/datasets/amedcj/10452_kurmanji-corrected-sentences.1_pattern_10Kplus_myanmar_sentences
🧠 1_pattern_10Kplus_myanmar_sentences
A structured dataset of 11,452 Myanmar sentences generated from a single, powerful grammar pattern:
📌 Pattern:
Verb လည်း Verb တယ်။
A natural way to express repetition, emphasis, or causal connection in Myanmar.
💡 About the Dataset
This dataset demonstrates how applying just one syntactic pattern to a curated verb list — combined with syllable-aware rules — can produce a high-quality corpus of over 10,000 valid… See the full description on the dataset page: https://huggingface.co/datasets/freococo/1_pattern_10Kplus_myanmar_sentences.portuguese-sentences-synthetic-g2p
Dataset Card for 'TigreGotico/portuguese_g2p'
Dataset Description
Dataset Summary
TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants.
It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.
