datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.Myanmar-Written-Spoken-Parallel-Corpus
Myanmar Written-Spoken Parallel Corpus (MWSPC)
Dataset Description
Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language.
Curated by: Khant Sint Heinn (Kalix Louis)
Organization: DatarrX | ဒေတာ-အက်စ်
Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences
We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.spoken_telugumyanmar-written-spoken-text-pairs
📝 Myanmar written spoken text pairs`
Dataset Overview
Myanmar Written-Spoken Text Pairs is a parallel text dataset designed to support Natural Language Processing (NLP) research and applications for the Myanmar language. The dataset contains paired Myanmar sentences in two different styles:
written_style — Formal written Myanmar (ရေးဟန်)
spoken_style — Natural spoken Myanmar (ပြောဟန်)
This dataset is useful for tasks involving style transfer, text normalization… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/myanmar-written-spoken-text-pairs.en-my-written-spoken-parallel
English-Burmese Written and Spoken Parallel Dataset
An advanced, human-curated English-to-Burmese parallel dataset specifically designed to explicitly split Burmese translations into two distinct linguistic registers: Written Style (ရေးဟန် / Literary) and Spoken Style (ပြောဟန် / Colloquial). The content is predominantly focused on the technology and global news domains.
Creator: Khant Sint Heinn
Project Status: ⚠️ Active & Ongoing (Work in Progress) — This dataset is actively… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/en-my-written-spoken-parallel.Spoken2TSL
Dataset Description
This dataset is a collection of Turkish to Turkish Sign Language (TSL) grammar version translations. The dataset is designed to facilitate research and development in the field of sign language translation and understanding. It contains pairs of sentences in Turkish and their corresponding TSL translations, which have been curated to follow the grammatical structure of TSL.
Data Collection
The data was collected primarily from the website… See the full description on the dataset page: https://huggingface.co/datasets/ismaildlml/Spoken2TSL.primary-language-spoken-by-the-medicaid-and-chip-p
Primary language spoken by the Medicaid and CHIP population
Description
This data set includes annual counts and percentages of Medicaid and Children’s Health Insurance Program (CHIP) enrollees by primary language spoken (English, Spanish, and all other languages). Results are shown overall; by state; and by five subpopulation topics: race and ethnicity, age group, scope of Medicaid and CHIP benefits, urban or rural residence, and eligibility category.
These results were… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/primary-language-spoken-by-the-medicaid-and-chip-p.SpokenVisIT
