CoolFace
Datasetpublic

URajinda/myanmar_spoken_corpus

Credits and Acknowledgments This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn). Original Dataset Source: freococo/myanmar_spoken_corpus Modifications: * Curated and filtered for specific training needs of the ShweYon model. Integrated with 3% English subset for bilingual proficiency maintenance. Re-formatted into 36 shards for optimized Continued Pre-training (CPT). We are deeply grateful to freococo for their contribution to… See the full description on the dataset page: https://huggingface.co/datasets/URajinda/myanmar_spoken_corpus.

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes119downloads
Dataset Card

Credits and Acknowledgments

This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn).

  • —Original Dataset Source: freococo/myanmar_spoken_corpus
  • —Modifications: * Curated and filtered for specific training needs of the ShweYon model.
  • —Integrated with 3% English subset for bilingual proficiency maintenance.
  • —Re-formatted into 36 shards for optimized Continued Pre-training (CPT).

We are deeply grateful to freococo for their contribution to the Myanmar AI community by providing this high-quality spoken corpus.