URajinda/myanmar_spoken_corpus
Credits and Acknowledgments This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn). Original Dataset Source: freococo/myanmar_spoken_corpus Modifications: * Curated and filtered for specific training needs of the ShweYon model. Integrated with 3% English subset for bilingual proficiency maintenance. Re-formatted into 36 shards for optimized Continued Pre-training (CPT). We are deeply grateful to freococo for their contribution to… See the full description on the dataset page: https://huggingface.co/datasets/URajinda/myanmar_spoken_corpus.
Credits and Acknowledgments
This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn).
- Original Dataset Source: freococo/myanmar_spoken_corpus
- Modifications: * Curated and filtered for specific training needs of the ShweYon model.
- Integrated with 3% English subset for bilingual proficiency maintenance.
- Re-formatted into 36 shards for optimized Continued Pre-training (CPT).
We are deeply grateful to freococo for their contribution to the Myanmar AI community by providing this high-quality spoken corpus.
