CoolFace
Datasetpublic

Aratako/Magpie-Speech-Orpheus-125k

Magpie-Speech-Orpheus-125k A ~125k-sample synthetic speech dataset generated by applying the Magpie instruction-synthesis approach to the Orpheus-TTS LLM-based text-to-speech model, then decoding audio tokens with the SNAC 24 kHz codec. Blog (EN): https://huggingface.co/blog/Aratako/magpie-speech Blog (JA): https://zenn.dev/aratako_lm/articles/87d8988d44ba4d This dataset is entirely synthetic: text prompts and audio tokens were produced by Orpheus-TTS and decoded to waveforms… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Magpie-Speech-Orpheus-125k.

sourceHugging Facellama3.2updated 1y agoView on Hugging Face
11likes69downloads
5 commits on main
26d41821y ago

Update README.md

Aratako
5240fff1y ago

Update README.md

Aratako
6a5e3a21y ago

Upload dataset (part 00001-of-00002)

Aratako
d7087521y ago

Upload dataset (part 00000-of-00002)

Aratako
d1c92e01y ago

initial commit

Aratako