datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peaky-blinders-learning-purpose-only
TTS Dataset - Peaky Blinders
This dataset contains audio segments with transcriptions from Peaky Blinders for Text-to-Speech training.
Dataset Generation
This dataset was generated using the TTS-Dataset-Maker pipeline, which provides:
Silero VAD-based silence removal - Removes long silences while preserving natural speech gaps
DeepFilterNet denoising - CPU-optimized audio denoising with gentle attenuation (15dB)
AssemblyAI transcription - High-quality speech-to-text with… See the full description on the dataset page: https://huggingface.co/datasets/ahk-d/peaky-blinders-learning-purpose-only.ePark_xue_xi_ci_biao_learning_vocabulary
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary.
