datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librispeech-clean-lance
LibriSpeech clean (Lance Format)
A Lance-formatted version of the LibriSpeech ASR clean configuration, sourced from openslr/librispeech_asr. Each row is one utterance with inline FLAC audio bytes, the reference transcript, a sentence-transformers embedding of that transcript, and speaker/chapter metadata — all available directly from the Hub at hf://datasets/lance-format/librispeech-clean-lance/data.
Key features
Inline FLAC bytes in the audio column at 16 kHz mono… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/librispeech-clean-lance.emo_webds_2_formatted_batch_140nyan-jenny-format
Nyan Jenny Format
A Japanese text-to-speech (TTS) dataset in Jenny TTS format, derived from the
multilingual_toxicity_dataset
(Japanese split). Toxic words identified by DeepSeek V3 Pro are replaced with
「にゃん」(nyan) before audio synthesis, producing naturally-sounding speech
that masks harmful content.
Dataset Summary
This dataset was created through the following pipeline:
Source data — The Japanese split of
textdetox/multilingual_toxicity_dataset
(5 k samples:… See the full description on the dataset page: https://huggingface.co/datasets/RikkaBotan/nyan-jenny-format.emo_webds_2_formatted_batch_61vggsound_formatted_batch_5emo_webds_2_formatted_batch_109Afrivoice_Swahili-Voice_Instruct_Formatemo_webds_2_formatted_batch_10glue_fairseq_formatjenny_vibevoice_formattedThe Jenny (Dioco) dataset formatted for the VibeVoice format.
License:
Attribution is required in software/websites/projects/interfaces (including voice interfaces) that generate audio in response to user action using this dataset. Atribution means: the voice must be referred to as "Jenny", and where at all practical, "Jenny (Dioco)". Attribution is not required when distributing the generated clips (although welcome). Commercial use is permitted. Don't do unfair things like claim the dataset… See the full description on the dataset page: https://huggingface.co/datasets/vibevoice/jenny_vibevoice_formatted.emo_webds_2_formatted_batch_6emotive_formatted_audioemo_webds_2_formatted_batch_62emo_webds_2_formatted_batch_45emo_webds_2_formatted_batch_94emo_webds_2_formatted_batch_103vggsound_formatted_batch_4emo_webds_2_formatted_batch_46emo_webds_2_formatted_batch_77emo_webds_2_formatted_batch_108emo_webds_2_formatted_batch_143emo_webds_2_formatted_batch_14emo_webds_2_formatted_batch_71emo_webds_2_formatted_batch_99vggsound_formatted_batch_16emo_webds_2_formatted_batch_54emo_webds_2_formatted_batch_146emo_webds_2_formatted_batch_79emo_webds_2_formatted_batch_97emo_webds_2_formatted_batch_83
