datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PenaldoCR7pendakwah_teknologi_yt_stt_datasetDuDE-Stage-III
DuDE-Stage-III
Training data for synchronized duplex speech modeling, derived from Seamless Interaction by Meta.
full_conversations contains 7,387 complete conversations (488.973 conversation-hours; 977.946 participant-track hours), with 7,190 train and 197 dev examples. Both participant tracks retain their original shared clock. The longest example is 2,008 seconds. Internal ASR/codec chunks were reassembled; there is no example duration cap.
short_windows preserves the earlier… See the full description on the dataset page: https://huggingface.co/datasets/penguinfish1688/DuDE-Stage-III.pendakwah_teknologi_yt_stt_datasetnsc-imda-part6gigaspeech_subset_270hLibritts_p_dataset_20260129
Contribution
This dataset is a processed version of the original LibriTTS-P dataset, optimized for use on the Hugging Face platform. I've uploaded this version to make it more accessible to the community. All credit for the original data goes to the creators of LibriTTS-P.
Changes
Make a new column combined_prompt.
The combined_prompt is a concatenation of the style_prompt and speaker_prompt, using the connector: "The speaker's identity can be described as ".
In the… See the full description on the dataset page: https://huggingface.co/datasets/pengine/Libritts_p_dataset_20260129.accented_englishLIbritts_p_dataset_20260127
Contribution
This dataset is a processed version of the original LibriTTS-P dataset, optimized for use on the Hugging Face platform. I've uploaded this version to make it more accessible to the community. All credit for the original data goes to the creators of LibriTTS-P.
Changes
Make a new column combined_prompt.
The combined_prompt is a concatenation of the style_prompt and speaker_prompt, using the connector: "The speaker's identity can be described as ".
In the… See the full description on the dataset page: https://huggingface.co/datasets/pengine/LIbritts_p_dataset_20260127.wenetspeech-subset-SDS3500
English
中文
I. Basic Information of the Dataset
Dataset Name: Underwater Acoustic Target Radiated Noise Dataset (including the original ShipsEar dataset and the enhanced DS3500 dataset)
Dataset Version: V1.0
Release Date: July 2025 (based on the paper submission date)
Update Records: First release, no updates yet
Source and Contributors:
Original ShipsEar dataset: Collected along the Atlantic coast of Spain from 2012 to 2013
Enhanced DS3500 dataset: Generated by institutions such… See the full description on the dataset page: https://huggingface.co/datasets/peng7554/DS3500.penguinmod-vm-prain
PenguinMod/PenguinMod-Vm
Modified Scratch VM with a JIT compiler and more features.
This is a drop-in replacement for LLK/scratch-vm.
Setup
See https://github.com/TurboWarp/scratch-gui/wiki/Getting-Started to setup the complete TurboWarp environment.
If you just want to play with the VM then it's the same process as upstream scratch-vm.
Extension authors
If you only use the standard reporter, boolean, and command block types, everything should just work without… See the full description on the dataset page: https://huggingface.co/datasets/soiz1/penguinmod-vm-prain.NaiLong-Voice-Clone
奶龙语音克隆数据集
完整项目与 Demo 效果可参见 GitHub
如果这个数据集对你有帮助,欢迎在 GitHub 上点个 Star ⭐ 支持一下!
数据集介绍
数据集按处理阶段分为以下四部分:
1. raw_audio (原始采样)
处理方式:使用 Audacity 直接对视频素材进行录音,格式为 44.1kHz, 16-bit, Stereo。
说明:包含背景音、特效及多角色对话的非结构化原片素材,是整个流水线的起点。
2. vocal_only (人声分离)
处理方式:从 raw_audio 中使用 UVR5 的 MDX-Net 模型剥离背景音乐与噪音。
说明:利用 MDX-Net 模型提取出干净的人声轨道,为后续切片提供高信噪比素材。
3. sliced_vocal (自动化切片)
处理方式:基于停顿检测、音色突变及总时长控制,将 vocal_only 自动化切分为一系列短音频。… See the full description on the dataset page: https://huggingface.co/datasets/pengyichen/NaiLong-Voice-Clone.DuDE-Stage-II
DuDE-Stage-II
Committed Stage II teacher self-distillation records: 337,389 train utterances (703.460 hours of synthetic codec targets) and 128 validation utterances.
Each row is one independent utterance. The trainer samples independent utterances into two densely interleaved channels; this dataset does not assemble conversations.
Targets were generated by Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice (weights revision 0c0e3051f131929182e2c023b9537f8b1c68adfe, upstream code revision… See the full description on the dataset page: https://huggingface.co/datasets/penguinfish1688/DuDE-Stage-II.IALP-2026-data
IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets
Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC
Adaptation." This repository holds the fixed target-query, validation, and
evaluation sets used across all experiments. Each part is a self-contained
.tar.gz.
All audio is 16 kHz mono. Each split ships with:
audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16)
wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.ESD-SubsetsESD-Subsets_answerPennyNTU-refined-NSC-2021-evaluationDaniel_Penincustom-datasettara_penaltypenny456SWB1simran_penaltypeneraabritahub5_english_eval_2000_swb1EN-MALAY-CSEN-INDON-CSemotion-penyisihan_satria_data_2025-dataset
