datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
advanced-soundscapes-stage-1
Advanced Soundscapes Stage 1 — Raw Components (5M)
This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan.
Contents
5,000 shards containing 5,000,000 soundscape recipes with raw audio components
Each soundscape row includes:
recipe.json — full recipe with timeline, events, loudness, speaker IDs, overlap/density settings
spkN.flac / spkN.json — raw speech components + full source metadata
musicN.flac /… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/advanced-soundscapes-stage-1.dacvae-tts-tr-stage2-b
DACVAE-TTS Turkish stage 2 (from run B 40k, clean data)
Generated audio of every evaluated checkpoint of the training run tr-stage2-b (Turkish zero-shot voice-cloning TTS,
dacvae-tts, frozen Meta DACVAE latents, 48 kHz). This repository holds
model outputs and metrics, not training data. Training data: Vyvo/tr-dataset-12 (Turkish podcast segments).
Stage 2: warm start (--init-from) from run B's 40k checkpoint (VoiceHub/dacvae-tts-tr-nano-b-ke4), trained 20k more updates on the… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/dacvae-tts-tr-stage2-b.stage1-audio-evidence-shards
BP-LM stage1 音声 shard
KazukiY/bplm-stage1-registry-v1
が参照する音声です。47 ソース / 2.33 TB / 1,216 ファイル。
台帳(registry)と合わせて この 2 つだけで完結します。
使い方
台帳側の取得スクリプトが両方を面倒みます。こちらを直接触る必要はありません。
pip install huggingface_hub
huggingface-cli download KazukiY/bplm-stage1-registry-v1 \
fetch_stage1_v1.py lane_defs.py --repo-type dataset --local-dir .
python fetch_stage1_v1.py --dest ./bplm # 全部
python fetch_stage1_v1.py --dest ./bplm --stage audio… See the full description on the dataset page: https://huggingface.co/datasets/KazukiY/stage1-audio-evidence-shards.voice_conversion_stage0_v2_4M-4.1Mdacvae-tts-tr-w512-stage2-hq
DACVAE-TTS Turkish run C stage 2 (width 512, high-quality subset)
Generated audio of every evaluated checkpoint of the training run tr-w512-stage2-hq (Turkish zero-shot voice-cloning TTS,
dacvae-tts, frozen Meta DACVAE latents, 48 kHz). This repository holds
model outputs and metrics, not training data. Training data: Vyvo/tr-dataset-12 (Turkish podcast segments).
Warm start from VoiceHub/dacvae-tts-tr-w512-clean (step 60k, 66.5M parameters), 10k more updates on the high-quality… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/dacvae-tts-tr-w512-stage2-hq.bplm-stage1-registry-v1
BP-LM stage1 registry v1
音声 9,301,859 clip に 8 種類の道具注釈を付けた台帳です。
never_run は 0。 全 clip・全レーンについて「値がある」か「無い根拠がある」かの
どちらかで、未処理はありません。
2 つのリポジトリで完結します
中身
大きさ
KazukiY/bplm-stage1-registry-v1
台帳(このリポジトリ)
14 GB
KazukiY/stage1-audio-evidence-shards
音声 47 ソース
2.33 TB
使い方
1. 取得する
pip install huggingface_hub
# 取得スクリプトを落とす
huggingface-cli download KazukiY/bplm-stage1-registry-v1 \
fetch_stage1_v1.py lane_defs.py… See the full description on the dataset page: https://huggingface.co/datasets/KazukiY/bplm-stage1-registry-v1.dacvae-tts-tr-stage3-hq
tr-stage3-hq
Generated audio of every evaluated checkpoint of the training run tr-stage3-hq (Turkish zero-shot voice-cloning TTS,
dacvae-tts, frozen Meta DACVAE latents, 48 kHz). This repository holds
model outputs and metrics, not training data. Training data: Vyvo/tr-dataset-12 (Turkish podcast segments).
How to listen
monitor/prompt-<uid>.wav: the reference voice given to the model (a real validation recording of an unseen speaker,
decoded through the DACVAE… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/dacvae-tts-tr-stage3-hq.advanced-soundscapes-stage-1
Advanced Soundscapes Stage 1 — Raw Components
This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan.
Contents
0 shard(s) containing 0 soundscape recipes with raw audio components
Each soundscape row includes:
recipe.json — full recipe with timeline, events, loudness, speaker IDs, overlap/density settings
spkN.flac / spkN.json — raw speech components + full source metadata
musicN.flac / musicN.json —… See the full description on the dataset page: https://huggingface.co/datasets/ChristophSchuhmann/advanced-soundscapes-stage-1.stage1a_smoke_data
stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh)
Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP →
frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend:
WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT
encoder — no raw-audio decoding at train time.
113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.gigaspeech-tiny-stage1Parler-TTS-Datadriven-100h-44.1kHz_stage1stage3_previewbengali-tts-folderized-parquet-stage1
Bengali TTS Folderized Parquet Stage 1
This is the intermediate organized parquet layer before the final fully row-wise Bengali TTS dataset.
Generated metadata refresh: 2026-06-18T21:44:49Z
Layout
<speaker>/part-00000.parquet
<speaker>/part-00001.parquet
<speaker>/metadata/stats.json
Columns
speaker
video_id
chunk_file
audio_file
duration
transcription
uuid
audio as Hugging Face Audio feature backed by parquet struct<bytes,path>… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-tts-folderized-parquet-stage1.bengali-tts-folderized-parquet-stage2
Bengali TTS Folderized Parquet Stage 2
Final filtered (keep=True) Bengali TTS dataset, with merged/combined chunks.
Layout
<speaker>.parquet (single shard, audio <= ~1GB)
<speaker>_001.parquet, _002.parquet, ... (multiple shards, split by audio byte size)
Audio sourcing convention
COMBINED == False -> sourced from extracted_audio/<speaker>/<video_id>/<chunk_file>
COMBINED == True -> sourced from… See the full description on the dataset page: https://huggingface.co/datasets/dipit099/bengali-tts-folderized-parquet-stage2.tts_stage2_extra
Bengali TTS Stage 2 Extra
Final filtered (keep=True) Bengali TTS dataset — extra speakers.
Layout
<speaker>.parquet (single shard, audio <= ~1GB)
<speaker>_001.parquet, _002.parquet, ... (multiple shards, split by audio byte size)
Audio sourcing convention
audio bytes sourced by uuid from bengali-tts-stage1-extra/<speaker>/part-*.parquet
See stat.json for per-speaker row/duration counts.
stage2_previewgigaspeech-tiny-stage4gigaspeech-tiny-stage2voice_conversion_stage0_v2_3.5M-3.7Mvnaska-two-stage-tagsgigaspeech-tiny-stage3voice_conversion_stage0_v2_3M-3.5Mvoice_conversion_stage0_v2_3.7M-3.9Mvoice_conversion_stage0_v2_4.2M-4.3Mvoice_conversion_stage0_v2_2M-2.5Malign_stage_2_data_ensnorTTS-stage2-audio-samplesstage0-codec
