CoolFace
Datasetpublic

argo11/natori-irodori-tts-dataset

natori-irodori-tts-dataset High-quality single-speaker Japanese speech dataset prepared for Irodori-TTS LoRA training from multiple long-form さなちゃんねる videos featuring 名取さな. Dataset Summary This repository is a merged export of four curated subsets derived from public YouTube playlists on さなちゃんねる. Current top-level merged export: 40,995 total utterances 33,398 training examples 7,597 validation examples 43.4277 hours of speech speaker id: natori_sana The merged… See the full description on the dataset page: https://huggingface.co/datasets/argo11/natori-irodori-tts-dataset.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes456downloads
Dataset Card

natori-irodori-tts-dataset

High-quality single-speaker Japanese speech dataset prepared for Irodori-TTS LoRA training from multiple long-form さなちゃんねる videos featuring 名取さな.

Dataset Summary

This repository is a merged export of four curated subsets derived from public YouTube playlists on さなちゃんねる.

Current top-level merged export:

  • —40,995 total utterances
  • —33,398 training examples
  • —7,597 validation examples
  • —43.4277 hours of speech
  • —speaker id: natori_sana

The merged dataset currently contains these subsets:

  • —paranormasight_2026: 5.6632h
  • —chats_2024: 16.4336h
  • —chats_2025: 15.5682h
  • —chats_2026: 5.7628h

Source Data

All source media comes from the YouTube channel さなちゃんねる.

Playlists used in the current release:

  • —Paranormasight gameplay/commentary:
  • —https://youtube.com/playlist?list=PLPDt0GwV6dYdk79D0obyy0IOtFogXB-6R&si=0n2JxJatr8dtdaiU
  • —2026 chat streams:
  • —https://youtube.com/playlist?list=PLPDt0GwV6dYdx80xch7a9IKpM6leRn8XJ&si=1fOo79y2_aImnq8R
  • —2025 chat streams:
  • —https://youtube.com/playlist?list=PLPDt0GwV6dYdvClZPJOS9aupubucXm6nh&si=wJNgnCbt4Qp79v48
  • —2024 chat streams:
  • —https://youtube.com/playlist?list=PLPDt0GwV6dYfUJrt3deX9yoisWDlJ2pUk&si=nQf7U_aRhMrMgvJ8

Source format:

  • —long-form gameplay commentary
  • —long-form free-talk / chat streams

Derived assets in this repository:

  • —segmented WAV clips
  • —merged and subset JSONL manifests
  • —dataset statistics
  • —audit samples for manual review

Dataset Structure

Top-level files:

  • —train.jsonl: merged training split
  • —valid.jsonl: merged validation split
  • —dataset_stats.json: merged counts, durations, and subset summary
  • —dataset_audit_sample.csv: merged manual-review sample
  • —README.md: dataset card

Subset directories:

  • —subsets/paranormasight_2026/
  • —subsets/chats_2024/
  • —subsets/chats_2025/
  • —subsets/chats_2026/

Each subset contains:

  • —train.jsonl
  • —valid.jsonl
  • —dataset_stats.json
  • —dataset_audit_sample.csv
  • —playlist_metadata.json
  • —videos.csv
  • —audio/

Note on audio layout:

  • —paranormasight_2026 and chats_2026 store audio directly under audio/
  • —chats_2024 and chats_2025 are sharded under audio/s000/, audio/s001/, ... to stay within Hugging Face directory file-count limits

Each JSONL row contains:

  • —audio: repository-relative path to the WAV file
  • —text: machine-generated transcript
  • —speaker: speaker id
  • —video_id: source YouTube video id
  • —segment_id: segment identifier
  • —quality_score: combined filtering score
  • —duration_sec: utterance duration in seconds
  • —subset: subset name

Processing Pipeline

The current merged release was prepared with the natori-irodori-tts-pipeline workflow.

High-level steps:

  1. 1.collect playlist metadata
  2. 2.download and normalize source audio
  3. 3.detect speech / music / singing events with FireRedVAD
  4. 4.apply speech enhancement to speech candidates
  5. 5.refine speech boundaries with Silero VAD
  6. 6.filter segments by speaker similarity
  7. 7.generate transcripts with ASR
  8. 8.score and keep high-quality segments
  9. 9.export subset-level train / valid manifests
  10. 10.merge subsets into the top-level dataset export

Front-end stack used for the chat-stream subsets:

  • —FireRedVAD for event-aware front-end filtering
  • —SpeechBrain speech enhancement
  • —Silero VAD for refined utterance boundaries
  • —speaker embedding filtering
  • —Faster-Whisper large-v3 for Japanese ASR

Representative filtering rules used in the merged release:

  • —speaker_score >= 0.70
  • —asr_score >= 0.42
  • —2s <= duration <= 10s
  • —6 <= text_len <= 120
  • —japanese ratio >= 0.75
  • —repeated-character noise removed

Intended Uses

  • —single-speaker Japanese TTS adaptation
  • —LoRA fine-tuning for Irodori-TTS-compatible pipelines
  • —voice cloning experiments from commentary / chat style speech
  • —controlled experiments on long-form streaming audio turned into TTS supervision

Limitations

  • —Transcripts are machine-generated and may contain ASR errors.
  • —The dataset is single-speaker and domain-specific; it is not a general-purpose speech corpus.
  • —Source material includes commentary, chat-style speech, proper nouns, stream-specific context, and occasional platform / game references.
  • —Some subsets originate from gameplay streams and may still reflect the source domain even after filtering.
  • —Users should independently confirm that their use complies with the rights and terms governing the original videos and any third-party content appearing in them.

License and Rights

This dataset card is marked as license: other because no standard open license is asserted for the underlying source media.

This repository does not grant rights to:

  • —the original videos
  • —character likeness
  • —voice performance
  • —any third-party copyrighted content present in the source material

Those rights remain with their respective owners.


追記: argo11 管理用カード

管理メモ

この dataset は、名取さな / さなちゃんねる由来の長尺動画から作成された、日本語単一話者 TTS / Irodori-TTS LoRA 学習向け音声 dataset です。既存カードに記載の通り、merged export は 40,995 utterances、train 33,398、validation 7,597、約 43.4277 時間、speaker id は natori_sana です。

確認した内容

  • —audio は audio/*.wav として多数保存されています。
  • —text / metadata は JSON 系ファイルで管理されています。
  • —tag は text-to-speech, language:ja, audio, tts, japanese, single-speaker です。

推定した内容

  • —Irodori-TTS / voice clone 系の学習・検証に使うため、長尺動画から VAD / ASR / text cleanup / split を経て作成された dataset と推定しています。

注意

  • —元動画・話者・キャラクター権利に関わる利用制限があります。公開利用、再配布、商用利用、声質模倣用途では元コンテンツと権利者の条件を必ず確認してください。
  • —学習済み model は argo11/natori-irodori-tts-model と対応します。