CoolFace
Datasetpublic

HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview

Japanese Casual Conversational Speech Golden Dataset (Preview) 💼 Commercial License & Full Access This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning. To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business 🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.

sourceHugging Faceotherupdated 27d agoView on Hugging Face
2likes319downloads
Dataset Card

Japanese Casual Conversational Speech Golden Dataset (Preview)

💼 Commercial License & Full Access

This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning.

To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business


🌟 4 Reasons to Choose This Dataset

  1. 1.Authentic & "Rough" Natural Japanese Real, unscripted casual conversations. It captures the true essence of everyday Japanese, including frequent backchannels (Aizuchi), fillers, and complex conversational overlaps.
  1. 1.Advanced Speaker Separation (Stereo) For dialogues, we provide channel-separated stereo audio (Speaker A on Left / Speaker B on Right). This allows perfect isolation of voices even during overlapping speech, making it ideal for training Speaker Diarization and high-precision ASR models.
  1. 1.Ready-to-Train: Pre-segmented & Human-Verified The audio is already perfectly sliced into short utterance-level segments, precisely mapped to transcripts in JSONL format. 100% manually checked for extreme accuracy. You can feed this directly into Whisper or other ASR models without any preprocessing pain.
  1. 1.Highly Cost-Effective We offer premium, human-annotated quality at a fraction of the cost of large enterprise data agencies.

🔊 Technical Specifications

  • —Format & Channels:
  • —Monologues: WAV (44.1kHz, 16-bit), Mono
  • —Dialogues (Call-based): WAV (48kHz, 16-bit), Channel-separated Stereo (Independent L/R channels for each speaker)
  • —Total Duration: 120+ Hours (Full Dataset) / 5 Minutes (Preview)

📂 Data Structure & Metadata (metadata.jsonl)

The dataset provides utterance-level segmented audio files. Each line in the JSONL corresponds to one sliced audio file.

FieldTypeDescription
file_namestringRelative path to the sliced audio file
transcriptionstringLabeled and transcribed text
durationfloatAudio length in seconds
speaker_idstringUnique Speaker ID
languagestringLanguage code (ja)
genderstringmale / female / other
age_groupstringe.g., 20s, 50s

🏷️ Annotation & Labeling Rules

To preserve natural conversational phenomena, we embed specific tags in the transcription. This makes the dataset perfectly suited for training natural Spoken Dialogue Models.

TagMeaningExample (Japanese)
(F ...)Filler / Hesitation(F えー) そうですね。
(D ...)Disfluency / Correction(D 東京に) 大阪へ行きました。
(?)Inaudible明日は (?) に行きます。
(? ...)Unsure / Guessed(? 舞い) 上がっちゃいまして
[PERSON_NN]Anonymization / Masking[PERSON_01] に連絡します。

Filler (F) Tag — Detection Scope

Only えー-type and あのー-type fillers and their variants are tagged:

Variant GroupTagged Forms
えー familyえー, えーと, えーっと, えっと, えっとー
あの familyあのー, あのう, あの

Other filler-like words (e.g., まあ, なんか, うーん, はい) are transcribed as-is without (F) tags.


👥 Speaker Attributes (Sample)

Audio IDTypeSpeakersGenderAge GroupSegments
6ac6e83fMonologue1male30s4
70d0fb05Monologue1female30s3
a85693deMonologue1female30s22
jhConversation2male + female30s40

日本語自然会話音声データセット (プレビュー版)

💼 商用ライセンス・フルデータへのアクセス

本リポジトリはプレビュー版(約5分)です。アプリ「Kataro」を通じて収集された全60時間以上のフルデータセットは、商用利用、音声認識(ASR)のベンチマーク、音声対話モデルの学習用として販売しております。

フルセットの購入をご希望の方は、以下までお問い合わせください: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business


🌟 本データセットが選ばれる4つの理由

  1. 1.圧倒的にリアルな「生」の日本語 台本のない、極めて自然な日常会話を収録。相槌、フィラー(えー、あの)、発話の被りなど、従来のデータセットでは抜け落ちがちな「生の対話現象」を網羅しています。
  1. 1.高度な話者分離(ステレオ収録) 対話データはLchとRchで話者ごとに分離して収録されています。発話が重なっている箇所でも各話者の音声を完璧に抽出できるため、話者分離(Diarization)モデルの学習に最適です。
  1. 1.前処理不要:セグメント分割&人間による書き起こし済み 音声は発話単位で短くカットされ、JSONL形式のテキストと1対1で紐付けられています。専門スタッフが100%手作業で校正しているため、Whisper等のモデルにそのまま学習・評価用として投入可能です。
  1. 1.高いコストパフォーマンス 大手データベンダーと比較し、高品質なアノテーション済みデータを圧倒的な低コストでご提供します。

🔊 技術仕様

  • —ファイル形式とチャンネル構成:
  • —独話: WAV (44.1kHz, 16-bit) / モノラル
  • —対話(通話型): WAV (48kHz, 16-bit) / 話者別ステレオ(L/Rチャンネル分離)
  • —収録時間: 60時間以上(フルセット) / 5分(プレビュー版)