datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-casual-conversational-speech-golden-dataset-preview
Japanese Casual Conversational Speech Golden Dataset (Preview)
💼 Commercial License & Full Access
This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning.
To purchase the full dataset, please contact us:
👉 Email: info@hth-inc.com
👉 Website: https://hth-inc.com/business
🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.visualears-golden-6669
🗂️ visualears-golden-6669
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Held-out VisualEars6669 / Golden6669 evaluation dataset.
مجموعهٔ ارزیابی نگهداشتهشدهٔ Golden6669 با شرایط پاک، دورمیدان و مسدود برای سنجش واقعگرایانهٔ ASR فارسی.
🧩 Role
evaluation and benchmarking asset
مصنوع ارزیابی و بنچمارک
📦 Snapshot
13 files; approximately 916.66 MB
13 فایل؛ حدود… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-golden-6669.german-golden-audio_speech-IPA
🌟 German Golden Speech & IPA Corpus (FLEURS + Multilingual TEDx)
An ultra-clean, high-standard curated German speech dataset combining Google FLEURS (de_de) and Multilingual TEDx German (mTEDx), fully embedded with 16kHz WAV audio bytes, normalized orthographic text, and pre-computed International Phonetic Alphabet (IPA) transcriptions.
📊 Dataset Summary
Total Samples: 1,354 high-quality audio recordings.
Total Size: ~419 MB (Compressed Parquet format).
Audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-golden-audio_speech-IPA.gol-dataset
GOL Dataset — audited metadata repair
This card is based on a read-only audit of revision
4b43743fbe5f84589130587c30d625a8b0d95694. The repository is gated and contains
visual-novel speech, but the source card does not identify the included titles, rights
holders, extraction procedure, or license. Access approval is not a license; maintainers
must document the lawful use and redistribution basis before downstream use.
Exact repository inventory
607 files totaling… See the full description on the dataset page: https://huggingface.co/datasets/midralab/gol-dataset.visualears-benchmark-269-gold
🗂️ visualears-benchmark-269-gold
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
269-record gold/noisy benchmark dataset.
معیار طلایی ۲۶۹ نمونهای برای بررسی سریع خطاهای گفتار نویزی و مقایسهٔ نسخههای مدل.
🧩 Role
evaluation and benchmarking asset
مصنوع ارزیابی و بنچمارک
📦 Snapshot
276 files; approximately 44.35 MB
276 فایل؛ حدود 44.35 MB
🧱 Packaging
1 Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-benchmark-269-gold.golden-eval-set
Golden Eval Set
Private Vietnamese audio evaluation set with audio and transcription columns.
alconost-multilingual-speech-gold
Multilingual Speech & Translation Dataset — EN↔JA/AR-EG/PL/RU (10 phrases, dual-take)
Description
10 English source phrases with expert human translations into Japanese, Egyptian
Arabic (ar-EG), and Polish. Each target phrase is recorded by native speakers (two
takes each). Audio files are WAV 48 kHz mono, 16‑bit PCM format.
Translations are produced and QA'd by professional linguists; recordings follow
consistent orthography/style (AR-EG: Egyptian dialect; JA/PL: standard). All… See the full description on the dataset page: https://huggingface.co/datasets/alconost/alconost-multilingual-speech-gold.gol-dataset-2k-ljspeech
GOL 2K LJSpeech metadata — audited repair
This gated repository contains one 320.54 GB tar archive and a pipe-delimited metadata
file. The source repository did not document provenance, selection rules, audio format,
license, or the meaning of “2K”. This card records only properties verified at revision
23747a89469c5487262604efb21d72bd7beef41f; it does not fill those gaps by inference.
Verified contents
metadata.csv: 1,652,985 logical records with contiguous IDs… See the full description on the dataset page: https://huggingface.co/datasets/midralab/gol-dataset-2k-ljspeech.
