datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-casual-conversational-speech-golden-dataset-preview
Japanese Casual Conversational Speech Golden Dataset (Preview)
💼 Commercial License & Full Access
This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning.
To purchase the full dataset, please contact us:
👉 Email: info@hth-inc.com
👉 Website: https://hth-inc.com/business
🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.german-golden-audio_speech-IPA
🌟 German Golden Speech & IPA Corpus (FLEURS + Multilingual TEDx)
An ultra-clean, high-standard curated German speech dataset combining Google FLEURS (de_de) and Multilingual TEDx German (mTEDx), fully embedded with 16kHz WAV audio bytes, normalized orthographic text, and pre-computed International Phonetic Alphabet (IPA) transcriptions.
📊 Dataset Summary
Total Samples: 1,354 high-quality audio recordings.
Total Size: ~419 MB (Compressed Parquet format).
Audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-golden-audio_speech-IPA.visualears-golden-6669
🗂️ visualears-golden-6669
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Held-out VisualEars6669 / Golden6669 evaluation dataset.
مجموعهٔ ارزیابی نگهداشتهشدهٔ Golden6669 با شرایط پاک، دورمیدان و مسدود برای سنجش واقعگرایانهٔ ASR فارسی.
🧩 Role
evaluation and benchmarking asset
مصنوع ارزیابی و بنچمارک
📦 Snapshot
13 files; approximately 916.66 MB
13 فایل؛ حدود… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-golden-6669.gold-treasures-artgol-dataset
GOL Dataset — audited metadata repair
This card is based on a read-only audit of revision
4b43743fbe5f84589130587c30d625a8b0d95694. The repository is gated and contains
visual-novel speech, but the source card does not identify the included titles, rights
holders, extraction procedure, or license. Access approval is not a license; maintainers
must document the lawful use and redistribution basis before downstream use.
Exact repository inventory
607 files totaling… See the full description on the dataset page: https://huggingface.co/datasets/midralab/gol-dataset.gopt-vh-gold
VuiHoc GOPT Gold — audio + consensus labels (Arrow)
6,361 cau IELTS read-aloud thuan sach (~35h) tu he thong VuiHoc, moi row gom
audio + nhan dong thuan 3 vendor (SpeechAce, SpeechSuper, iFlytek ISE),
thang nghiep vu [0.0, 100.0], chia 4 split zero-leakage.
Splits
Split
Mau
Speakers
Phone valid
Word valid
train
4,643
812
94.25%
98.25%
val
581
100
94.28%
98.32%
test_unseen_speakers
580
94
94.2%
98.13%
test_unseen_prompts
557
190
93.2%
97.77%… See the full description on the dataset page: https://huggingface.co/datasets/tiennguyenbnbk/gopt-vh-gold.neyshekar-koochik-gold-27kgolden-dataset-2.1visualears-benchmark-269-gold
🗂️ visualears-benchmark-269-gold
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
269-record gold/noisy benchmark dataset.
معیار طلایی ۲۶۹ نمونهای برای بررسی سریع خطاهای گفتار نویزی و مقایسهٔ نسخههای مدل.
🧩 Role
evaluation and benchmarking asset
مصنوع ارزیابی و بنچمارک
📦 Snapshot
276 files; approximately 44.35 MB
276 فایل؛ حدود 44.35 MB
🧱 Packaging
1 Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-benchmark-269-gold.zant-echo-golden
license: cc0-1.0
task_categories:
- audio-classification
language:
- en
tags:
- speaker-diarization
- test-dataset
size_categories:
- n<1K
ZantOS Golden Test Set
Human-recorded meeting audio with ground truth speaker annotations for acceptance testing.
Dataset Details
Version: 1.0.0
Clips: 3 meetings (2-4 minutes each)
Speakers: 2-4 per clip
Format: 16kHz mono WAV
Annotation: RTTM format (Rich Transcription Time Marked)… See the full description on the dataset page: https://huggingface.co/datasets/zant-os/zant-echo-golden.golden-eval-set
Golden Eval Set
Private Vietnamese audio evaluation set with audio and transcription columns.
gold-treasures-audioalconost-multilingual-speech-gold
Multilingual Speech & Translation Dataset — EN↔JA/AR-EG/PL/RU (10 phrases, dual-take)
Description
10 English source phrases with expert human translations into Japanese, Egyptian
Arabic (ar-EG), and Polish. Each target phrase is recorded by native speakers (two
takes each). Audio files are WAV 48 kHz mono, 16‑bit PCM format.
Translations are produced and QA'd by professional linguists; recordings follow
consistent orthography/style (AR-EG: Egyptian dialect; JA/PL: standard). All… See the full description on the dataset page: https://huggingface.co/datasets/alconost/alconost-multilingual-speech-gold.golha-asr-gold-69
🗂️ golha-asr-gold-69
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Golha gold-69 evaluation dataset.
مجموعهٔ طلایی ۶۹ نمونهای گلها برای ارزیابی دستی و تحلیل دقیق خطای ASR.
🧩 Role
evaluation and benchmarking asset
مصنوع ارزیابی و بنچمارک
📦 Snapshot
4 files; approximately 11.92 MB
4 فایل؛ حدود 11.92 MB
🧱 Packaging
1 Parquet files and 0 standalone audio files
1… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/golha-asr-gold-69.Thanosgol-dataset-2k-ljspeech
GOL 2K LJSpeech metadata — audited repair
This gated repository contains one 320.54 GB tar archive and a pipe-delimited metadata
file. The source repository did not document provenance, selection rules, audio format,
license, or the meaning of “2K”. This card records only properties verified at revision
23747a89469c5487262604efb21d72bd7beef41f; it does not fill those gaps by inference.
Verified contents
metadata.csv: 1,652,985 logical records with contiguous IDs… See the full description on the dataset page: https://huggingface.co/datasets/midralab/gol-dataset-2k-ljspeech.urdu-gold-audioGowgol-dala-cluster
GOL voice clusters — audited repair
This audit covers midralab/gol-dala-cluster revision
89e1c5982086207f0a8de22cdb880e5ea52789f6. The original repository has no dataset
card and stores its files under Windows-style backslash paths. Its 89-byte
cluster_statistics.json ends inside the cluster_statistics object and is invalid
JSON.
Verified source layout
voice_clusters.csv: 7,362,684 strictly parsed rows, 596 folders, 19,245 speakers,
150 cluster IDs (0–149), and… See the full description on the dataset page: https://huggingface.co/datasets/midralab/gol-dala-cluster.LouisBeaters
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Goldyhghoul/LouisBeaters.KnySpidermanGOLDENINFANTIL3spk-cetuc-golden-dataset3spk-cetuc-golden-with-evaluichanJinxSelimhanSenoglu
