datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audioset-dasheng-0.6b-emb
AudioSet DaSheng-0.6B embeddings
Mean-pooled, float16 embeddings of
danjacobellis/audioset_opus_24kbps
from mispeech/dasheng-0.6B.
Columns
path: source clip path (string)
label: source AudioSet label indices (list of int64)
emb: 1,280-dimensional fixed-size list of float16
Audio is decoded from the source Opus bytes, mixed to mono, and resampled to
16 kHz. The embedding is the model's documented outputdim=None output:
sigmoid applied to the mean of the final… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/audioset-dasheng-0.6b-emb.qwen3-tts-0.6b-voice-zenless-100-2026-07-22
《绝区零》中文音色数据集(cat3)
先说明:这是中国游戏
《绝区零》本身就是由中国上海的米哈游开发的国产游戏。可参阅米哈游官网的公司介绍和《绝区零》中国大陆官方网站。
本目录里出现的大量英文,并不是在把游戏当成外国游戏,而是因为上游 Hugging Face 数据集 simon3000/zenless-voice 使用了英文或拼音形式的 speaker 标识。为了兼容已有脚本和模型,metadata.jsonl 中的 name 字段仍保留上游标识;下面专门给出完整中文对照,方便人读。
数据集概况
说话人标识:100 个
音频文件:100 个,每个标识对应一条拼接后的 WAV
总时长:约 1624.9 秒(约 27.1 分钟)
单条时长:约 10.0~33.2 秒
音频格式:24 kHz、单声道、16 位 PCM WAV
标注文件:metadata.jsonl
原始样本目录:../raw3
上游数据集:simon3000/zenless-voice
注意:早先抓取 raw3 时生成的是… See the full description on the dataset page: https://huggingface.co/datasets/MigoXV/qwen3-tts-0.6b-voice-zenless-100-2026-07-22.qwen3-tts-0.6b-voice-zenless-100-2026-07-23
《绝区零》中文音色数据集(cat3)
《绝区零》是由中国上海的米哈游开发的国产游戏。可参阅米哈游官网的公司介绍和《绝区零》中国大陆官方网站。
本目录里出现的大量英文,并不是在把游戏当成外国游戏,而是因为上游 Hugging Face 数据集 simon3000/zenless-voice 使用了英文或拼音形式的 speaker 标识。为了兼容已有脚本和模型,metadata.jsonl 中的 name 字段仍保留上游标识;下面专门给出完整中文对照,方便人读。
数据集概况
说话人标识:100 个
音频文件:100 个,每个标识对应一条拼接后的 WAV
总时长:约 1624.9 秒(约 27.1 分钟)
单条时长:约 10.0~33.2 秒
音频格式:24 kHz、单声道、16 位 PCM WAV
标注文件:metadata.jsonl
原始样本目录:../raw3
上游数据集:simon3000/zenless-voice
注意:早先抓取 raw3 时生成的是 16 kHz 单声道 WAV;为了与 cat2… See the full description on the dataset page: https://huggingface.co/datasets/MigoXV/qwen3-tts-0.6b-voice-zenless-100-2026-07-23.eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926
Evaluation Results: parakeet-tdt-0.6b-v3
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/parakeet-tdt-0.6b-v3
11.34%
3.63%
Source Data
Evaluation Dataset: Trelis/medical-terms-2025
Model Evaluated: nvidia/parakeet-tdt-0.6b-v3
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926.eval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920
Evaluation Results: parakeet-tdt-0.6b-v3
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/parakeet-tdt-0.6b-v3
37.59%
20.64%
Source Data
Evaluation Dataset: Trelis/eka-hard
Model Evaluated: nvidia/parakeet-tdt-0.6b-v3
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920.eval-parakeet-tdt-0.6b-v3-multimed-hard-20260408-1930
Evaluation Results: parakeet-tdt-0.6b-v3
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/parakeet-tdt-0.6b-v3
15.94%
10.13%
Source Data
Evaluation Dataset: Trelis/multimed-hard
Model Evaluated: nvidia/parakeet-tdt-0.6b-v3
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-multimed-hard-20260408-1930.
