datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
joyo-kanji-yomi-benchmark-parakeet
日本語 | English
常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet)
常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。
このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。
このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。
概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.joyo-kanji-yomi-benchmark-parakeet
日本語 | English
常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet)
常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。
このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。
このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。
概要… See the full description on the dataset page: https://huggingface.co/datasets/0x3/joyo-kanji-yomi-benchmark-parakeet.parakeet-hindi-asr
Parakeet Hindi-English Bilingual ASR
Fine-tuning NVIDIA Parakeet TDT 0.6B for bilingual Hindi-English automatic speech recognition.
Quick Start
# Download
pip install huggingface_hub
huggingface-cli download ketav/parakeet-hindi-asr --repo-type dataset --local-dir ./parakeet-hindi-asr
# Install dependencies
pip install nemo_toolkit[asr] bitsandbytes sentencepiece
# Train (after updating paths in config)
cd parakeet-hindi-asr/scripts
python ft_0.6B_hi_v3.py… See the full description on the dataset page: https://huggingface.co/datasets/ketav/parakeet-hindi-asr.J-HARD-TTS-Eval
J-HARD-TTS-Eval
[!NOTE]
For full documentation, detailed benchmark results, and methodology, please refer to the GitHub Repository.
Overview
J-HARD-TTS-Eval is a benchmark designed to evaluate the robustness of autoregressive Japanese Text-To-Speech (TTS) models.
It focuses on specific failure modes such as stability in short sequences, repetition handling, and context completion.
Usage
You can easily load the dataset using the Hugging Face datasets… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/J-HARD-TTS-Eval.neuro-parakeet-food
neuro-whisper-v1
Dataset Description
This is a synthetic dataset for German medical speech recognition, specifically designed for fine-tuning ASR models on neuro-oncology and neurology terminology. The dataset provides a comprehensive coverage of German medical terminology in the neurology and neuro-oncology domains.
Data Generation
Voice Data: Synthetically generated using Resemble AI Chatterbox TTS
Text Data: Medical text generated with Qwen/Qwen3-30B-A3B… See the full description on the dataset page: https://huggingface.co/datasets/NeurologyAI/neuro-parakeet-food.parakeet-tdt-blind-spots
Blind Spots of nvidia/parakeet-tdt-0.6b-v2
This dataset documents 14 systematically identified blind spots in NVIDIA's parakeet-tdt-0.6b-v2 automatic speech recognition model. The errors span 8 distinct categories and reveal a consistent pattern: the model struggles with inputs outside the distribution of its Western English-centric training data.
Model Under Test
Property
Value
Model
nvidia/parakeet-tdt-0.6b-v2
Parameters
600M
Architecture… See the full description on the dataset page: https://huggingface.co/datasets/TieIncred/parakeet-tdt-blind-spots.parakeet-indic-audioparakeet-whisper-divergence
Parakeet vs Whisper Divergence Samples
This dataset contains 40 audio samples where Parakeet TDT v3 and Whisper (Granary pseudolabels) show significant divergence.
Purpose
Investigate why Whisper outputs blank or very short transcriptions on French VoxPopuli audio.
Issues Found
whisper_short: Whisper outputs < 2 chars/sec while Parakeet outputs > 5 chars/sec
high_wer: WER > 80% between Whisper and Parakeet transcriptions
Common Whisper Hallucinations… See the full description on the dataset page: https://huggingface.co/datasets/toth235a/parakeet-whisper-divergence.parakeet-indic-dataeval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920
Evaluation Results: parakeet-tdt-0.6b-v3
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/parakeet-tdt-0.6b-v3
37.59%
20.64%
Source Data
Evaluation Dataset: Trelis/eka-hard
Model Evaluated: nvidia/parakeet-tdt-0.6b-v3
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920.eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926
Evaluation Results: parakeet-tdt-0.6b-v3
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/parakeet-tdt-0.6b-v3
11.34%
3.63%
Source Data
Evaluation Dataset: Trelis/medical-terms-2025
Model Evaluated: nvidia/parakeet-tdt-0.6b-v3
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926.parakeet-stt-redone
parakeet-stt-redone
What this is
108,276 raw→clean transcript pairs sourced from
aldigobbler/stt-correction,
re-labeled using GLM-5.1-FP8 as the teacher model with our production
cleanup prompt.
How it differs from the source dataset
aldigobbler/stt-correction
this dataset
Target
Verbatim transcript restoration (lowercase, no punctuation, fillers kept/restored)
Polished readable text — punctuated, paragraphed, fillers selectively removed… See the full description on the dataset page: https://huggingface.co/datasets/rdsm/parakeet-stt-redone.eval-parakeet-tdt-0.6b-v3-multimed-hard-20260408-1930
Evaluation Results: parakeet-tdt-0.6b-v3
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/parakeet-tdt-0.6b-v3
15.94%
10.13%
Source Data
Evaluation Dataset: Trelis/multimed-hard
Model Evaluated: nvidia/parakeet-tdt-0.6b-v3
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-multimed-hard-20260408-1930.mmm_project_parakeetmmm_project_parakeet_intermdatamoe-speech-plus-cache-parakeet-v1
moe-speech-plus cache (parakeet transcription・cache_version=1)
otoha m5_f5_trainer の Path B' precache 生成物。
内訳
cache_version: 1
transcription_source: parakeet
n_samples: 354,110
speakers: 424 (moe-speech-plus 全 473 UUID から hold-out 48 話者除外・fair 分布)
hold-out UUID: 48 (GAP-11 で確定・commit 3f3a703)
使い方
from huggingface_hub import snapshot_download
cache_dir = snapshot_download(
repo_id="otoha-project/moe-speech-plus-cache-parakeet-v1"… See the full description on the dataset page: https://huggingface.co/datasets/otoha-project/moe-speech-plus-cache-parakeet-v1.parakeet_inferencehinglish-parakeet-tarred
