qwen3_asr
qwen3-asr-lg-atypical-promptless-specaug-results
Qwen3-ASR Promptless SpecAugment Ablation — Luganda
Private reproducibility archive for the Luganda promptless architectural
ablation with train-only SpecAugment.
Experimental controls
Base model: KasuleTrevor/cdli-qwen3-asr-lg-typical-1p7b-base-finetune
Dataset: cdli/ugandan_luganda_nonstandard_speech_v1.0
LR: 0.0001
Scheduler: cosine
Seed: 42
SpecAugment: enabled
Prompt: none
Checkpoint selection: corpus
Primary reporting: mean normalized per-utterance WER/CER… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/qwen3-asr-lg-atypical-promptless-specaug-results.Qwen3-ASR
Qwen3-ASR
🤗 Hugging Face | 🤖 ModelScope | 📑 Blog | 📑 Paper
🖥️ Hugging Face Demo | 🖥️ ModelScope Demo | 💬 WeChat (微信) | 🫨 Discord | 📑 API
We release Qwen3-ASR, a family that includes two powerful all-in-one speech recognition models that support language identification and ASR for 52 languages and dialects, as well as a novel non-autoregressive speech forced-alignment model that can align text–speech pairs in 11 languages.… See the full description on the dataset page: https://huggingface.co/datasets/echodict/Qwen3-ASR.Qwen3-ASR-PostTrain-Complete-Medical-French-Fullqwen3asr-shoken-conv
qwen3asr-shoken-conv
証券・投資ドメインの日本語会話YouTube音声(対談/インタビュー中心)。Qwen3-ASR 追加学習用。
音声/ : opus 38本 / 計19.4h / 多様チャンネル(最大3本/ch)
文字起こし_scribe/ : ElevenLabs Scribe diarized JSON(words[]にspeaker_id+start/end)+txt
cut_plan_shoken.jsonl : 無音境界≤120s の627チャンク。text=プレーン / text_spk=話者タグ付き
curated_20h.tsv : 収集元リスト(vid/dur/channel/title)
出典は公開YouTube。話者分離ラベルはScribe自動(2話者は信頼性高、3+は要検証)。
qwen3-asr-en-atypical-specaug-prompt-ablation-resultsqwen3-asr-hebrew-100k
