timcuhk/NRC-HeritageLab-ASR-Demo
NRC-HeritageLab ASR Demo
This public Space provides password-gated access to a browser-based automatic speech recognition demo for Inuktitut HeritageLab audio. The default model is the best primary model from the HeritageLab benchmark: a Qwen3-ASR 0.6B checkpoint fine-tuned on the HeritageLab dataset and selected by development-set word error rate. The app also includes the off-the-shelf Qwen/Qwen3-ASR-0.6B model so users can compare recognition before and after HeritageLab fine-tuning.
The app supports two model choices:
- HeritageLab fine-tuned Qwen3-ASR 0.6B: recommended for this Inuktitut demo.
- Off-the-shelf Qwen3-ASR 0.6B: original public model, included as a pre-fine-tuning comparison baseline.
The app supports two inference modes:
- GPU: uses Hugging Face ZeroGPU when available and is recommended for normal use.
- CPU: available as a portability fallback for short clips; it is substantially slower.
The app runs models locally inside the Space and does not call an external ASR API.
Access is controlled by an in-app password gate using the DEMO_USERNAME and DEMO_PASSWORD Space secrets. The password is not stored in the repository.
See ACCESS.md for the access handoff procedure and the location of the restricted cluster-only password note.
For Inuktitut audio, the recommended language setting is Auto / no forced language.
The recognition context is handled automatically from the language setting and is not exposed as a user-editable field.
Maximum transcript length (tokens) caps transcript length. The default is intended for short demo clips; increase it only for longer audio.
Included demo audio
The included demo buttons run ASR directly on short public Inuktitut pronunciation files from Wikimedia Commons/Tusaalanga:
illu:File:Iu-illu.ogg, Open Government Licence – Canada 2.0, source credited to Tusaalanga.siqiniq:File:Siqiniq.ogg, Open Government Licence – Canada 2.0, source credited to Tusaalanga.
