CoolFace
Modelpublic

AtomGradient/Qwen3-ASR-0.6B-int8-ov

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
0likes25downloads
Model Card

Qwen3-ASR-0.6B for Intel Core Ultra

Maintained by AtomGradient / 质子梯度(北京)科技有限公司.

Model artifacts used by AtomGradient ASR Runtime for local speech recognition on Intel hardware. The evaluated configuration uses Qwen3-ASR-0.6B with INT8 weight storage and floating-point components. INT8 does not describe every computation in the system.

Model and device specifications

ItemSpecification
Base modelQwen3-ASR-0.6B
Model size class0.6B
Delivered runtimeAtomGradient ASR Runtime, supplied separately from this model repository
Evaluated input16 kHz, mono, 16-bit PCM
Recognition modeText returned after processing a complete audio segment
Validated deviceIntel Core Ultra 5 225H, 32 GiB DDR5
Evaluation focusChinese dialogue; public Chinese and English audio samples

AtomGradient ASR Runtime

AtomGradient ASR Runtime is the local speech-recognition runtime in the AtomGradient speech delivery package. It provides model execution, recognition-request handling, cancellation and dialogue integration, using these model artifacts and third-party inference libraries.

The runtime and the model artifacts are complementary delivery components. This repository distributes the model artifacts; AtomGradient supplies the runtime, installation package and customer deployment documentation separately. The base model and third-party inference libraries retain their respective attribution.

Observed performance

Public audio tests on the device above, 2026-09-11:

SampleAudio durationProcessing time
Chinese4.204 s0.391 s
Short English2.461 s0.296 s
Longer English15.051 s1.378 s

Each file was run twice. The table shows the second, warm observation, including audio preprocessing and excluding model loading. These are small-sample measurements, not latency percentiles or a general speed guarantee.

A separate AtomGradient ASR Runtime field test on 2026-09-12 completed 10 of 10 recognition requests, with median local request processing time 0.236 s and range 0.172–1.031 s. This included one wake-phrase verification and nine dialogue requests. Timing excludes the user's speaking time, audio upload and end-of-speech detection. Concurrent speech synthesis can increase recognition latency.

Scope and limitations

  • The recorded tests establish that the model runs on the specified device. They do not establish recognition accuracy across arbitrary speakers, languages or environments.
  • No held-out WER/CER benchmark or lossless-quantization claim is made. Transcriptions differed between precision variants on the longer English sample.
  • The public audio tests and the field-test requests are different inputs with different timing boundaries; they must not be combined into one latency distribution.
  • This repository contains model artifacts and associated configuration. AtomGradient ASR Runtime, the installation package and product integration are delivered separately. Model files alone are not the complete demonstrated assistant.

For deployment and integration, contact [AtomGradient](https://huggingface.co/AtomGradient).

License and attribution

The base model is [Qwen/Qwen3-ASR-0.6B](https://huggingface.co/Qwen/Qwen3-ASR-0.6B). Model artifacts are distributed under Apache-2.0, following the base model's license. AtomGradient's copyright covers its own contributions and does not extend to the upstream Qwen model or third-party components.