CoolFace
Modelpublic

kevinqz/MOSS-Transcribe-Diarize-Decoder-CoreAI

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes
Model Card
Canonical: `kevinqz/MOSS-Transcribe-Diarize-Decoder-CoreAI` — source of truth.

MOSS-Transcribe-Diarize Qwen3 Decoder (fabric)

Apple Core AI chat model — runs fully on-device on Apple Silicon (iPhone / iPad / Mac, macOS/iOS 27+).

A quantized stateful KV-cache chat .aimodel — an Apple Core AI conversion of OpenMOSS-Team/MOSS-Transcribe-Diarize, with an embedded tokenizer + chat template. Produced by coreai-fabric and indexed by coreai-catalog.

Model facts

FieldValue
Parameters0.6B
Architecturetransformer
Capabilitiesspeech-to-text, text-generation
Quantization / precisionnone / float32
Context length
On-disk size2.2 GB
Asset kindstateful KV-cache chat bundle; embedded tokenizer + chat template
assetVersion2.0

Use it

Install via the catalog, then run it with Apple's Foundation Models runtime:

bash
pip install coreai-catalog && coreai-catalog install moss-transcribe-diarize-decoder
swift
import CoreAILanguageModels
import FoundationModels

// modelURL = the installed macos/ bundle directory for this model
let model = try await CoreAILanguageModel(resourcesAt: modelURL)
let session = LanguageModelSession(model: model)
let reply = try await session.respond(to: "Explain on-device AI in one sentence.")
print(reply)

A complete, buildable example lives at coreai-catalog/examples/llm-chat.

Requirements

  • Deployment: macOS 27.0+ / iOS 27.0+, Xcode 27+. The asset serializes with minimum_os v27, so the on-device Swift runtime requires macOS/iOS 27+.
  • A Mac on macOS 26 can convert and inspect the asset but cannot run it on-device (the Swift runtime needs the 27 SDK).
  • Apple Silicon.

Intended use & limitations

  • Intended use: general on-device chat / text generation. Inherits the base model's capabilities, languages, and biases.
  • Limitations: uncompressed (fp16) — full precision. See the Evaluation section for the measured greedy fidelity vs the fp16 reference.

Evaluation (parity)

  • Gate A (structure): passed — the bundle's layout + metadata were validated on real hardware (Apple Silicon); the asset loads and generates.
  • Gate B (numeric accuracy): passed. Task-accuracy evaluation (e.g. tinyMMLU) is pending upstream: Apple's coreai.llm.eval is a stub in coreai-models 0.1.0 that cannot score a stateful KV-cache asset. Greedy fidelity vs fp32 can be measured on-device via the parity runner. fabric never fakes a parity number.
  • Runtime throughput (tok/s): to be published once measured on the on-device (macOS/iOS 27) Swift runtime. Not estimated — real numbers or none.

Provenance

FieldValue
Base modelOpenMOSS-Team/MOSS-Transcribe-Diarize @ d7231bbae2587a4af278735eb765b318c4f64edd
Converted bymodels/moss_transcribe/export_decoder.py (version not reported)
Recipemoss-transcribe-diarize-decoder (recipe_source: fabric)
Precision / quantizationfloat32 / none
Conversion date2026-07-10

Machine-readable, in this repo: `parity-report.json` (gate results) · `reproduce-manifest.json` (exact tool + stack + pinned revision to reproduce this conversion) · `LICENSE` (upstream terms).

License and attribution

Weights licensed apache-2.0 — see the bundled LICENSE. This artifact is a converted + quantized derivative of the base model (the Apache-2.0 §4(b) change notice): weights were converted to Apple Core AI format and quantized to uncompressed (fp16). The conversion itself is community work.

Links

The on-device Core AI ecosystem

This conversion is part of a broader open ecosystem for running models on Apple's on-device stack — useful references if you're building here:

Not affiliated with Apple

Community conversion. Not produced, hosted, or endorsed by Apple. Apple and Core AI are trademarks of Apple Inc., used here only to describe the target runtime/format. This is an independent community conversion.