XReyRobert/granite-embedding-97m-multilingual-r2-coreml-ane-experimental
Experimental Granite 97M Multilingual R2 Core ML ANE Packages
Experimental Status
These artifacts are personal experimental Core ML conversions of `ibm-granite/granite-embedding-97m-multilingual-r2`. They are intended for macOS Core ML / Apple Neural Engine experimentation and should be treated as research/engineering artifacts rather than an official model release.
For the official model description, supported languages, intended use, training details, license, and limitations, refer to the original IBM Granite model card: <https://huggingface.co/ibm-granite/granite-embedding-97m-multilingual-r2>.
What This Repository Contains
This repository contains fixed-shape Apple Core ML .mlpackage variants derived from the original Granite Embedding 97M Multilingual R2 model:
The packages are conversion artifacts only. Consumers still need the original Granite tokenizer and must preserve the same pooling and normalization semantics used by the consuming application.
Evaluation and Validation Snapshot
The table below mirrors the coverage of the official IBM Granite model card's evaluation section, but separates upstream model-quality scores from local Core ML conversion evidence. Not run means this Core ML package has not yet been evaluated on that benchmark family.
The local 57.73 value is the simple unweighted task-level mean from this fixed-shape Core ML run multiplied by 100. It is useful as a caveated comparison point for this Core ML artifact, but it is not a recipe-equivalent reproduction of the official 60.3 model-card result.
For throughput context, IBM reports 2,534 docs/s on a single H100 using 512-token chunks. On a local Apple M4 Mac mini, the recommended sequence-512 Core ML ANE package measured 121.1 512-token windows/s in the 1024-chunk release run. A later targeted quiet check using three 256-chunk repeats measured a mean of 122.2 512-token windows/s, with a range of 121.5 to 123.6. This is roughly 4.8% of the H100 throughput, or about 21x slower, while running locally on Apple Silicon without a datacenter GPU.
Core ML Artifact Validation
Core ML / Hugging Face Parity
The sequence-512 package was also checked against the original Hugging Face model on a small local ranking fixture:
Local Multilingual Retrieval Details
Full local Core ML run:
Per-task local main scores:
These measurements are local validation results for the conversion artifacts, not official benchmark claims. Validate placement and throughput in the target application process before using the packages for performance comparisons.
License and Attribution
The original model card lists the base model license as Apache 2.0. This conversion repository follows that license metadata and attributes the base model to IBM Granite. For the official model, documentation, and limitations, refer to the original IBM Granite model card.
