litert-community/granite-embedding-311m-multilingual-r2
3156
Card: LiteRT-LM EmbeddingEngine bundles section (litert-lm >= 0.17.0) — files, Python/Kotlin usage, insert_special_tokens note, measured latency on a Galaxy S26 and a Mac
Add LiteRT-LM EmbeddingEngine bundle (granite-embedding-311m-r2_fp16.litertlm) — litert-lm >= 0.17.0; same weights and quantization as the .tflite, vectors identical
Add LiteRT-LM EmbeddingEngine bundle (granite-embedding-311m-r2_wi8fc.litertlm) — litert-lm >= 0.17.0; same weights and quantization as the .tflite, vectors identical
Card: base_model_relation: quantized (list under the base model's Quantizations)
Add measured Galaxy S26 NPU/GPU section
granite-embedding-311m-multilingual-r2 converted to LiteRT (int8 + fp16, iPhone-verified bit-exact)
initial commit
