CoolFace
Modelpublic

atjsh/llmlingua-2-js-mobilebert-meetingbank-onnx-v4

sourceHugging Faceunknownupdated 1mo agoView on Hugging Face
0likes175downloads
Model Card

MobileBERT MeetingBank — ONNX v4 conversion archive

This repository is a neutral format-conversion archive of `JueZhang/lingua2_mobilebert_meetingbank_only` at source revision af78299502cbbb512ce8a664ff4b7ce7e6ff7e0a.

The source model card does not declare a license, so this repository records license: unknown and does not add license permissions.

Conversion provenance

  • —Official converter: `onnx-community/convert-to-onnx` at af0ef3070cef5e863b371628ad3dcdc10de4d091
  • —Export: optimum-cli export onnx --model <pinned-local-source> --task token-classification <output-directory>
  • —Variant generation: the pinned converter's ModelConverter._apply_quantizations
  • —Resolved environment: `provenance/environment.txt`
  • —Recorded commands and source-staging disclosure: `provenance/commands.txt`
  • —Converter source and requirements: `provenance/converter/`
  • —Complete export and variant logs: `provenance/logs/`

The source model.safetensors was 98,470,112 bytes with SHA-256 7c6e38183447c5d7846a10fbac680c942ee852584f95e9a0ad3077bbb1257849. All source-file hashes are in `provenance/source.json`. The pinned source also contains a redundant state_dict.pth; its hash is recorded there, while the export selected model.safetensors.

Emitted ONNX files

Every row below was emitted by a successful pinned converter invocation.

dtypefilenamebytesSHA-256
fp32onnx/model.onnx98,960,7608902795483843cd2645ec5842980be4371a341bd774b5ec45f0804622b010d38
fp16onnx/model_fp16.onnx50,043,69026a46969486ab482c5595fb0d0552e65362ef1dd35d8860f5b6126105876985d
int8onnx/model_int8.onnx26,470,22075ac8650441941c07caa3241fcb11baeedff54514e04fe8a05343c5c7e8ca309
uint8onnx/model_uint8.onnx26,470,4168963754a429a9e9c30dd04281303c1f89d4a061fe0010e4857b811be3293de74
q8onnx/model_quantized.onnx26,470,22075ac8650441941c07caa3241fcb11baeedff54514e04fe8a05343c5c7e8ca309
q4onnx/model_q4.onnx30,665,261c3f7b96daf2eae0c3c4025c646dabb321c8c54676584f9f3faf7444be77e5f73
q4f16onnx/model_q4f16.onnx20,917,291f69b66eb622b26aaefbaa39874992d21e5f0914087c3593c45fd373655d70b91
bnb4onnx/model_bnb4.onnx30,665,261c3f7b96daf2eae0c3c4025c646dabb321c8c54676584f9f3faf7444be77e5f73

Observed runtime result

An ONNX Runtime 1.28.0 CPU smoke on Apple M4 loaded fp32, int8, uint8, q8, q4, and bnb4. Each returned finite [1, 14, 2] logits and repeated exactly three times. fp16 and q4f16 failed to load because a converted float16 output did not match an expected float tensor type. Exact timings and errors are in `measurements/runtime-cpu.json`.

The official exporter also recorded a maximum logits difference warning for the fp32 graph. The complete text is preserved in `provenance/logs/export.log`. These observations are measurements only. The files are published because the pinned official converter emitted them successfully.

Integrity

`manifest.json` and `SHA256SUMS` record the published file inventory. Conversion logs include all converter warnings.

Recorded demo Pareto selection

The static demo dtype selection produced by the recorded formula is bnb4 for WebGPU and bnb4 for WASM. This is a product-selection record, not a qualification, recommendation, commercial-suitability statement, or backend-support claim.

Inputs, raw measurements, the exact harness, and the concise selector output are under `measurements/pareto/`. The measured model bytes came from repository revision 5bd77cc8a968650b2e5fb6bffaf75f66759feb0f.