CoolFace
Modelpublic

atjsh/llmlingua-2-js-xlm-roberta-large-meetingbank-onnx-v4

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes561downloads
Model Card

LLMLingua-2 XLM-R ONNX v4 conversion archive

This is a neutral, reproducible format-conversion archive of `microsoft/llmlingua-2-xlm-roberta-large-meetingbank` at source revision ebaba9b0e874dadd3003ffcff828e4397e568089.

The source repository declares MIT. The retained license and conversion notice are in LICENSE and NOTICE. These files were not retrained.

Provenance

  • —Converter reference: onnx-community/convert-to-onnx@af0ef3070cef5e863b371628ad3dcdc10de4d091
  • —Export: optimum-cli export onnx --model <pinned-local-source> --task token-classification <output-directory>
  • —Variants: python conversion/quantize.py onnx/model.onnx
  • —Resolved environment: conversion/requirements.txt
  • —Converter output: conversion/export.log and conversion/quantize.log
  • —Machine-readable file sizes and SHA-256 hashes: conversion/manifest.json

The export command exited successfully and reported a maximum reference/ONNX logit difference of 7.2479248046875e-05, above its 1e-05 validation tolerance. The warning is preserved in conversion/export.log and does not alter this archive's converter-trust publication rule.

The q8 dtype maps to onnx/model_quantized.onnx, an exact copy of the official recipe's signed int8 output.

Artifacts

DtypeFileBytesSHA-256
fp32 graphonnx/model.onnx488,5182cddfd767604e06699123e2640178ef3231699b6362fda364fa5152ff7baae4e
fp32 external dataonnx/model.onnx_data2,235,785,232a957df2691421facdac43eb4df22598980345f2e8c7da5a467a91e18b674ce63
fp16onnx/model_fp16.onnx27648eb81c4e9fef7b51879c6fcd6fb4b8863eea0cb2a978c508711dc52c80041
int8onnx/model_int8.onnx560,745,894e2ee14513c0527822c3701f4adbae16e3e95463cb05d7058cbef79ab3b396e31
uint8onnx/model_uint8.onnx560,745,966716baafeff90f70a94e945c7824e5c0456508ba8a744d30c07dbb3b997ef02ec
q8onnx/model_quantized.onnx560,745,894e2ee14513c0527822c3701f4adbae16e3e95463cb05d7058cbef79ab3b396e31
q4onnx/model_q4.onnx1,217,047,880cb14ee1a4145f7388aae32130ff528c498507a455dcb19b3206639f8893d7a57
q4f16onnx/model_q4f16.onnx684,410,6037ce164c4ac2b5f9d0b06d52a364efaedaff05b2224816bd8134b71e7d0ec17a7
bnb4onnx/model_bnb4.onnx1,217,047,880cb14ee1a4145f7388aae32130ff528c498507a455dcb19b3206639f8893d7a57

Observed runtime result

conversion/runtime-cpu.json records a three-run ONNX Runtime 1.28.0 CPU smoke on an Apple M4. fp32, int8, uint8, q8, q4, and bnb4 loaded, returned finite [1, 13, 2] logits, and repeated byte-identically. fp16 failed because its two-byte emitted graph has no opset; q4f16 failed with an ONNX type mismatch. These are observations from that run, not a runtime-support, numerical-fidelity, or product-suitability claim.