atjsh/llmlingua-2-js-xlm-roberta-large-meetingbank-onnx-v4
LLMLingua-2 XLM-R ONNX v4 conversion archive
This is a neutral, reproducible format-conversion archive of `microsoft/llmlingua-2-xlm-roberta-large-meetingbank` at source revision ebaba9b0e874dadd3003ffcff828e4397e568089.
The source repository declares MIT. The retained license and conversion notice are in LICENSE and NOTICE. These files were not retrained.
Provenance
- Converter reference:
onnx-community/convert-to-onnx@af0ef3070cef5e863b371628ad3dcdc10de4d091 - Export:
optimum-cli export onnx --model <pinned-local-source> --task token-classification <output-directory> - Variants:
python conversion/quantize.py onnx/model.onnx - Resolved environment:
conversion/requirements.txt - Converter output:
conversion/export.logandconversion/quantize.log - Machine-readable file sizes and SHA-256 hashes:
conversion/manifest.json
The export command exited successfully and reported a maximum reference/ONNX logit difference of 7.2479248046875e-05, above its 1e-05 validation tolerance. The warning is preserved in conversion/export.log and does not alter this archive's converter-trust publication rule.
The q8 dtype maps to onnx/model_quantized.onnx, an exact copy of the official recipe's signed int8 output.
Artifacts
Observed runtime result
conversion/runtime-cpu.json records a three-run ONNX Runtime 1.28.0 CPU smoke on an Apple M4. fp32, int8, uint8, q8, q4, and bnb4 loaded, returned finite [1, 13, 2] logits, and repeated byte-identically. fp16 failed because its two-byte emitted graph has no opset; q4f16 failed with an ONNX type mismatch. These are observations from that run, not a runtime-support, numerical-fidelity, or product-suitability claim.
