CoolFace
Modelpublic

FredrikKarlssonSpeech/megatron-bert-large-swedish-cased-165k-onnx

sourceHugging Faceunknownupdated 2mo agoView on Hugging Face
0likes
Model Card

megatron-bert-large-swedish-cased-165k - ONNX

ONNX export of `KBLab/megatron-bert-large-swedish-cased-165k`.

This package contains two task-specific exports:

  • —fill-mask/onnx/model.onnx for masked language modeling logits
  • —feature-extraction/onnx/model.onnx for last_hidden_state and cls_embedding

Available variants

Each task folder contains:

  • —model.onnx - FP32 baseline
  • —model_fp16.onnx - FP16, recommended
  • —model_int8.onnx - dynamic INT8
  • —model_uint8.onnx - dynamic UINT8
  • —model_q4.onnx - 4-bit MatMul quantization

File sizes

Fill-mask

  • —model.onnx: 1741.6 MB
  • —model_fp16.onnx: 871.0 MB
  • —model_int8.onnx: 437.2 MB
  • —model_uint8.onnx: 437.2 MB
  • —model_q4.onnx: 497.2 MB

Feature-extraction

  • —model.onnx: 1474.4 MB
  • —model_fp16.onnx: 737.4 MB
  • —model_int8.onnx: 370.2 MB
  • —model_uint8.onnx: 370.2 MB
  • —model_q4.onnx: 455.2 MB

Accuracy summary

FP32 parity vs PyTorch

  • —Fill-mask max logit diff: 0.000237
  • —Fill-mask top-5 tokens: exact match
  • —Feature last_hidden_state max diff: 0.000011
  • —Feature cls_embedding cosine similarity: 1.0

Quantized variants vs FP32 ONNX

Fill-mask
  • —fp16: top-5 exact match, max diff 0.0173
  • —int8: top-5 drift after rank 2
  • —uint8: top-5 drift after rank 2
  • —q4: top-5 drift at rank 5
Feature-extraction
  • —fp16: CLS cosine 0.9999998
  • —q4: CLS cosine 0.9882
  • —int8: CLS cosine 0.9282
  • —uint8: CLS cosine 0.9298

Recommendation

  • —Use model_fp16.onnx by default.
  • —For feature extraction on tight storage budgets, model_q4.onnx is plausible but lower fidelity than fp16.
  • —Avoid int8 and uint8 for fill-mask if token ranking fidelity matters.

Layout

text
megatron-bert-large-swedish-cased-165k/
├── fill-mask/
│   ├── config.json
│   ├── tokenizer.json
│   ├── tokenizer_config.json
│   ├── special_tokens_map.json
│   └── onnx/
│       ├── model.onnx
│       ├── model_fp16.onnx
│       ├── model_int8.onnx
│       ├── model_uint8.onnx
│       └── model_q4.onnx
└── feature-extraction/
    ├── config.json
    ├── tokenizer.json
    ├── tokenizer_config.json
    ├── special_tokens_map.json
    └── onnx/
        ├── model.onnx
        ├── model_fp16.onnx
        ├── model_int8.onnx
        ├── model_uint8.onnx
        └── model_q4.onnx

Export notes

  • —Native optimum 2.1.0 task export does not support megatron-bert.
  • —These graphs were exported with direct torch.onnx.export wrappers using legacy TorchScript exporter (dynamo=False).
  • —Feature extraction export uses cls_embedding = last_hidden_state[:, 0, :].
  • —No pooler output included, because loading MegatronBertModel from this MLM checkpoint would introduce randomly initialized pooler weights.