FredrikKarlssonSpeech/megatron-bert-large-swedish-cased-165k-onnx
0
megatron-bert-large-swedish-cased-165k - ONNX
ONNX export of `KBLab/megatron-bert-large-swedish-cased-165k`.
This package contains two task-specific exports:
fill-mask/onnx/model.onnxfor masked language modeling logitsfeature-extraction/onnx/model.onnxforlast_hidden_stateandcls_embedding
Available variants
Each task folder contains:
model.onnx- FP32 baselinemodel_fp16.onnx- FP16, recommendedmodel_int8.onnx- dynamic INT8model_uint8.onnx- dynamic UINT8model_q4.onnx- 4-bit MatMul quantization
File sizes
Fill-mask
model.onnx: 1741.6 MBmodel_fp16.onnx: 871.0 MBmodel_int8.onnx: 437.2 MBmodel_uint8.onnx: 437.2 MBmodel_q4.onnx: 497.2 MB
Feature-extraction
model.onnx: 1474.4 MBmodel_fp16.onnx: 737.4 MBmodel_int8.onnx: 370.2 MBmodel_uint8.onnx: 370.2 MBmodel_q4.onnx: 455.2 MB
Accuracy summary
FP32 parity vs PyTorch
- Fill-mask max logit diff:
0.000237 - Fill-mask top-5 tokens: exact match
- Feature
last_hidden_statemax diff:0.000011 - Feature
cls_embeddingcosine similarity:1.0
Quantized variants vs FP32 ONNX
Fill-mask
fp16: top-5 exact match, max diff0.0173int8: top-5 drift after rank 2uint8: top-5 drift after rank 2q4: top-5 drift at rank 5
Feature-extraction
fp16: CLS cosine0.9999998q4: CLS cosine0.9882int8: CLS cosine0.9282uint8: CLS cosine0.9298
Recommendation
- Use
model_fp16.onnxby default. - For feature extraction on tight storage budgets,
model_q4.onnxis plausible but lower fidelity than fp16. - Avoid
int8anduint8for fill-mask if token ranking fidelity matters.
Layout
megatron-bert-large-swedish-cased-165k/
├── fill-mask/
│ ├── config.json
│ ├── tokenizer.json
│ ├── tokenizer_config.json
│ ├── special_tokens_map.json
│ └── onnx/
│ ├── model.onnx
│ ├── model_fp16.onnx
│ ├── model_int8.onnx
│ ├── model_uint8.onnx
│ └── model_q4.onnx
└── feature-extraction/
├── config.json
├── tokenizer.json
├── tokenizer_config.json
├── special_tokens_map.json
└── onnx/
├── model.onnx
├── model_fp16.onnx
├── model_int8.onnx
├── model_uint8.onnx
└── model_q4.onnxExport notes
- Native
optimum 2.1.0task export does not supportmegatron-bert. - These graphs were exported with direct
torch.onnx.exportwrappers using legacy TorchScript exporter (dynamo=False). - Feature extraction export uses
cls_embedding = last_hidden_state[:, 0, :]. - No pooler output included, because loading
MegatronBertModelfrom this MLM checkpoint would introduce randomly initialized pooler weights.
