bh-healthcare/distilbart-mnli-12-3-int8-onnx
BART-Large-MNLI INT8 ONNX (bh-sentinel pinned artifact)
Quantized ONNX export of `facebook/bart-large-mnli` at commit d7645e127eaf1aefc7862fd59a17a5aa8558b8ce. Hosted as the canonical Layer 2 model for `bh-sentinel-ml` ≥ 0.2.1.
The repository namedistilbart-mnli-12-3-int8-onnxis historical — chosen during the v0.2.1 planning phase when the intended source was the smaller distilled variantvalhalla/distilbart-mnli-12-3. During the release's license verification gate (see provenance doc), that source was found to have no declared license and was rejected; the teacher modelfacebook/bart-large-mnli(explicit MIT) was used as the fallback. The repo name is preserved so downstream config references inbh-sentinel-mlcontinue to resolve. The actual source model is BART-Large fine-tuned on MultiNLI — see Provenance below.
Clinical Use Notice
This is clinical decision support software. It is not a diagnostic tool, not FDA-cleared, and not a substitute for clinical judgment. The bh-sentinel pipeline that consumes this artifact is intended only for flagging signals for clinician review — never for autonomous clinical action. All outputs are signals for clinician review. See the main repository's clinical disclaimer for the full notice.
License Chain
The redistribution is permitted under the upstream MIT terms. The original copyright notice is preserved verbatim in LICENSE. The Python package that loads this artifact is Apache-2.0 — different copyrightable works can carry different terms; this is a common open-source pattern and is not contradictory.
Provenance
To reproduce locally, install optimum[onnxruntime]>=1.16 and onnx>=1.15 and run the export script against the source revision above. The artifact is fully reproducible from those inputs.
Size & Performance
Quantization compresses MatMul weights from FP32 to INT8 (~4x for the quantized ops). BART-Large has ~406M parameters; ONNX FP32 export is ~1.6GB, INT8 quantized is 391 MB.
These numbers are operator-reported approximations for capacity planning. Real-world latency varies with CPU class, ARM vs x86, ONNX Runtime threading config, and concurrent load. The bh-sentinel-ml pipeline supports BH_SENTINEL_ML_OFFLINE=1 for fully air-gapped VPC deployments where the model is pre-baked into the container image.
Why not a smaller model? The originally-intended source, valhalla/distilbart-mnli-12-3 (~140MB INT8, ~30ms/pair), was rejected by the license verification gate. We will revisit smaller MIT-licensed NLI alternatives (e.g. fine-tuned DeBERTa-v3-base variants) when validated clinical calibration is in scope (bh-sentinel v0.3+).Files
Intended Use
Consumed by bh-sentinel-ml ≥ 0.2.1 via Pipeline(enable_transformer=True). The bh-sentinel pipeline performs zero-shot NLI inference against a curated set of clinical hypothesis templates (see `config/ml/zero_shot_hypotheses.yaml`) to surface candidate clinical safety flags.
Use cases this artifact is designed for:
- Layer 2 (semantic) detection in the bh-sentinel multi-layer pipeline
- Catching clinical signals that pattern-matching alone misses: indirect language, implied distress, contextual meaning
- Reproducible local evaluation of bh-sentinel against the shared corpus in `bh-sentinel-examples`
Out-of-Scope Use
This artifact is not appropriate for:
- Direct clinical diagnosis or treatment decisions
- Autonomous clinical action of any kind
- Use as the sole basis for any clinical decision
- Inference tasks other than zero-shot NLI (it has not been retrained for other tasks)
- Production deployments without independent clinical validation against your organization's de-identified data
- Fine-tuning starting points for high-stakes clinical applications without further validation
Limitations
- Trained on general-purpose MNLI, not clinical text. Confidence calibration on clinical inputs is dampened via
FixedDiscount(0.85)inbh-sentinel-ml's Phase A calibrator, but is not clinically validated. Validated calibration against clinician-labeled data is a bh-sentinel v0.3 deliverable. - INT8 dynamic quantization introduces small numerical drift relative to the FP32 source. The bh-sentinel test suite includes a corpus-level round-trip check (see `bh-sentinel-examples`) but does not guarantee bit-exact equivalence to FP32.
- English only. The MNLI training data is English-language; behavior on other languages is undefined and unsafe.
- No clinical labels in training. The model has never seen clinician-labeled examples of behavioral health safety signals. Its outputs are best understood as plausibility scores against pre-defined hypothesis templates, not as clinical probabilities.
Citation
If you use this artifact, please cite both the upstream source and bh-sentinel:
@misc{bart-large-mnli,
title = {bart-large-mnli},
author = {Facebook AI Research},
year = {2019},
publisher = {Hugging Face},
url = {https://huggingface.co/facebook/bart-large-mnli}
}
@article{lewis2019bart,
title = {{BART}: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension},
author = {Lewis, Mike and Liu, Yinhan and Goyal, Naman and Ghazvininejad, Marjan and Mohamed, Abdelrahman and Levy, Omer and Stoyanov, Veselin and Zettlemoyer, Luke},
journal = {arXiv preprint arXiv:1910.13461},
year = {2019},
url = {https://arxiv.org/abs/1910.13461}
}
@article{yin2019benchmarking,
title = {Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach},
author = {Yin, Wenpeng and Hay, Jamaal and Roth, Dan},
journal = {arXiv preprint arXiv:1909.00161},
year = {2019},
url = {https://arxiv.org/abs/1909.00161}
}
@software{bh_sentinel,
title = {bh-sentinel: Open-source clinical safety signal detection for behavioral health systems},
year = {2026},
publisher = {bh-healthcare},
url = {https://github.com/bh-healthcare/bh-sentinel}
}Acknowledgments
- The source model `facebook/bart-large-mnli` was authored by Facebook AI Research (now Meta AI) and is redistributed here under its original MIT license.
- The zero-shot NLI classification method was proposed by Yin et al. 2019; HF's blog post by Joe Davison (May 2020) popularized the pattern this artifact serves.
- The
optimumlibrary by Hugging Face provides the ONNX export path used here.
