Masterx/cohere-transcribe-arabic-07-2026-ONNX
Add cross_bias cross-attention mask input to all decoders (bias-only injection)
Add fp16 variant: encoder + static/dyn decoders (verified packer; fp32 transcript parity; ~3.5x faster wall on DML)
Drop q4 quant: int8 strictly dominates (smaller AND more accurate) for this checkpoint
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Upload README.md with huggingface_hub
Add files using upload-large-folder tool
Upload README.md with huggingface_hub
Delete onnx/decoder_model_merged.onnx_data with huggingface_hub
Delete onnx/decoder_model_merged.onnx with huggingface_hub
Delete onnx/encoder_model.onnx with huggingface_hub
Upload onnx/encoder_model.onnx with huggingface_hub
Upload onnx/decoder_model_merged.onnx_data with huggingface_hub
Upload verify.py with huggingface_hub
Upload quantize.py with huggingface_hub
Upload export_merged.py with huggingface_hub
Upload onnx/decoder_model_merged.onnx with huggingface_hub
Upload README.md with huggingface_hub
Upload special_tokens_map.json with huggingface_hub
Upload tokenizer_config.json with huggingface_hub
Upload tokenizer.json with huggingface_hub
Upload preprocessor_config.json with huggingface_hub
Upload generation_config.json with huggingface_hub
Upload config.json with huggingface_hub
Upload onnx/decoder_model_merged_q4.onnx_data with huggingface_hub
Upload onnx/decoder_model_merged_q4.onnx with huggingface_hub
Upload onnx/encoder_model_q4.onnx_data with huggingface_hub
Upload onnx/encoder_model_q4.onnx with huggingface_hub
initial commit
