datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SALMon_Flow-SLM-1B-Extended
SALMon Normalized Dataset
This repo preserves the SALMon per-config folder layout while normalizing
mismatched schema details across model families.
SALMon_Flow-SLM-1B
SALMon Normalized Dataset
This repo preserves the SALMon per-config folder layout while normalizing
mismatched schema details across model families.
SALMon_Flow-SLM-1B-depasr-1b0e733bSALMon_Flow-SLM-1B-Extended-dephow-people-make-money-csm1beval-canary-1b-v2-eka-hard-20260408-1921
Evaluation Results: canary-1b-v2
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/canary-1b-v2
39.85%
22.43%
Source Data
Evaluation Dataset: Trelis/eka-hard
Model Evaluated: nvidia/canary-1b-v2
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for this sample
cer… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-canary-1b-v2-eka-hard-20260408-1921.eval-canary-1b-v2-medical-terms-2025-20260408-1926
Evaluation Results: canary-1b-v2
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/canary-1b-v2
10.11%
3.26%
Source Data
Evaluation Dataset: Trelis/medical-terms-2025
Model Evaluated: nvidia/canary-1b-v2
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for this sample… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-canary-1b-v2-medical-terms-2025-20260408-1926.xlsr2_1b_v2_kmeans_10k_Llama-3.2-3B-Instruct_all_LJSpeech_TTS_v9xlsr2_1b_v2_kmeans_10k_Llama-3.2-11B-Vision-Instruct_all_ls960_TTS_v11eval-canary-1b-v2-multimed-hard-20260408-1931
Evaluation Results: canary-1b-v2
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/canary-1b-v2
15.02%
9.27%
Source Data
Evaluation Dataset: Trelis/multimed-hard
Model Evaluated: nvidia/canary-1b-v2
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for this sample
cer… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-canary-1b-v2-multimed-hard-20260408-1931.xlsr2_1b_v2_kmeans_10k_Llama-3.2-11B-Vision-Instruct_interleaf_last_5_selfattn_ls960_TTS_v12dac_inference_10200_1Bdac_inference_1B_GRPO_V30dac_inference_1B_SUPERdac_inference_1B_TBD-LLaMA-DAC-Denoiser-checkpoint-7000dac_inference_1B_17200dac_inference_1B_14400inference_long_1B_DAC-SE2_RoPE100K_finalinference_long_1B_DAC-SE2_RoPE100K_final_TEST1B-DAC-SE2_1B_np_UPSAMPLE_checkpoint_op_v3inference_long_1B_DAC-SE2_RoPE100K_final_TEST2inference_long_1B_DAC-SE2_RoPE100K_final_TEST2_no_overlapinference_long_1B_DAC-SE2_RoPE100K_final_TEST2_01_overlapinference_long_1B_DAC-SE2_RoPE100K_final_TEST2_no_overlap_silence_detinference_long_1B_DAC-SE2_RoPE100K_final_17_no_overlap_silence_det1B1B_natural_noise_fine_tuned-checkpoint-2600_correct_codebook
