datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
callcc-2k-hours-train-soniox
ErfanRou/callcc-2k
Persian call-centre ASR silver set (2,000 audio-hours, 2,288 h shipped): the ErfanRou/callcc-ft150-soniox schema plus two auxiliary columns.
Machine-labelled, not ground truth. Built by callcc-silver-5k/soniox_pipeline.py from markmuller/call-center-prod-data:
8 kHz stereo telephony split client-side into agent = channel 0 and customer = channel 1, each channel transcribed
separately by Soniox stt-async-v5, words packed into 5-28 s windows (15 s mean, natural… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/callcc-2k-hours-train-soniox.callcc-test-1k-samples-soniox
ErfanRou/callcc-test-1k
Full-channel benchmark set for Persian call-centre ASR: one row per channel of a call (0 = agent, 1 = customer),
with the complete 16 kHz mono channel audio and the complete Soniox stt-async-v5 transcript of that channel,
rebuilt from the raw tokens of ErfanRou/callcc-test-1k-windowed (no re-transcription). Use it to evaluate the serving path
(whole-channel input, model-side VAD/chunking) with corpus-level WER/CER — see eval_full_call.py in the kit.
text… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/callcc-test-1k-samples-soniox.callcc-150h-soniox
