soniox
Datasets
All datasets matching “soniox”callcc-2k-hours-train-soniox
ErfanRou/callcc-2k
Persian call-centre ASR silver set (2,000 audio-hours, 2,288 h shipped): the ErfanRou/callcc-ft150-soniox schema plus two auxiliary columns.
Machine-labelled, not ground truth. Built by callcc-silver-5k/soniox_pipeline.py from markmuller/call-center-prod-data:
8 kHz stereo telephony split client-side into agent = channel 0 and customer = channel 1, each channel transcribed
separately by Soniox stt-async-v5, words packed into 5-28 s windows (15 s mean, natural… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/callcc-2k-hours-train-soniox.details_soniox__Soniox-7B-v1.0
Dataset Card for Evaluation run of soniox/Soniox-7B-v1.0
Dataset automatically created during the evaluation run of model soniox/Soniox-7B-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_soniox__Soniox-7B-v1.0.callcc-test-1k-samples-soniox
ErfanRou/callcc-test-1k
Full-channel benchmark set for Persian call-centre ASR: one row per channel of a call (0 = agent, 1 = customer),
with the complete 16 kHz mono channel audio and the complete Soniox stt-async-v5 transcript of that channel,
rebuilt from the raw tokens of ErfanRou/callcc-test-1k-windowed (no re-transcription). Use it to evaluate the serving path
(whole-channel input, model-side VAD/chunking) with corpus-level WER/CER — see eval_full_call.py in the kit.
text… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/callcc-test-1k-samples-soniox.youtube-300h-movies-soniox
Persian Speech Corpus — full three-pool release
309.72 hours · 157,279 clips · 736 source videos · Soniox transcripts on every clip.
This is the complete quality-gated output of the persian-expressive-corpus
pipeline. It is organised into three mutually exclusive pools. Read the pool
column before using a clip — they are not interchangeable.
pool
clips
hours
transcript
emotion label
QC status
recommended use
A
66,108
108.34
yes
yes, 7-class
passed all gates
expressive… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/youtube-300h-movies-soniox.persian-youtube-330h-soniox
persian-youtube-330h-soniox
Persian conversational / voice-search / meeting ASR fine-tuning set built from
YouTube audio, labelled by Soniox stt-async-v5. Machine-labelled. Not ground
truth. Built 2026-09-22 by ytcrawl/build_dataset.py.
What this is
74,538 training segments (331.5 h) and 7,442 validation
segments (32.9 h), 16 kHz mono FLAC, 5–28 s each, cut from
2,708 videos on 106 channels across 9 domains.
Validation is held out by channel (a creator is never on… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/persian-youtube-330h-soniox.callcc-150h-soniox
