datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ciempiess_balance
Dataset Card for ciempiess_balance
Dataset Summary
The CIEMPIESS BALANCE Corpus is designed to match with the CIEMPIESS LIGHT Corpus (LDC2017S23). So, "Balance" means that if the CIEMPIESS BALANCE is combined with the CIEMPIESS LIGHT, one will get a gender balanced corpus. To appreciate this, one need to know that the CIEMPIESS LIGHT is by itself, a gender unbalanced corpus of approximately 25% of female speakers and 75% of male speakers. So, the CIEMPIESS BALANCE is a… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_balance.qwen3-omni-open-source-balanced-1200
Qwen3-Omni 30A3 open-source balanced subset
This dataset contains 1,200 samples selected from public-source-labelled portions of the Qwen3-Omni 30A3 posttrain recipe. The 8 topics are balanced at 150 samples each. Every item includes topic, public_dataset, public_dataset_confidence, source_id, and provenance_json fields. public_dataset is the canonical per-item public-dataset label.
Loading
The data/train-*.jsonl shards are ordinary Hugging Face JSONL data files… See the full description on the dataset page: https://huggingface.co/datasets/Transl/qwen3-omni-open-source-balanced-1200.phoneme-ctc-english-60h-balanced
Phoneme CTC — English 60h (Balanced & Normalized)
A cleaned, normalized and phoneme-balanced version of
bobboyms/phoneme-ctc-english-60h-noisy,
for training phoneme recognition models (CTC) — e.g. as the native acoustic
model behind pronunciation-feedback systems.
What's different from the source dataset
Label noise removed
Roman numerals dropped — eSpeak reads ii/iv/… as "Roman two/four",
producing labels that don't match the audio.
Non-English phonemes dropped… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/phoneme-ctc-english-60h-balanced.
