datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moshi-on-policy-prompts-3kBased on https://huggingface.co/datasets/yhytoto12/behavior-sd
moshi-on-policy-dpo-margin3manifest-policy-unified
Unified manifest_policy ASR dataset
Combined Vietnamese ASR dataset from multiple manifest_policy sources for Qwen3-ASR.
Audio mode: embed
Target total hours: 1000.0
Splits
Split
Examples
Duration (h)
train
731,135
916.4
test
43,032
83.6
Sampling
test: all test.jsonl rows from included datasets
train: random sample from all other JSONL splits until train + test reaches target hours
Load
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/tbntbn/manifest-policy-unified.moshi-on-policy-dpo-v17-partialmoshi-on-policy-dpo-v20-kyutai-smokemoshi-on-policy-dpo-v20-kyutai-alignedmoshi-on-policy-dpo-9kmoshi-on-policy-dpo-tts-v18-fullprompt-smokemoshi-on-policy-dpo-9k-v17moshi-on-policy-dpo-tts-v18moshi-on-policy-dpo-tts-v18-fullpromptmoshi-on-policy-dpo-v20-kyutai-aligned-smoke
