datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Belle_1.4M-SLAM-Omni
Belle_1.4M
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round Chinese spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/Belle_1.4M-SLAM-Omni.UltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/UltraChat-300K-SLAM-Omni.Belle_1.4M-SLAM-Omni
Belle_1.4M
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round Chinese spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/Belle_1.4M-SLAM-Omni.VoiceAssistant-400K-SLAM-Omni
VoiceAssistant-400K (Modified)
This dataset is prepared for the reproduction of SLAM-Omni.
This is a single-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/VoiceAssistant-400K-SLAM-Omni.UltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/UltraChat-300K-SLAM-Omni.VoiceAssistant-400K-SLAM-Omni
VoiceAssistant-400K (Modified)
This dataset is prepared for the reproduction of SLAM-Omni.
This is a single-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These… See the full description on the dataset page: https://huggingface.co/datasets/mwei/VoiceAssistant-400K-SLAM-Omni.
