datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fish-audio-s2-sglang-omni-models
Fish Audio S2 models
This Kaggle dataset contains a fishaudio/s2-pro checkpoint snapshot for sglang-omni PHRunner Fish Audio S2 inference.
Source repository: fishaudio/s2-pro
Revision: main
Layout: s2-pro-sglang
Required backend: sglang-omni
Default model file: model-00001-of-00002.safetensors
Default codec file: codec.pth
Files: 10
The runner expects this directory to be mounted as a Kaggle input dataset. Official Python/SGLang layouts are validated through config.json… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/fish-audio-s2-sglang-omni-models.higgs-audio-v3-tts-4b-vllm-omni-models
Higgs Audio v3 TTS model payload
This dataset was prepared by HiggsAudioModelsUploader.
Source repository: bosonai/higgs-tts-3-4b
Revision: main
Layout: vllm-omni
Required backend: higgs-audio-v3-vllm-omni
Default model file: model.safetensors.index.json
Contains Transformers trust_remote_code files: False
Files: 9
Expected Kaggle runner environment:
HIGGS_AUDIO_MODEL_DATASET_REF=/higgs-audio-v3-tts-4b-vllm-omni-models
HIGGS_AUDIO_BACKEND=vllm-omni
Reference-speaker audio is… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/higgs-audio-v3-tts-4b-vllm-omni-models.higgs-audio-v3-tts-4b-sglang-omni-models
Higgs Audio v3 TTS model payload
This dataset was prepared by HiggsAudioModelsUploader.
Source repository: bosonai/higgs-tts-3-4b
Revision: main
Layout: sglang-omni
Required backend: higgs-audio-v3-sglang-omni
Default model file: model.safetensors.index.json
Contains Transformers trust_remote_code files: False
Files: 9
Expected Kaggle runner environment:
HIGGS_AUDIO_MODEL_DATASET_REF=/higgs-audio-v3-tts-4b-sglang-omni-models
HIGGS_AUDIO_BACKEND=sglang-omni
Reference-speaker… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/higgs-audio-v3-tts-4b-sglang-omni-models.modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO
OminiGAIA-DPO-data
This dataset contains the final DPO training pairs used to train
ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO.
The pairs are derived from 7B-native rollouts on OmniGAIA train questions:
Roll out the SFT model on answer-hidden train inputs.
Audit each rollout with Gemini using the private reference answer and annotated solution.
Locate the first erroneous assistant sub-step.
Convert the corrected prefix (tau_win) and the original erroneous prefix… See the full description on the dataset page: https://huggingface.co/datasets/ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO.qwen4b-thinking-model_rewrite-omni-l1_4-dpo-pairs
