CoolFace
Datasetpublic

ga642381/TiCo-Bench-outputs-Qwen3-Omni-GRPO-ckpt600

TiCo (Qwen3-Omni-30B-A3B) GRPO checkpoint-600: TiCo-Bench outputs Inference outputs of the TiCo model built on Qwen3-Omni-30B-A3B-Instruct (SFT LoRA ga642381/TiCo-Qwen3-Omni-SFT-LoRA, then GRPO + CHORD on the 4,000 exact-duration prompts, checkpoint at step 600, LoRA merged) on the 2,000-sample TiCo-Bench speech-query benchmark (WeiChihChen/TiCo-Bench). For each benchmark id there are two files: <id>.wav: the generated speech (24 kHz mono). <id>.txt: the full prompt and the… See the full description on the dataset page: https://huggingface.co/datasets/ga642381/TiCo-Bench-outputs-Qwen3-Omni-GRPO-ckpt600.

sourceHugging Facecc-by-nc-4.0updated 13d agoView on Hugging Face
0likes339downloads
Dataset Card

TiCo (Qwen3-Omni-30B-A3B) GRPO checkpoint-600: TiCo-Bench outputs

Inference outputs of the TiCo model built on Qwen3-Omni-30B-A3B-Instruct (SFT LoRA ga642381/TiCo-Qwen3-Omni-SFT-LoRA, then GRPO + CHORD on the 4,000 exact-duration prompts, checkpoint at step 600, LoRA merged) on the 2,000-sample TiCo-Bench speech-query benchmark (WeiChihChen/TiCo-Bench).

For each benchmark id there are two files:

  • —<id>.wav: the generated speech (24 kHz mono).
  • —<id>.txt: the full prompt and the generated text response, including the <X.X seconds> time markers the thinker emits (the markers are stripped before the talker and are not spoken).

Scores (MAPE %, wav duration vs instructed duration, lower is better)

S-QAS-REAS-CRES-SUML-QAL-REAL-CREL-SUMOverall
13.0711.9112.7817.0811.2211.6511.3210.2712.18

MAE 3.91 s over 2,000 samples (1,992 with time markers). For reference, the paper's TiCo (Qwen2.5-Omni-7B) scores 16.19 overall and vanilla Qwen3-Omni-30B scores 42.06.

Generated 2026-09-14 with vllm-omni (thinker/talker batch 8, code2wav 1, talker max_tokens 1024).