4-step
Datasets
All datasets matching “4-step”step35-en2pl-conv-pass4-jsonlconversations: 1,251,034
chat-template tokens (role+content, incl. special tokens): 2,664,206,408
reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429
avg tokens/conversation: 2129.6
used tokenizer: APT4
ssh4-step3000
Model Card for Model ID
Model Details
Model Description
Developed by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Model type: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/ssh4-step3000.audio-quality-dataset-nfe4-30-step2
Audio Quality Dataset: NFE 4-30 Step 2
Overview
This dataset publishes synthetic speech artifacts and derived spectrograms used for repo-local audio-quality experiments.
At a glance:
2800 synthetic runs
200 short English prompt sentences
14 NFE settings: 4, 6, 8, ..., 30
fixed seed 1024
Each row represents one synthetic run and includes:
prompt text
raw synthetic WAV
processed synthetic WAV
spectrogram PNG
NFE value
procedural weak label
Here, NFE means the number of… See the full description on the dataset page: https://huggingface.co/datasets/TashaSkyUp/audio-quality-dataset-nfe4-30-step2.qwen3_4b_instruct_lcbv6_rsa_pop_32_k_4_steps_10_s_65_e_1310_25_rsa_pop_32_k_4_steps_10_v225_50_rsa_pop_32_k_4_steps_10_v2
