datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EVE-Instruct-Dspark-training-data
EVE-Instruct D-Spark Training Data
Prepared speculative-decoding training data for eve-esa/EVE-Instruct.
Source: eve-esa/synth
Source splits: qa and long_qa only
Rows: 614,960
Format: EVE system prompt + question (input) + answer (output)
Source context and file_path fields omitted
Maximum sequence length: 4096 tokens
Minimum trainable assistant tokens: 16
Columns: input_ids, loss_mask, seq_len
loss_mask is 1 only for assistant response tokens (including EOS). token_freq.pt… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/EVE-Instruct-Dspark-training-data.Qwen3.6-27B-DSpark-data
Qwen3.6-27B DSpark training data (on-policy, clean)
On-policy conversations generated by Avesed/Qwen3.6-27B-W4A16
on a sha256-verified checkpoint, used to train Avesed/Qwen3.6-27B-DSpark.
Each file is its own dataset config (they use different id schemes — integer vs zh_*
string — so the viewer must keep them separate rather than merge into one table).
config / file
convs
lang
prompt source
pb_pool94k_clean
~86k
en
PerfectBlend-style instruction mix
general_onpolicy… See the full description on the dataset page: https://huggingface.co/datasets/Avesed/Qwen3.6-27B-DSpark-data.
