dspark
Datasets
All datasets matching “dspark”EVE-Instruct-Dspark-training-data
EVE-Instruct D-Spark Training Data
Prepared speculative-decoding training data for eve-esa/EVE-Instruct.
Source: eve-esa/synth
Source splits: qa and long_qa only
Rows: 614,960
Format: EVE system prompt + question (input) + answer (output)
Source context and file_path fields omitted
Maximum sequence length: 4096 tokens
Minimum trainable assistant tokens: 16
Columns: input_ids, loss_mask, seq_len
loss_mask is 1 only for assistant response tokens (including EOS). token_freq.pt… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/EVE-Instruct-Dspark-training-data.Qwen3.8-DSpark-PerfectBlend-10M-Stride16-H32-FeaturesQwen3.8-DSpark-PerfectBlend-5M-Paired-BF16
Qwen3.8 DSpark PerfectBlend 5M paired BF16 features
Private, checksum-closed paired feature corpus for the Qwen3.8 Flash / 27B
DSpark transplant project. Repository: MJPansa/Qwen3.8-DSpark-PerfectBlend-5M-Paired-BF16.
The Hugging Face DatasetDict rows are a compact index. Each row points into
three immutable SafeTensor files in tensors/shard-NNNNN/ using exact token
and anchor offsets. This keeps the ~180 GB dense BF16 corpus resumable and
memory-mappable instead of duplicating… See the full description on the dataset page: https://huggingface.co/datasets/MJPansa/Qwen3.8-DSpark-PerfectBlend-5M-Paired-BF16.Qwen3.6-27B-DSpark-data
Qwen3.6-27B DSpark training data (on-policy, clean)
On-policy conversations generated by Avesed/Qwen3.6-27B-W4A16
on a sha256-verified checkpoint, used to train Avesed/Qwen3.6-27B-DSpark.
Each file is its own dataset config (they use different id schemes — integer vs zh_*
string — so the viewer must keep them separate rather than merge into one table).
config / file
convs
lang
prompt source
pb_pool94k_clean
~86k
en
PerfectBlend-style instruction mix
general_onpolicy… See the full description on the dataset page: https://huggingface.co/datasets/Avesed/Qwen3.6-27B-DSpark-data.dspark-e2b-heretic-datadspark-regen-input-opb-10000
