DaoCloud/Muse-Glimmer-OPB-100K
Muse Glimmer OPB 100K On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark. Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B. The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths. Reasoning strength Conversations Train-turn rows low 64,997 96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.
Muse Glimmer OPB 100K
On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark.
Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B.
The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths.
Schema
Every row has trainable tokens. Row IDs are unique, source conversation sets are disjoint, and generation ordinals are contiguous.
File
muse_opb_onpolicy_100k.jsonl
1,646,393,887 bytes
SHA-256 70d68dff971d1fe00d90a356b1a1eaeac845f9a3cc2069ce525043312444439aThis release contains tokenized training examples. Use the Muse Glimmer 30B tokenizer when inspecting or converting input_ids back to text.
