CoolFace
Datasetpublic

Jnx03/kanitakorn-deepseek-v40-option-permutation-micro

Kanitakorn DeepSeek v40 Option Permutation Micro Option-order robustness continuation data for the Kanitakorn <=14B campaign. Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B Intended parent: best v39 checkpoint, not the raw base Model name taught in identity rows: kanitakorn / คณิตกรณ์ Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ Size: 956 rows = 458 original MCQ anchors + 458 option permutations 40 identity anchors Remote audit: all 458… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v40-option-permutation-micro.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes15downloads
Dataset Card

Kanitakorn DeepSeek v40 Option Permutation Micro

Option-order robustness continuation data for the Kanitakorn <=14B campaign.

  • —Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
  • —Intended parent: best v39 checkpoint, not the raw base
  • —Model name taught in identity rows: kanitakorn / คณิตกรณ์
  • —Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
  • —Size: 956 rows = 458 original MCQ anchors + 458 option permutations
  • —40 identity anchors
  • —Remote audit: all 458 permutations changed the answer label, valid final-answer format, no missing (a)-(e) option labels, no mojibake markers
  • —Constraints: single-model training data; no BoN, self-consistency, routing, ensemble, Thai-family base, or benchmark prompt/gold leakage

The goal is to reduce answer-position bias by teaching that the answer follows the option content after deterministic option reordering, not the old letter.