Jnx03/kanitakorn-deepseek-v40-option-permutation-micro
Kanitakorn DeepSeek v40 Option Permutation Micro Option-order robustness continuation data for the Kanitakorn <=14B campaign. Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B Intended parent: best v39 checkpoint, not the raw base Model name taught in identity rows: kanitakorn / คณิตกรณ์ Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ Size: 956 rows = 458 original MCQ anchors + 458 option permutations 40 identity anchors Remote audit: all 458… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v40-option-permutation-micro.
Kanitakorn DeepSeek v40 Option Permutation Micro
Option-order robustness continuation data for the Kanitakorn <=14B campaign.
- Target base:
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B - Intended parent: best v39 checkpoint, not the raw base
- Model name taught in identity rows:
kanitakorn/คณิตกรณ์ - Developer taught in identity rows:
Chawabhon Netisingha/ชวภณ เนตสิงหะ - Size:
956rows =458original MCQ anchors +458option permutations 40identity anchors- Remote audit: all
458permutations changed the answer label, valid final-answer format, no missing(a)-(e)option labels, no mojibake markers - Constraints: single-model training data; no BoN, self-consistency, routing, ensemble, Thai-family base, or benchmark prompt/gold leakage
The goal is to reduce answer-position bias by teaching that the answer follows the option content after deterministic option reordering, not the old letter.
