JingweiNi/ocr2_hardest10k_k2_lowmed_gpt55pos_qwen35neg_seed20260529
OCR2 hardest10k K2 low/medium aggregated labels This dataset follows the row-level aggregated label format of JingweiNi/ocr2_cf1900_k2_qwen35_gpt55_aggregated_10k_seed20260513, with extra K2 generation metadata. Each row is one generated K2-Think trace from the 50/50 low+medium mix over the hardest OCR2 questions. Prompt Columns question: the exact K2 completion prefix used for generation, rendered from raw_question according to k2_reasoning_effort. raw_question:… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_hardest10k_k2_lowmed_gpt55pos_qwen35neg_seed20260529.
OCR2 hardest10k K2 low/medium aggregated labels
This dataset follows the row-level aggregated label format of `JingweiNi/ocr2_cf1900_k2_qwen35_gpt55_aggregated_10k_seed20260513`, with extra K2 generation metadata.
Each row is one generated K2-Think trace from the 50/50 low+medium mix over the hardest OCR2 questions.
Prompt Columns
question: the exact K2 completion prefix used for generation, rendered fromraw_questionaccording tok2_reasoning_effort.raw_question: the original unformatted OCR2 problem statement.k2_reasoning_effort: K2 generation reasoning mode, eitherlowormedium.
The K2 prompt templates are:
low:configs/k2_think_low_completion_prefix.txt, ending in<think_faster>.medium:configs/k2_think_medium_completion_prefix.txt, ending in<think_fast>.
For every row, tokenizing question with the K2 tokenizer matches the prefix of input_ids.
Label Columns
Step labels are stored as aligned arrays:
qwen35_verified: original Qwen3.5-122B-A10B-FP8 step labels.gpt55_verified: GPT-5.5 medium re-annotations for completed Qwen3.5-positive steps, withNaNelsewhere.verified: final aggregated labels. Qwen3.5 negatives are0; GPT-5.5-confirmed positives are1; GPT-5.5-filtered Qwen positives are0; unreannotated Qwen positives remainNaN.
Aggregation rule: GPT-5.5 confirmed errors are errors; all Qwen3.5 negatives and GPT-5.5 rejected Qwen positives are correct.
Summary
- Rows:
19936 - K2 reasoning effort counts:
low=9968,medium=9968 - Qwen3.5 positive steps:
5110 - Qwen3.5 negative steps:
14867 - Valid GPT-5.5 reannotations in this mixed dataset:
1191 - GPT-5.5 confirmed positive steps:
763 - Final verified negative steps:
15295 - Final
NaNsteps:4349885
Validation
The uploaded dataset was checked against the source mixed dataset, K2 tokenizer, and GPT-5.5 annotation records:
raw_questionis the original source problem statement.questionis the exact low/medium K2 completion prefix, and its token IDs match the prefix ofinput_idsfor every row.k2_reasoning_effortis copied row-wise from the mixed source dataset and has9968low rows and9968medium rows.qwen35_verifiedis an exact copy of the source Qwen3.5 labels.- GPT-5.5 labels are aligned by stable row/step identity, with
1191valid GPT positions in this mixed dataset. claims,qwen35_verified,gpt55_verified, andverifiedlengths match for every row.- All checked
claims[*].aligned_token_idsare within the row'sinput_idsrange.
solution and qwq_critique are empty strings because these columns are required by the reference schema but absent from the mixed source dataset.
See summary.json for source paths and tuple counts.
