CoolFace
Datasetpublic

JingweiNi/ocr2_hardest10k_k2_lowmed_gpt55pos_qwen35neg_seed20260529

OCR2 hardest10k K2 low/medium aggregated labels This dataset follows the row-level aggregated label format of JingweiNi/ocr2_cf1900_k2_qwen35_gpt55_aggregated_10k_seed20260513, with extra K2 generation metadata. Each row is one generated K2-Think trace from the 50/50 low+medium mix over the hardest OCR2 questions. Prompt Columns question: the exact K2 completion prefix used for generation, rendered from raw_question according to k2_reasoning_effort. raw_question:… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_hardest10k_k2_lowmed_gpt55pos_qwen35neg_seed20260529.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
1likes15downloads
Dataset Card

OCR2 hardest10k K2 low/medium aggregated labels

This dataset follows the row-level aggregated label format of `JingweiNi/ocr2_cf1900_k2_qwen35_gpt55_aggregated_10k_seed20260513`, with extra K2 generation metadata.

Each row is one generated K2-Think trace from the 50/50 low+medium mix over the hardest OCR2 questions.

Prompt Columns

  • —question: the exact K2 completion prefix used for generation, rendered from raw_question according to k2_reasoning_effort.
  • —raw_question: the original unformatted OCR2 problem statement.
  • —k2_reasoning_effort: K2 generation reasoning mode, either low or medium.

The K2 prompt templates are:

  • —low: configs/k2_think_low_completion_prefix.txt, ending in <think_faster>.
  • —medium: configs/k2_think_medium_completion_prefix.txt, ending in <think_fast>.

For every row, tokenizing question with the K2 tokenizer matches the prefix of input_ids.

Label Columns

Step labels are stored as aligned arrays:

  • —qwen35_verified: original Qwen3.5-122B-A10B-FP8 step labels.
  • —gpt55_verified: GPT-5.5 medium re-annotations for completed Qwen3.5-positive steps, with NaN elsewhere.
  • —verified: final aggregated labels. Qwen3.5 negatives are 0; GPT-5.5-confirmed positives are 1; GPT-5.5-filtered Qwen positives are 0; unreannotated Qwen positives remain NaN.

Aggregation rule: GPT-5.5 confirmed errors are errors; all Qwen3.5 negatives and GPT-5.5 rejected Qwen positives are correct.

Summary

  • —Rows: 19936
  • —K2 reasoning effort counts: low=9968, medium=9968
  • —Qwen3.5 positive steps: 5110
  • —Qwen3.5 negative steps: 14867
  • —Valid GPT-5.5 reannotations in this mixed dataset: 1191
  • —GPT-5.5 confirmed positive steps: 763
  • —Final verified negative steps: 15295
  • —Final NaN steps: 4349885

Validation

The uploaded dataset was checked against the source mixed dataset, K2 tokenizer, and GPT-5.5 annotation records:

  • —raw_question is the original source problem statement.
  • —question is the exact low/medium K2 completion prefix, and its token IDs match the prefix of input_ids for every row.
  • —k2_reasoning_effort is copied row-wise from the mixed source dataset and has 9968 low rows and 9968 medium rows.
  • —qwen35_verified is an exact copy of the source Qwen3.5 labels.
  • —GPT-5.5 labels are aligned by stable row/step identity, with 1191 valid GPT positions in this mixed dataset.
  • —claims, qwen35_verified, gpt55_verified, and verified lengths match for every row.
  • —All checked claims[*].aligned_token_ids are within the row's input_ids range.

solution and qwq_critique are empty strings because these columns are required by the reference schema but absent from the mixed source dataset.

See summary.json for source paths and tuple counts.