JingweiNi/ocr2_hardest10k_k2_lowmed_gpt55pos_qwen35neg_seed20260529
OCR2 hardest10k K2 low/medium aggregated labels This dataset follows the row-level aggregated label format of JingweiNi/ocr2_cf1900_k2_qwen35_gpt55_aggregated_10k_seed20260513, with extra K2 generation metadata. Each row is one generated K2-Think trace from the 50/50 low+medium mix over the hardest OCR2 questions. Prompt Columns question: the exact K2 completion prefix used for generation, rendered from raw_question according to k2_reasoning_effort. raw_question:… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_hardest10k_k2_lowmed_gpt55pos_qwen35neg_seed20260529.
Document K2-prefixed question column
Format question column as K2 completion prefix
Document k2_reasoning_effort column
Add k2_reasoning_effort column
Update dataset card and summary for fixed dataset
Push fixed row-level aggregated dataset
Add row-level dataset card and summary; remove stale step table
Overwrite with row-level aggregated GPT-5.5/Qwen3.5 labels
Upload README.md with huggingface_hub
Upload labels.csv with huggingface_hub
Upload summary.json with huggingface_hub
Upload dataset
initial commit
