datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama31_8b_math_new_prompt_filtered_no_self_correction_sftupdated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v2updated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v1llama3_8b_math_new_prompt_filtered_no_self_correction_sftupdated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v2filteredself-consistency-correction-ambigqa
self-consistency-correction — AmbigQA
Run of the 5-score self-consistency-correction comparison on 10 hand-clustered AmbigQA questions from the rankalign longform branch data/ambigqa/with_negatives/. Paper: Correcting Generator Scores via Self-Consistency.
Setup
Each AmbigQA question has multiple valid answer entities (ambiguous interpretations) and rankalign-generated distractors (strategies: plausible, clearly-wrong, uninformed). We hand-clustered surface forms from the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/self-consistency-correction-ambigqa.qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v1updated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v1filteredqwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v2updated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v2self-consistency-correction-exp1-assumptions
self-consistency-correction — Experiment 1 (assumption checks)
Paper §7.2. Tests Assumptions 2 and 3 of the self-consistency framework using the same 20 hand-crafted cases as Exps 2/3. No new inference — this is an analysis pass over the exp23 artifact.
Assumption 2 — semantic validation
The validator should judge meaning, not surface form: V'(X, y) ≈ V'(X, y') whenever y ∼ y'. Test: within each case's correct equivalence class, compute the absolute difference |V'(y) −… See the full description on the dataset page: https://huggingface.co/datasets/latkes/self-consistency-correction-exp1-assumptions.updated_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212__self_correction_iter1_v1self-consistency-correction-exp23-correlation-discrimination
self-consistency-correction — Experiments 2 + 3 (combined)
Paper: Correcting Generator Scores via Self-Consistency §7.3 (correlation) and §7.4 (discrimination).
20 hand-crafted cases with known equivalence classes covering geography, literature, science, history, and pop culture. Each case has a correct equivalence class containing 2-6 surface-form paraphrases plus 2 singleton wrong-answer classes (mix of rare and common string frequencies). Equivalence classes are hardcoded, so… See the full description on the dataset page: https://huggingface.co/datasets/latkes/self-consistency-correction-exp23-correlation-discrimination.self-consistency-correction-exp0-sanity-canary
self-consistency-correction — Experiment 0 (sanity check)
Paper: Correcting Generator Scores via Self-Consistency: Theory, Practical Approximations, and Experiments.
§7.1 no-paraphrase baseline. 8 hand-crafted cases (L1-L4 low paraphrase richness, H1-H4 high paraphrase richness). For each case, all candidate answers are scored with:
G_raw = log P(y | X) via teacher-forced scoring (longest-common-prefix boundary handling to avoid whitespace retokenization bugs)
V' = log P(Yes |… See the full description on the dataset page: https://huggingface.co/datasets/latkes/self-consistency-correction-exp0-sanity-canary.self-consistency-correction-exp5-diagnostic
self-consistency-correction — Experiment 5 (Liechtenstein vs UK diagnostic)
Paper §7.6. Hand-crafted cases where PMI and cluster correction predictably disagree.
For each of 6 questions, candidates include multiple surface forms of a "rich" equivalence class (correct meaning with many paraphrases) plus singleton wrong answers with rare and common string frequencies. Equivalence classes are hardcoded (not estimated), so G_cluster_exact = log sum_{y' in [y]} exp(G(X, y')) is the exact… See the full description on the dataset page: https://huggingface.co/datasets/latkes/self-consistency-correction-exp5-diagnostic.
