HoangHa/meddies-title
Meddies Title SFT title_sft_v0 trains multilingual session-title generation. Each row stores a system instruction, the user query, and one title response in messages. The raw query pools remain in this public canonical repository for provenance, but are omitted from dataset viewer configurations. SFT generation progress Status: partial Mode: full Generated accepted rows: 52901 / 100000 primary + 23941 auxiliary Published rows: 52901 primary + 23941 auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/HoangHa/meddies-title.
Meddies Title SFT
title_sft_v0 trains multilingual session-title generation. Each row stores a system instruction, the user query, and one title response in messages.
The raw query pools remain in this public canonical repository for provenance, but are omitted from dataset viewer configurations.
SFT generation progress
Status: partial
Mode: full
Generated accepted rows: 52901 / 100000 primary + 23941 auxiliary
Published rows: 52901 primary + 23941 auxiliary
Excluded warnings: 40137
Excluded hard rejects: 6962
Prompt revision: title-only-v8.3-sidebar-fit
Curated training set
title_curated_train_v0 holds 53,804 judge-verified (query, title) pairs in SFT chat format (messages: system, user query, assistant title). Winners picked by rubric judges (curator-ranked), extra kept near-duplicates (curator-kept), and judge-written repairs of rejected candidates (curator-repair). Judging prompts: title-curator-rubric-v3 and title-curator-rubric-round2-v2. Ready to fine-tune on as-is.
title_curated_train_v1 merges every judged verdict (SFT, prune, QAT, repairs) into one 86,709-row SFT set with a minimal training prompt (34 words): write the title, name the distinguishing fact, use only stated facts. Ready to fine-tune on as-is.
title_curated_train_v2 rebuilds the same recipe from all current verdicts: 105,516 rows (96,147 ranked winners, 8,183 repairs, 1,186 kept near-duplicates), a strict superset of v1 with zero dropped rows. Row IDs carry the verdict suffix (.ranked./.kept./.repair. + title hash). Same 34-word prompt and schema as v1.
The viewer above lists training configs only (curated SFT + DPO). Raw generation pools (title_sft_v0, title_pruning_calibration_v0, title_qat_v0) and candidate query folders stay in the repository for provenance but are out of the viewer.
DPO preference pairs
title_dpo_v1 holds 3,824 (prompt, chosen, rejected) triples for preference training. Each row pairs one judge-rejected title (rejected) with its judge-written rewrite (chosen) from the round-2 repair prompt, restricted to singleton pools (one candidate, one repair) so the pairing is unambiguous. Three defect classes are removed: echo pairs (rewrite identical to the rejected title), gate-failing rewrites, and lane boilerplate (one provider lane emitted the same canned text across hundreds of unrelated queries — every repeated-across-queries rewrite is dropped). Multi-candidate repair records are excluded: measured alignment shows no positional or similarity structure to attribute a rewrite to its source candidate, so they stay out until the judge records alignment. Rows carry the same source/first_lineage/repair_lineage provenance as title_dpo_repair_v0, which it supersedes in coverage but does not replace. Judge-rewrite pairs distill judge style; on-policy (model output vs repair) pairs remain future work.
