CoolFace
Datasetpublic

cs552-the-expendables/mts-rl-training-data

Source coverage MTS-Dialog contains 1,701 source dialogues. Fact generation succeeded for 1,700; one training dialogue was excluded after no valid fact record could be generated. Fact records are available for all 200 dialogues in test1. This fact-generation exclusion is separate from the source-dialogue eligibility rule used by the current G-Eval comparison. PatientAgent MTS-Dialog training data Facts and preference-training data used by the current PatientAgent… See the full description on the dataset page: https://huggingface.co/datasets/cs552-the-expendables/mts-rl-training-data.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes67downloads
Dataset Card

Source coverage

MTS-Dialog contains 1,701 source dialogues. Fact generation succeeded for 1,700; one training dialogue was excluded after no valid fact record could be generated. Fact records are available for all 200 dialogues in test1. This fact-generation exclusion is separate from the source-dialogue eligibility rule used by the current G-Eval comparison.

PatientAgent MTS-Dialog training data

Facts and preference-training data used by the current PatientAgent SFT/DPO/KTO pipeline.

Current reproducible inputs

ConfigContentPublished splits
factsExtracted dialogue factstrain, validation, test1, test2
dpo_sft16DPO pairs generated from the rank-16 SFT modeltrain
kto_sft16KTO examples generated from the rank-16 SFT modeltrain
dpo_sft128DPO pairs generated from the rank-128 SFT modeltrain
kto_sft128KTO examples generated from the rank-128 SFT modeltrain

Only the training split was generated for the current preference datasets. The training entry points deterministically create held-out validation subsets with seed 42 when no validation file exists. The original MTS-Dialog test splits are used only for final generation/evaluation, not preference optimization.

The canonical eligibility decisions from fact extraction are stored at facts/approved_dialogues.jsonl.

Archived inputs

The archive/ directory contains earlier KTO/DPO data variants retained for provenance. They are not the current four-model comparison inputs.

Related artifacts

  • Current adapters: patientagent-sft-glm5, patientagent-sft-glm5-r128, patientagent-kto-sft16, patientagent-kto-sft128, patientagent-dpo-sft16, and patientagent-dpo-sft128.
  • Current complete G-Eval artifacts: `patientagent-eval-results`.
  • Historical 183-dialogue evaluation snapshot: `MTS-Dataset-preprocessed`.

Intended use

Research only. This dataset is for training and evaluating simulated patient dialogue. It is not medical advice, a medical device, or evidence of clinical safety.