datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B.
The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning.
We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details.
If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.RAQUEL-MUSE-News-Paraphrase
RAQUEL MUSE-News knowmem paraphrases
One reworded version of each question in the MUSE-News knowledge-memorization (knowmem) QA sets: 100 forget and
100 retain questions. The reference answer is unchanged, so a paraphrase is scored against the same answer as its
original. MUSE-News ships no paraphrased questions; this set fills that gap for the RAQUEL unlearning evaluation,
mirroring the paraphrased_question field that TOFU releases for its forget and retain sets.
Split… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/RAQUEL-MUSE-News-Paraphrase.
