datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LIMO_QFFT
📘 LIMO–QFFT
LIMO–QFFT is a question-free variant of the original GAIR/LIMO dataset, tailored for use in QFFT (Question-Free Fine-Tuning) pipelines.
🔍 Description
This dataset removes the original input questions and system prompts from the LIMO dataset, and keeps only the long-form reasoning responses. The goal is to enable training large language models to learn from reasoning traces alone, without depending on task-specific questions.
All entries are converted into… See the full description on the dataset page: https://huggingface.co/datasets/lwl-uestc/LIMO_QFFT.S1_QFFT
📘 S1–QFFT
S1–QFFT is a question-free version of the original simplescaling/s1K-1.1 dataset, designed for QFFT training workflows.
🔍 Description
This dataset discards the original questions and any system instructions, keeping only the reasoning completions as supervision. It is especially useful for models that aim to learn when and how to think, rather than just how to answer.
The dataset is fully converted into a format compatible with LLaMA-Factory training.… See the full description on the dataset page: https://huggingface.co/datasets/lwl-uestc/S1_QFFT.
