datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reason-RFT-CoT-Dataset
🤗 Reason-RFT CoT Dateset
The full dataset used in our project "Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning".
⭐️ Project │ 🌎 Github │ 🔥 Models │ 📑 ArXiv │ 💬 WeChat
🤖 RoboBrain: Aim to Explore ReasonRFT Paradigm to Enhance RoboBrain's Embodied Reasoning Capabilities.
♣️ Quick Start
Please refer to Dataset Preparation
🔥 Overview
Visual reasoning abilities play a crucial role in understanding complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/tanhuajie2001/Reason-RFT-CoT-Dataset.wildchat-creative-writing-3k-rft142_sft_rft_dpo_simpo_v2
webshopv_sft300_3hist_rft_dpo_v2_simpo2.0
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/LLaMA-Factory/models/qwen2_vl_lora_sft_rft_v2 on the vl_dpo_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_dpo_simpo_v2.R2E-Gym-Lite-RFT-no-thinkqwen3-8b-dpo-agentgym-rft-data4b_rft_response-2-custom_student_responseGraph-R1-RFT-COT-30K
Dataset Card: Graph-CoT-30k
Dataset Details
Dataset Name: Graph-CoT-30k
Dataset Creator: HKUST-DSAIL
Dataset Version: 1.0
Release Date: August 2025
Description
Graph-CoT-30k is a large-scale, high-quality instruction tuning dataset designed to enhance the reasoning capabilities of large language models (LLMs) on complex graph-theoretic problems. It contains 30,000 question-answer (QA) pairs, each featuring ultra-long chain-of-thought (CoT) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-RFT-COT-30K.R2E-Gym-Lite-RFT142_sft_rft_v2
sft_rft_v2
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/models/qwen2_vl_lora_sft_webshopv_300 on the vl_finetune_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 0.0001… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_v2.juno-landmark-rft142_sft_rft
sft_rft
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/models/qwen2_vl_lora_sft_webshopv_300 on the vl_finetune_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 0.0001… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft.x2x_rft_22k4b_rft_response-7-custom_student_response-verified4b_rft_response-5-custom_student_response-verified4b_rft_response-3-custom_student_response-verified4b_rft_response-4-custom_student_response-verified4b_rft_response-1-custom_student_response-verifiedGraphInstruct-RFT-72K142_sft_rft_dpo_v2
webshopv_sft300_3hist_rft_dpo_v2
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/LLaMA-Factory/models/qwen2_vl_lora_sft_rft_v2 on the vl_dpo_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_dpo_v2.rft-finetune-llama-3.2-1b-mathx2x_rft_16kdpv-rft-v3.1dpv-rft-v3.24b_rft_response-4-custom_student_response4b_rft_response-2-custom_student_response-verifieddpv-rft-v3.x-potentialsrft-finetune-llama-3.2-1b-math-k10kk-onpolicy-rftkk-rft-train4b_rft_response-2-student_logps
