datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reason-RFT-CoT-Dataset
🤗 Reason-RFT CoT Dateset
The full dataset used in our project "Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning".
⭐️ Project │ 🌎 Github │ 🔥 Models │ 📑 ArXiv │ 💬 WeChat
🤖 RoboBrain: Aim to Explore ReasonRFT Paradigm to Enhance RoboBrain's Embodied Reasoning Capabilities.
♣️ Quick Start
Please refer to Dataset Preparation
🔥 Overview
Visual reasoning abilities play a crucial role in understanding complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/tanhuajie2001/Reason-RFT-CoT-Dataset.Embodied-R1.5-RFT-Dataset
Embodied-R1.5-RFT-Dataset
🌐 Project Page |
📄 arXiv |
💻 Code |
🧰 EmbodiedEvalKit |
🤗 Models & Datasets
🗓️ Update — 2026-08-20 (20260820). All 28 Stage 2 RFT JSON annotation files have been uploaded to rft_datasets_json/. The complete JSON ↔ media archive mapping is documented in the Dataset composition table below.
⚠️ Partial release. This repository currently contains only a subset of the full Stage 2 RFT data… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-RFT-Dataset.3D-RFT-Reasoning142_sft_rft_dpo_simpo_v2
webshopv_sft300_3hist_rft_dpo_v2_simpo2.0
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/LLaMA-Factory/models/qwen2_vl_lora_sft_rft_v2 on the vl_dpo_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_dpo_simpo_v2.142_sft_rft_v2
sft_rft_v2
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/models/qwen2_vl_lora_sft_webshopv_300 on the vl_finetune_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 0.0001… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_v2.juno-landmark-rft142_sft_rft
sft_rft
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/models/qwen2_vl_lora_sft_webshopv_300 on the vl_finetune_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 0.0001… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft.x2x_rft_22k142_sft_rft_dpo_v2
webshopv_sft300_3hist_rft_dpo_v2
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/LLaMA-Factory/models/qwen2_vl_lora_sft_rft_v2 on the vl_dpo_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_dpo_v2.x2x_rft_16kdpv-rft-v3.1dpv-rft-v3.2dpv-rft-v3.x-potentialsdpv-rft-v1dpv-rft-v3image-rft-subsetsdpv-rft-v2dpv-rft-v3-devimage-rftask-model-anything-rft
