datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reason-RFT-CoT-Dataset
🤗 Reason-RFT CoT Dateset
The full dataset used in our project "Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning".
⭐️ Project │ 🌎 Github │ 🔥 Models │ 📑 ArXiv │ 💬 WeChat
🤖 RoboBrain: Aim to Explore ReasonRFT Paradigm to Enhance RoboBrain's Embodied Reasoning Capabilities.
♣️ Quick Start
Please refer to Dataset Preparation
🔥 Overview
Visual reasoning abilities play a crucial role in understanding complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/tanhuajie2001/Reason-RFT-CoT-Dataset.Embodied-R1.5-RFT-Dataset
Embodied-R1.5-RFT-Dataset
🌐 Project Page |
📄 arXiv |
💻 Code |
🧰 EmbodiedEvalKit |
🤗 Models & Datasets
🗓️ Update — 2026-08-20 (20260820). All 28 Stage 2 RFT JSON annotation files have been uploaded to rft_datasets_json/. The complete JSON ↔ media archive mapping is documented in the Dataset composition table below.
⚠️ Partial release. This repository currently contains only a subset of the full Stage 2 RFT data… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-RFT-Dataset.RFT-reference-trajectorywildchat-creative-writing-3k-rft3D-RFT-Reasoning142_sft_rft_dpo_simpo_v2
webshopv_sft300_3hist_rft_dpo_v2_simpo2.0
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/LLaMA-Factory/models/qwen2_vl_lora_sft_rft_v2 on the vl_dpo_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_dpo_simpo_v2.R2E-Gym-Lite-RFT-no-thinkqwen3-8b-dpo-agentgym-rft-data4b_rft_response-2-custom_student_responserftransThis repo contains the dataset used in RFTrans, generated by the Data Generator, powered by RFUniverse.
The train and val folder contain the proposed synthetic dataset.
You can also generate your own dataset with the tools and assets provided by us.
Resources.zip contains the assets we used to generate the dataset.
To note, we do not claim to own the copyright of these assets.
The cleargrasp folder contains the example train set generated with the models from ClearGrasp.
It's the train set we… See the full description on the dataset page: https://huggingface.co/datasets/robotflow/rftrans.Graph-R1-RFT-COT-30K
Dataset Card: Graph-CoT-30k
Dataset Details
Dataset Name: Graph-CoT-30k
Dataset Creator: HKUST-DSAIL
Dataset Version: 1.0
Release Date: August 2025
Description
Graph-CoT-30k is a large-scale, high-quality instruction tuning dataset designed to enhance the reasoning capabilities of large language models (LLMs) on complex graph-theoretic problems. It contains 30,000 question-answer (QA) pairs, each featuring ultra-long chain-of-thought (CoT) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-RFT-COT-30K.R2E-Gym-Lite-RFT142_sft_rft_v2
sft_rft_v2
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/models/qwen2_vl_lora_sft_webshopv_300 on the vl_finetune_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 0.0001… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_v2.juno-landmark-rft142_sft_rft
sft_rft
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/models/qwen2_vl_lora_sft_webshopv_300 on the vl_finetune_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 0.0001… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft.x2x_rft_22kQwen-2.5-Math-1.5B-RFT-DPO-Data4b_rft_response-7-custom_student_response-verified4b_rft_response-5-custom_student_response-verified4b_rft_response-3-custom_student_response-verified4b_rft_response-4-custom_student_response-verified4b_rft_response-1-custom_student_response-verifiedGraphInstruct-RFT-72K142_sft_rft_dpo_v2
webshopv_sft300_3hist_rft_dpo_v2
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/LLaMA-Factory/models/qwen2_vl_lora_sft_rft_v2 on the vl_dpo_data dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_sft_rft_dpo_v2.rft-finetune-llama-3.2-1b-mathx2x_rft_16kdpv-rft-v3.1dpv-rft-v3.2libero-rft-spatial-full14b_rft_response-4-custom_student_response
