datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mulberry-SFTPlease check our GitHub for more details.: https://github.com/HJYao00/Mulberry
Training
We use LLaMA-Factory to fine-tune the Mulberry models. We provide the training instructions and configs here.
First, install LLaMA-Factory according to the official_instruction.
Then, refer here and update the following customized dataset into dataset_info.json in LLaMA-Factory.
"mulberry": {
"file_name": "./mulberry_sft.json",
"formatting": "sharegpt",
"columns": {
"messages":… See the full description on the dataset page: https://huggingface.co/datasets/HuanjinYao/Mulberry-SFT.UnifiedReward-Flex-SFT-90K
UnifiedReward-Flex-SFT-90K
This repository releases 90K SFT data of UnifiedReward-Flex.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/abs/2602.02380
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/flex
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-flex
🤗 Dataset: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-Flex-SFT-90K
👋 Point of Contact: Yibin Wang
Citation… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-Flex-SFT-90K.Mobile-O-SFT
Mobile-O SFT Data
Supervised Fine-Tuning · ~105K Curated Prompt-Image Pairs
📌 Overview
This dataset is used for Stage 2: Supervised Fine-Tuning (SFT) of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to improve image generation quality by fine-tuning on high-quality curated prompt-image pairs.
📊 Dataset Composition
Source
Samples
Description
BLIP3o
60K
High-quality prompt-image pairs… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-SFT.Mobile-GUI-Worldmodel-SFT
Mobile-GUI-Worldmodel-SFT
This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text.
Repository Layout
.
├── GUI-agent-main/ # Data annotation scripts and examples
├── eval/ # Evaluation assets
│ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.ABot-PhysWorld_SFT_Training_Data_v1VLM-SFTCapTTS-SFT-Audio
CapTTS-SFT Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapTTS-SFT-Audio.ABot-PhysWorld_SFT_Training_Data_v1_RoboMINDsft-q1GoClick_sft_datasftsft_ruleUIPro_Mobile_SFTDataNL-To1_sft_datasetsnewsm-sft-images-tmpEmoAct-SFT-Data
