datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Step-3.5-Flash-SFT
Step-3.5-Flash-SFT
Step-3.5-Flash-SFT is a general-domain supervised fine-tuning release for chat models.
This repository keeps the full training interface in one place:
json/: canonical raw training data
tokenizers/: tokenizer snapshots for Step-3.5-Flash and Qwen3, released to preserve chat-template alignment
compiled/: tokenizer-specific compiled shards for StepTronOSS training
Data Format
Each raw shard is a JSON file whose top level is a list of examples.… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/Step-3.5-Flash-SFT.GEdit-BenchDataset for Step1X-Edit: A Practical Framework for General Image Editing.
This dataset is a new benchmark, grounded in real-world usages is developed to support more authentic and comprehensive evaluation of image editing models.
Code
PaCoRe-Train-8k
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
Read the Paper | GitHub Repository | Download Models | Training Data
📖 Overview
We introduce PaCoRe (Parallel Coordinated Reasoning), a framework that shifts the driver of inference from sequential depth to coordinated parallel breadth, breaking the model context limitation and massively scaling test time compute:
Think in Parallel: PaCoRe launches massive parallel exploration… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/PaCoRe-Train-8k.StepEval-Audio-Paralinguistic
StepEval-Audio-Paralinguistic Dataset
Paper: Step-Audio 2 Technical ReportCode: https://github.com/stepfun-ai/Step-Audio2Project Page: https://www.stepfun.com/docs/en/step-audio2
Overview
StepEval-Audio-Paralinguistic is a speech-to-speech benchmark designed to evaluate AI models' understanding of paralinguistic information in speech across 11 distinct dimensions. The dataset contains 550 carefully curated and annotated speech samples for assessing capabilities beyond… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-Paralinguistic.StepEval-Audio-Toolcall
StepEval-Audio-Toolcall
Paper: Step-Audio 2 Technical ReportCode: https://github.com/stepfun-ai/Step-Audio2Project Page: https://www.stepfun.com/docs/en/step-audio2
Dataset Description
StepEval Audio Toolcall evaluates the invocation performance of four tool types. For each tool, the benchmark contains approximately 200 multi-turn dialogue sets for both positive and negative scenarios:
Positive samples: The assistant is required to invoke the specified tool in the… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-Toolcall.Step-Video-T2V-EvalThis dataset contains the data of the paper Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.
Code: https://github.com/stepfun-ai/Step-Video-T2V
Project page: https://yuewen.cn/videos
GEBench
GEBench: Comprehensive Benchmark for Evaluating Dynamic Interaction and Temporal Coherence in GUI Generation
Overview
Recent advancements in image generation models enable the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general domain visual fidelity, leaving evaluation of state transitions and temporal coherence in GUI-specific contexts underexplored.
To address this gap… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/GEBench.Step1X-3D-obj-dataThhis is the training subdataset for Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
StepFun-Formalizer-Training
StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
Introduction
This repository includes the Stage-1 SFT and RL training data of StepFun-Formalizer:
NuminaMath-Formal-SFT-183K: Informal-formal problem pairs translated from NuminaMath-1.5 by Kimina-Autoformalizer-7B, used to supplement the model’s domain knowledge in formal language.
StepFun-Formalizer-RL-6K: We collected informal math… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepFun-Formalizer-Training.AndroidDaily
AndroidDaily Dataset
This repository hosts the AndroidDaily dataset, a benchmark grounded in real-world mobile usage patterns, introduced in the paper Step-GUI Technical Report.
The AndroidDaily benchmark comprises 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios. It is specifically designed to assess whether GUI agents can handle authentic everyday usage, providing a robust evaluation for GUI automation capabilities.
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/AndroidDaily.NextStep-dataCF-Div2-Stepfun
CF-Div2-Stepfun Evaluation Benchmark
Offline benchmark of 53 Div.2 CodeForces problems.
Introduction
We introduce CF-Div2-Stepfun, a dataset curated to benchmark the competitive programming capabilities of Large Language Models (LLMs). We evaluate our proprietary Step 3.5 Flash alongside several frontier models on this benchmark.
The benchmark comprises 53 problems sourced from official CodeForces Division 2 contests held between September 2024 and February 2025. We… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/CF-Div2-Stepfun.StepEval-Audio-360
StepEval-Audio-360
Dataset Description
StepEval Audio 360 is a comprehensive dataset that evaluates the ability of multi-modal large language models (MLLMs) in human-AI audio interaction. This audio benchmark dataset, sourced from professional human annotators, covers a full spectrum of capabilities: singing, creativity, role-playing, logical reasoning, voice understanding, voice instruction following, gaming, speech emotion control, and language ability.… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-360.stepfunstepfun3.5-flash-zh22w.jsonldrtulu_v2_stepfun_scientific_knowledge_0415mapalo-stepfun-metadata-real-2stepfun3.5-flash-zh1w.jsonlmapalo-stepfun-metadata2dee-stepfun-metadata-real-2trevor-stepfun-metadata2-real-v1mapalo-stepfun-metadata-real-1
