CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stepfun-ai /Step-3.5-Flash-SFT Step-3.5-Flash-SFT Step-3.5-Flash-SFT is a general-domain supervised fine-tuning release for chat models. This repository keeps the full training interface in one place: json/: canonical raw training data tokenizers/: tokenizer snapshots for Step-3.5-Flash and Qwen3, released to preserve chat-template alignment compiled/: tokenizer-specific compiled shards for StepTronOSS training Data Format Each raw shard is a JSON file whose top level is a list of examples.… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/Step-3.5-Flash-SFT.text-generation1M<n<10M347 likes7.3k downloads6mo agoHugging Face02stepfun-ai /GEdit-BenchDataset for Step1X-Edit: A Practical Framework for General Image Editing. This dataset is a new benchmark, grounded in real-world usages is developed to support more authentic and comprehensive evaluation of image editing models. Code imageimage-to-image1K<n<10K32 likes2.6k downloads1y agoHugging Face03stepfun-ai /PaCoRe-Train-8k PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning Read the Paper | GitHub Repository | Download Models | Training Data 📖 Overview We introduce PaCoRe (Parallel Coordinated Reasoning), a framework that shifts the driver of inference from sequential depth to coordinated parallel breadth, breaking the model context limitation and massively scaling test time compute: Think in Parallel: PaCoRe launches massive parallel exploration… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/PaCoRe-Train-8k.texttext-generation1K<n<10K80 likes724 downloads8mo agoHugging Face04stepfun-ai /StepEval-Audio-Paralinguistic StepEval-Audio-Paralinguistic Dataset Paper: Step-Audio 2 Technical ReportCode: https://github.com/stepfun-ai/Step-Audio2Project Page: https://www.stepfun.com/docs/en/step-audio2 Overview StepEval-Audio-Paralinguistic is a speech-to-speech benchmark designed to evaluate AI models' understanding of paralinguistic information in speech across 11 distinct dimensions. The dataset contains 550 carefully curated and annotated speech samples for assessing capabilities beyond… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-Paralinguistic.audion<1K12 likes282 downloads1y agoHugging Face05stepfun-ai /StepEval-Audio-Toolcall StepEval-Audio-Toolcall Paper: Step-Audio 2 Technical ReportCode: https://github.com/stepfun-ai/Step-Audio2Project Page: https://www.stepfun.com/docs/en/step-audio2 Dataset Description StepEval Audio Toolcall evaluates the invocation performance of four tool types. For each tool, the benchmark contains approximately 200 multi-turn dialogue sets for both positive and negative scenarios: Positive samples: The assistant is required to invoke the specified tool in the… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-Toolcall.audio-text-to-text7 likes281 downloads1y agoHugging Face06stepfun-ai /Step-Video-T2V-EvalThis dataset contains the data of the paper Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model. Code: https://github.com/stepfun-ai/Step-Video-T2V Project page: https://yuewen.cn/videos videotext-to-videon<1K1 likes278 downloads2y agoHugging Face07stepfun-ai /GEBench GEBench: Comprehensive Benchmark for Evaluating Dynamic Interaction and Temporal Coherence in GUI Generation Overview Recent advancements in image generation models enable the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general domain visual fidelity, leaving evaluation of state transitions and temporal coherence in GUI-specific contexts underexplored. To address this gap… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/GEBench.text-to-imagen<1K12 likes162 downloads7mo agoHugging Face08stepfun-ai /Step1X-3D-obj-dataThhis is the training subdataset for Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets 10 likes148 downloads1y agoHugging Face09stepfun-ai /StepFun-Formalizer-Training StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion       Introduction This repository includes the Stage-1 SFT and RL training data of StepFun-Formalizer: NuminaMath-Formal-SFT-183K: Informal-formal problem pairs translated from NuminaMath-1.5 by Kimina-Autoformalizer-7B, used to supplement the model’s domain knowledge in formal language. StepFun-Formalizer-RL-6K: We collected informal math… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepFun-Formalizer-Training.texttext-generation100K<n<1M5 likes138 downloads11mo agoHugging Face10stepfun-ai /AndroidDaily AndroidDaily Dataset This repository hosts the AndroidDaily dataset, a benchmark grounded in real-world mobile usage patterns, introduced in the paper Step-GUI Technical Report. The AndroidDaily benchmark comprises 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios. It is specifically designed to assess whether GUI agents can handle authentic everyday usage, providing a robust evaluation for GUI automation capabilities. Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/AndroidDaily.tabularimage-text-to-textn<1K17 likes138 downloads9mo agoHugging Face11stepfun-ai /NextStep-data8 likes117 downloads7mo agoHugging Face12stepfun-ai /CF-Div2-Stepfun CF-Div2-Stepfun Evaluation Benchmark Offline benchmark of 53 Div.2 CodeForces problems. Introduction We introduce CF-Div2-Stepfun, a dataset curated to benchmark the competitive programming capabilities of Large Language Models (LLMs). We evaluate our proprietary Step 3.5 Flash alongside several frontier models on this benchmark. The benchmark comprises 53 problems sourced from official CodeForces Division 2 contests held between September 2024 and February 2025. We… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/CF-Div2-Stepfun.tabulartext-generationn<1K7 likes92 downloads7mo agoHugging Face13stepfun-ai /StepEval-Audio-360 StepEval-Audio-360 Dataset Description StepEval Audio 360 is a comprehensive dataset that evaluates the ability of multi-modal large language models (MLLMs) in human-AI audio interaction. This audio benchmark dataset, sourced from professional human annotators, covers a full spectrum of capabilities: singing, creativity, role-playing, logical reasoning, voice understanding, voice instruction following, gaming, speech emotion control, and language ability.… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-360.textn<1K29 likes90 downloads2y agoHugging Face14PGCodeLLM /stepfun0 likes41 downloads4mo agoHugging Face15miugod /stepfun3.5-flash-zh22w.jsonltext100K<n<1M0 likes18 downloads4mo agoHugging Face16rl-rag /drtulu_v2_stepfun_scientific_knowledge_0415tabular1K<n<10K0 likes17 downloads5mo agoHugging Face17vietnhat /mapalo-stepfun-metadata-real-2audion<1K0 likes16 downloads10mo agoHugging Face18miugod /stepfun3.5-flash-zh1w.jsonltext10K<n<100K0 likes14 downloads4mo agoHugging Face19vietnhat /mapalo-stepfun-metadata2audion<1K0 likes13 downloads10mo agoHugging Face20vietnhat /dee-stepfun-metadata-real-2audion<1K0 likes11 downloads10mo agoHugging Face21vietnhat /trevor-stepfun-metadata2-real-v1audion<1K0 likes10 downloads10mo agoHugging Face22vietnhat /mapalo-stepfun-metadata-real-1audion<1K0 likes8 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.