datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AndroidDaily
AndroidDaily Dataset
This repository hosts the AndroidDaily dataset, a benchmark grounded in real-world mobile usage patterns, introduced in the paper Step-GUI Technical Report.
The AndroidDaily benchmark comprises 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios. It is specifically designed to assess whether GUI agents can handle authentic everyday usage, providing a robust evaluation for GUI automation capabilities.
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/AndroidDaily.CF-Div2-Stepfun
CF-Div2-Stepfun Evaluation Benchmark
Offline benchmark of 53 Div.2 CodeForces problems.
Introduction
We introduce CF-Div2-Stepfun, a dataset curated to benchmark the competitive programming capabilities of Large Language Models (LLMs). We evaluate our proprietary Step 3.5 Flash alongside several frontier models on this benchmark.
The benchmark comprises 53 problems sourced from official CodeForces Division 2 contests held between September 2024 and February 2025. We… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/CF-Div2-Stepfun.drtulu_v2_stepfun_scientific_knowledge_0415
