datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-bench-Science
SWE-bench Science
SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers.
GitHub release repository: OpenMOSS/SWE-bench-Science
Runtime images: Docker Hub, pinned by immutable linux/amd64 digests
Evaluation framework: Pier, compatible with Harbor task format
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.VehicleWorld
📚 Introduction
VehicleWorld is the first comprehensive multi-device environment for intelligent vehicle interaction that accurately models the complex, interconnected systems in modern cockpits. This environment enables precise evaluation of agent behaviors by providing real-time state information during execution. This dataset is specifically designed to evaluate the capabilities of Large Language Models (LLMs) as in-car intelligent assistants in understanding and executing… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/VehicleWorld.
