CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alibabagroup /terminal-bench-pro Terminal-Bench Pro Overview Terminal-Bench Pro is a systematic extension of the original Terminal-Bench, designed to address key limitations in existing terminal-agent benchmarks. 400 tasks (200 public + 200 private) across 8 domains: data processing, games, debugging, system admin, scientific computing, software engineering, ML, and security Expert-designed tasks derived from real-world scenarios and GitHub issues High test coverage with ~28.3 test cases per… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/terminal-bench-pro.texttext-generationn<1K5 likes3.3k downloads9mo agoHugging Face02Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob           🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid. Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M63 likes2.2k downloads8mo agoHugging Face03Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M352 likes1.4k downloads8mo agoHugging Face04Alibaba-Aone /aacr-bench Dataset for Running AACR-Bench English | 简体中文 This is a test set designed for automated code review reflection models, primarily aiming to evaluate the extent to which a model can intercept low-quality review comments. The dataset contains 2,145 code review comments, consisting of 1,505 expert-verified correct comments and 640 incorrect comments. This data is part of the AACR-Bench project and is provided by the Alibaba Aone team. Data Sample Each sample in the… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Aone/aacr-bench.tabulartext-generation1K<n<10K9 likes604 downloads8mo agoHugging Face05Alibaba-AAIG /XGuard-Train-Open-200K YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models   🤗 HuggingFace   |   🤖 ModelScope   |   📄 Paper   🐬 Introduction XGuard-Train-Open-200K is an open-source subset of the training corpus developed for the YuFeng-XGuard-Reason guardrail model series. YuFeng-XGuard-Reason is engineered to accurately identify security risks in user requests, model responses, and general text, while… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/XGuard-Train-Open-200K.texttext-classification100K<n<1M4 likes277 downloads6mo agoHugging Face06alibaba-multimodal-industrial-ai /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese. Overview Dimension Details Total questions 2,049 Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.textquestion-answering1K<n<10K30 likes211 downloads5mo agoHugging Face07alibabagroup /SKYLENAGE-GameCodeGym V-GameGym: Visual Game Generation for Code Large Language Models Abstract Code large language models have demonstrated remarkable capabilities in programming tasks, yet current benchmarks primarily focus on single modality rather than visual game development. Most existing code-related benchmarks evaluate syntax correctness and execution accuracy, overlooking critical game-specific metrics such as playability, visual aesthetics, and user engagement that are… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/SKYLENAGE-GameCodeGym.texttext-generation1K<n<10K4 likes100 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.