CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Team-ACE /ToolACE ToolACE ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. More details… See the full description on the dataset page: https://huggingface.co/datasets/Team-ACE/ToolACE.texttext-generation10K<n<100K199 likes28k downloads2y agoHugging Face02nvidia /AceReason-1.1-SFT AceReason-1.1-SFT AceReason-1.1-SFT is a diverse and high-quality supervised fine-tuning (SFT) dataset focused on math and code reasoning. It serves as the SFT training data for AceReason-Nemotron-1.1-7B, with all responses in the dataset generated by DeepSeek-R1. AceReason-1.1-SFT contains 2,668,741 math samples and 1,301,591 code samples, covering the data sources from OpenMathReasoning, NuminaMath-CoT, OpenCodeReasoning, MagicoderEvolInstruct, opc-sft-stage2, leetcode… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceReason-1.1-SFT.texttext-generation1M<n<10M101 likes3.6k downloads1y agoHugging Face03nvidia /AceReason-Math AceReason-Math Dataset Overview AceReason-Math is a high quality, verfiable, challenging and diverse math dataset for training math reasoning model using reinforcement leraning. This dataset contains 49K math problems and answer sourced from NuminaMath and DeepScaler-Preview applying filtering rules to exclude unsuitable data (e.g., multiple sub-questions, multiple-choice, true/false, long and complex answers, proof, figure) this dataset was used to train… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceReason-Math.texttext-generation10K<n<100K57 likes1.5k downloads1y agoHugging Face04xiaobing11 /ACE-SQL ACE-SQL Training Data This repository contains the curated supervised fine-tuning (SFT), reinforcement learning (RL), and empirical-pool data released with ACE-SQL: Adaptive Co-Optimization via Empirical Credit Assignment for Text-to-SQL. ACE-SQL trains a shared language-model policy in two roles: a schema retriever that selects the minimum required database columns, and a SQL generator that operates on the resulting pruned schema. The SFT data provides a cold start for both… See the full description on the dataset page: https://huggingface.co/datasets/xiaobing11/ACE-SQL.texttext-generation10K<n<100K1 likes92 downloads3mo agoHugging Face05jiachengzhg /ACE-Bench ACE-Bench: Agent Coding Evaluation Benchmark Dataset Description ACE-Bench is a comprehensive benchmark designed to evaluate AI agents' capabilities in end-to-end feature-level code generation. Unlike traditional benchmarks that focus on function-level or algorithm-specific tasks, ACE-Bench challenges agents to implement complete features within real-world software projects. Key Characteristics Feature-Level Tasks: Each task requires implementing a complete… See the full description on the dataset page: https://huggingface.co/datasets/jiachengzhg/ACE-Bench.texttext-generationn<1K0 likes81 downloads10mo agoHugging Face06wangbing1416 /MSMS-AceReason-20K-SFT A Multi-Source Multi-Solution Long CoT SFT Dataset from 20K AceReason Questions texttext-generation100K<n<1M1 likes37 downloads9mo agoHugging Face07Acesif /blind-spot-fatima-institute-qwen3.5-0.8b Evaluation Report of Qwen3.5-0.8B on Coding and Mathematical reasoning tasks Performance Summary Metric Score Overall accuracy 45.8% Coding accuracy 33.3% Math accuracy 58.3% Total tests evaluated 24 Coding tests (12 total) Result Count Correct 4 Partially correct 1 Incorrect 7 Math tests (12 total) Result Count Correct 7 Partially correct 2 Incorrect 3 What the Model Did Well… See the full description on the dataset page: https://huggingface.co/datasets/Acesif/blind-spot-fatima-institute-qwen3.5-0.8b.texttext-generation0 likes27 downloads6mo agoHugging Face08sungyub /acecode-87k-verl AceCode-87K (VERL Format) Overview AceCode-87K dataset converted to VERL-compatible format for reinforcement learning training with code generation tasks. Original Dataset: TIGER-Lab/AceCode-87K License: MIT Converted by: sungyub Conversion Date: 2025-11-03 Dataset Statistics Total Examples: 87,100 Split: train Format: Parquet (VERL-compatible) Data Sources: OSS: 25857 APPS: 0 MBPP: 0 Schema The dataset follows the VERL training format with the… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/acecode-87k-verl.texttext-generation10K<n<100K0 likes24 downloads11mo agoHugging Face09neurips-2026-submission-ACE /ACE-Dataset ACE Dataset This dataset is designed for the formal evaluation of mathematical autoformalization consistency. It contains pairs of formal statements (Lean 4) that have been formally verified for semantic equivalence or non-equivalence. Dataset Structure equal/: Pairs of statements that are logically equivalent. nonequal/: Pairs of statements that are logically non-equivalent. Anonymization This dataset is anonymized for double-blind review in NeurIPS 2026.… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-submission-ACE/ACE-Dataset.text-generation0 likes18 downloads5mo agoHugging Face10cs-giung /acereason-1.1-sft-math-mini AceReason 1.1 SFT Math Mini A math-only, length-bounded adaptation of nvidia/AceReason-1.1-SFT. Each row contains id, source, question, steps, and answer. Only category=math rows from OpenMathReasoning and NuminaMath-CoT are retained; every row satisfies the shared 3–50-step, explanatory-answer, 4,096-character, and 1,024-token contracts. Dataset statistics Metric Value Final records 31,737 Retention from 3,970,332 source rows 0.799% Math source… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/acereason-1.1-sft-math-mini.texttext-generation10K<n<100K0 likes18 downloads1mo agoHugging Face11BabyLM-community /babylm-acegated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: ace Script: Latn Tier: 1M Byte Premium Factor: 1.241957 Size (MB): 6.74 Expected Size (MB): 6.74 Number of Documents: 20,883 Total Tokens: 968,194 Tokenizer: separate by whitespace Tokens Per Category child-books: 242,613 tokens padding: 283,843 tokens padding-wikipedia: 441… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ace.texttext-generation10K<n<100K0 likes3 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.