datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToolACE
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/Team-ACE/ToolACE.AceReason-1.1-SFT
AceReason-1.1-SFT
AceReason-1.1-SFT is a diverse and high-quality supervised fine-tuning (SFT) dataset focused on math and code reasoning. It serves as the SFT training data for AceReason-Nemotron-1.1-7B, with all responses in the dataset generated by DeepSeek-R1.
AceReason-1.1-SFT contains 2,668,741 math samples and 1,301,591 code samples, covering the data sources from OpenMathReasoning, NuminaMath-CoT, OpenCodeReasoning, MagicoderEvolInstruct, opc-sft-stage2, leetcode… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceReason-1.1-SFT.AceReason-Math
AceReason-Math Dataset
Overview
AceReason-Math is a high quality, verfiable, challenging and diverse math dataset for training math reasoning model using reinforcement leraning. This dataset contains
49K math problems and answer sourced from NuminaMath and DeepScaler-Preview
applying filtering rules to exclude unsuitable data (e.g., multiple sub-questions, multiple-choice, true/false, long and complex answers, proof, figure)
this dataset was used to train… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/AceReason-Math.ACE-SQL
ACE-SQL Training Data
This repository contains the curated supervised fine-tuning (SFT),
reinforcement learning (RL), and empirical-pool data released with
ACE-SQL: Adaptive Co-Optimization via Empirical Credit Assignment for
Text-to-SQL.
ACE-SQL trains a shared language-model policy in two roles: a schema retriever
that selects the minimum required database columns, and a SQL generator that
operates on the resulting pruned schema. The SFT data provides a cold start for
both… See the full description on the dataset page: https://huggingface.co/datasets/xiaobing11/ACE-SQL.ACE-Bench
ACE-Bench: Agent Coding Evaluation Benchmark
Dataset Description
ACE-Bench is a comprehensive benchmark designed to evaluate AI agents' capabilities in end-to-end feature-level code generation. Unlike traditional benchmarks that focus on function-level or algorithm-specific tasks, ACE-Bench challenges agents to implement complete features within real-world software projects.
Key Characteristics
Feature-Level Tasks: Each task requires implementing a complete… See the full description on the dataset page: https://huggingface.co/datasets/jiachengzhg/ACE-Bench.MSMS-AceReason-20K-SFT
A Multi-Source Multi-Solution Long CoT SFT Dataset from 20K AceReason Questions
blind-spot-fatima-institute-qwen3.5-0.8b
Evaluation Report of Qwen3.5-0.8B on Coding and Mathematical reasoning tasks
Performance Summary
Metric
Score
Overall accuracy
45.8%
Coding accuracy
33.3%
Math accuracy
58.3%
Total tests evaluated
24
Coding tests (12 total)
Result
Count
Correct
4
Partially correct
1
Incorrect
7
Math tests (12 total)
Result
Count
Correct
7
Partially correct
2
Incorrect
3
What the Model Did Well… See the full description on the dataset page: https://huggingface.co/datasets/Acesif/blind-spot-fatima-institute-qwen3.5-0.8b.acecode-87k-verl
AceCode-87K (VERL Format)
Overview
AceCode-87K dataset converted to VERL-compatible format for reinforcement learning training with code generation tasks.
Original Dataset: TIGER-Lab/AceCode-87K
License: MIT
Converted by: sungyub
Conversion Date: 2025-11-03
Dataset Statistics
Total Examples: 87,100
Split: train
Format: Parquet (VERL-compatible)
Data Sources:
OSS: 25857
APPS: 0
MBPP: 0
Schema
The dataset follows the VERL training format with the… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/acecode-87k-verl.ACE-Dataset
ACE Dataset
This dataset is designed for the formal evaluation of mathematical autoformalization consistency. It contains pairs of formal statements (Lean 4) that have been formally verified for semantic equivalence or non-equivalence.
Dataset Structure
equal/: Pairs of statements that are logically equivalent.
nonequal/: Pairs of statements that are logically non-equivalent.
Anonymization
This dataset is anonymized for double-blind review in NeurIPS 2026.… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-submission-ACE/ACE-Dataset.acereason-1.1-sft-math-mini
AceReason 1.1 SFT Math Mini
A math-only, length-bounded adaptation of nvidia/AceReason-1.1-SFT. Each row contains id, source, question, steps, and answer. Only category=math rows from OpenMathReasoning and NuminaMath-CoT are retained; every row satisfies the shared 3–50-step, explanatory-answer, 4,096-character, and 1,024-token contracts.
Dataset statistics
Metric
Value
Final records
31,737
Retention from 3,970,332 source rows
0.799%
Math source… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/acereason-1.1-sft-math-mini.babylm-ace
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: ace
Script: Latn
Tier: 1M
Byte Premium Factor: 1.241957
Size (MB): 6.74
Expected Size (MB): 6.74
Number of Documents: 20,883
Total Tokens: 968,194
Tokenizer: separate by whitespace
Tokens Per Category
child-books: 242,613 tokens
padding: 283,843 tokens
padding-wikipedia: 441… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-ace.
