datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact but high-difficulty synthetic reasoning datasetwith long Chain-of-Thought (CoT) trajectories and broad STEM coverage, designed for reasoning post-training. All examples are fully LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
🔥 Why CHIMERA?
Recent reasoning advances rely heavily on high-quality… See the full description on the dataset page: https://huggingface.co/datasets/TianHongZXY/CHIMERA.reasoning-sft-CHIMERA
reasoning-sft-CHIMERA
Converted version of TianHongZXY/CHIMERA, filtered and reformatted for SFT/reasoning training. Both subsets (Qwen3-235B-2507 and Qwen3.5-397B) are included. No content was modified or regenerated, just reformatted the columns into a standard messages format.
Filtering
Kept only rows with correctness == True
Randomly dropped 50% of Mathematics rows to reduce math dominance
Both subsets combined into a single file
Format
Each row has three… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-CHIMERA.chimera-bench-v1
CHIMERA Bench v3.0 Mega
Comprehensive Hybrid Intelligence Metric for Excellence in Reasoning & Analysis
8503 articulated multi-step problems across 4 domains (larger than GSM8K).
Domain
Problems
Focus
MATH
3803
Multi-step word problems: shopping, speed/distance, geometry, combinatorics, algebra, number theory, calculus
CODE
1500
Code tracing, bug finding, algorithm design, complexity analysis, OOP, recursion
SCIENCE
1500
Physics (projectile, energy, circuits)… See the full description on the dataset page: https://huggingface.co/datasets/vectionlabs/chimera-bench-v1.CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
Overview
Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.chimera-training-data
Chimera Training Data
Training data for Project Chimera — research into widened I/O bandwidth for AI models.
The Goal
Train a single neural network to process multiple inputs and generate multiple outputs simultaneously. Not interleaving, not fast-switching — genuinely parallel cognitive streams from ONE brain.
┌─────────────────────────────────────────────┐
Input A ─┤ ├─ Output A
│ ONE BRAIN… See the full description on the dataset page: https://huggingface.co/datasets/AI-Foundation/chimera-training-data.chimera-curated-68k
CHIMERA Curated 68K
A quality-gated instruction dataset for reasoning, code, and post-training.
CHIMERA Curated 68K is a 68,101-sample instruction-following dataset built for SFT and post-training runs where data quality matters more than raw scale.
The dataset is focused on high-signal reasoning and code examples, with a smaller conversational component included for broader instruction-following behavior. Every admitted sample passes through a quality-gated curation process… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/chimera-curated-68k.
