datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chimera-bench
CHIMERA-Bench v1.0
A unified benchmark for epitope-specific antibody CDR sequence-structure co-design.
Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop)
Code: github.com/mansoorbaloch/chimera-bench
Dataset Summary
Property
Value
Complexes
2,922
PDB structures
2,721
Pre-computed features
2,941 .pt files
Splits
3 (epitope-group, antigen-fold, temporal)
Numbering schemes
IMGT, Chothia
Contact… See the full description on the dataset page: https://huggingface.co/datasets/Baoruixi/chimera-bench.chimera-bench
CHIMERA-Bench v1.0
A unified benchmark for epitope-specific antibody CDR sequence-structure co-design.
Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop)
Code: github.com/mansoorbaloch/chimera-bench
Dataset Summary
Property
Value
Complexes
2,922
PDB structures
2,721
Pre-computed features
2,941 .pt files
Splits
3 (epitope-group, antigen-fold, temporal)
Numbering schemesIMGT, Chothia
Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact but high-difficulty synthetic reasoning datasetwith long Chain-of-Thought (CoT) trajectories and broad STEM coverage, designed for reasoning post-training. All examples are fully LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
🔥 Why CHIMERA?
Recent reasoning advances rely heavily on high-quality… See the full description on the dataset page: https://huggingface.co/datasets/TianHongZXY/CHIMERA.reasoning-sft-CHIMERA
reasoning-sft-CHIMERA
Converted version of TianHongZXY/CHIMERA, filtered and reformatted for SFT/reasoning training. Both subsets (Qwen3-235B-2507 and Qwen3.5-397B) are included. No content was modified or regenerated, just reformatted the columns into a standard messages format.
Filtering
Kept only rows with correctness == True
Randomly dropped 50% of Mathematics rows to reduce math dominance
Both subsets combined into a single file
Format
Each row has three… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-CHIMERA.chimera-1000
CHIMERA-1000
The Benchmark That Refuses to Be Solved
1000 items · 10 cognitive dimensions · English + 中文
Designed when every existing benchmark is saturated or being saturated.
v1.0 — August 2026
Why CHIMERA-1000 exists (EN)
Every existing AI benchmark is saturated, contaminated, or structurally blind to the capabilities that actually matter for useful, safe AI:
Problem with 2026 benchmarks
CHIMERA's answer
MMLU / GPQA-Diamond / MATH are saturated —… See the full description on the dataset page: https://huggingface.co/datasets/VoidWalkercero/chimera-1000.chimera-bench-v1
CHIMERA Bench v3.0 Mega
Comprehensive Hybrid Intelligence Metric for Excellence in Reasoning & Analysis
8503 articulated multi-step problems across 4 domains (larger than GSM8K).
Domain
Problems
Focus
MATH
3803
Multi-step word problems: shopping, speed/distance, geometry, combinatorics, algebra, number theory, calculus
CODE
1500
Code tracing, bug finding, algorithm design, complexity analysis, OOP, recursion
SCIENCE
1500
Physics (projectile, energy, circuits)… See the full description on the dataset page: https://huggingface.co/datasets/vectionlabs/chimera-bench-v1.omnimind-human-mouse-chimera-dopamine-inventory
OmniMind — Human-mouse chimera dopamine neurons inventory
Source: https://zenodo.org/records/10215356
Paper: Dawson et al., Stem Cell Reports — Interspecies chimerism with human embryonic stem cells generates functional human dopamine neurons at low efficiency
This is an inventory of the 508 files in the RAR archive (702 MB), not a full extraction.
Most files are TIFF/JPG microscopy images and .fcs flow cytometry files.
Also extracted:
injection numbers.xlsx
pups death… See the full description on the dataset page: https://huggingface.co/datasets/fabricioslv/omnimind-human-mouse-chimera-dopamine-inventory.chimeraqwen3_4b_chimera_fixedtopics_questions_nofilterSome-RP-v2-R1T2-Chimera-allModel turns in ToastyPigeon/some-rp-v2 regenerated using tngtech/DeepSeek-TNG-R1T2-Chimera.
You should mask everything except the last turn when training. All previous model turns are the original dataset.
It's setup to be trained like R1:
NousResearch/Minos-v1 was used to avoid refusals. Only checked against <|user|>\n{latest_user_turn}\n<|assistant|>\n{response_without_thinking}, regenerating if not at least 80% confident it's a non-refusal.
CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
Overview
Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.Chimera-XTRM-Dataset-v1foundRP-R1T2-Chimera-allModel turns in BeaverAI/foundRP regenerated using tngtech/DeepSeek-TNG-R1T2-Chimera.
You should mask everything except the last turn when training. All previous model turns are the original dataset.
It's setup to be trained like R1:
NousResearch/Minos-v1 was used to avoid refusals. Only checked against <|user|>\n{latest_user_turn}\n<|assistant|>\n{response_without_thinking}, regenerating if not at least 80% confident it's a non-refusal.
mlabonne__ChimeraLlama-3-8B-v3-details
Dataset Card for Evaluation run of mlabonne/ChimeraLlama-3-8B-v3
Dataset automatically created during the evaluation run of model mlabonne/ChimeraLlama-3-8B-v3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlabonne__ChimeraLlama-3-8B-v3-details.qwen3_8b_chimera_fixedtopics_highschool_75k_nofilter_instill_n8_valredundancy5_round1utkmst__chimera-beta-test2-lora-merged-details
Dataset Card for Evaluation run of utkmst/chimera-beta-test2-lora-merged
Dataset automatically created during the evaluation run of model utkmst/chimera-beta-test2-lora-merged
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/utkmst__chimera-beta-test2-lora-merged-details.Gryphe-Sonnet3.5-Charcard-Roleplay-R1T2-Chimera-allqwen3_32b_chimera_fixedtopics_questions_nofilterqwen3_32b_chimera_fixedtopics_probval_instill_n8_valredundancy5_round1qwen3_8b_chimera_fixedtopics_highschool_75k_probval_instill_n8_valredundancy5_round1mlabonne__ChimeraLlama-3-8B-v2-details
Dataset Card for Evaluation run of mlabonne/ChimeraLlama-3-8B-v2
Dataset automatically created during the evaluation run of model mlabonne/ChimeraLlama-3-8B-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlabonne__ChimeraLlama-3-8B-v2-details.qwen3_8b_chimera_fixedtopics_75k_probval_instill_n8_valredundancy5_round1Gryphe-Aesir-RPG-Charcards-Opus-Mixed-R1T2-Chimera-allqwen3_8b_chimera_fixedtopics_questions_probvalqwen3_4b_chimera_treetopics_v1_questions_probvalChimeraLMSYS-Chat-1M-R1T2-Chimera-allModel turns in natong19/lmsys-chat-1m-filtered regenerated using tngtech/DeepSeek-TNG-R1T2-Chimera.
You should mask everything except the last turn when training. All previous model turns are the original dataset.
It's setup to be trained like R1:
NousResearch/Minos-v1 was used to avoid refusals. Only checked against <|user|>\n{latest_user_turn}\n<|assistant|>\n{response_without_thinking}, regenerating if not at least 80% confident it's a non-refusal.
CoSER-R1T2-Chimera-allStoryPlay_Gryphe-Sonnet3.5-Charcard-Roleplay-R1T2-Chimera-allchimera-training-data
Chimera Training Data
Training data for Project Chimera — research into widened I/O bandwidth for AI models.
The Goal
Train a single neural network to process multiple inputs and generate multiple outputs simultaneously. Not interleaving, not fast-switching — genuinely parallel cognitive streams from ONE brain.
┌─────────────────────────────────────────────┐
Input A ─┤ ├─ Output A
│ ONE BRAIN… See the full description on the dataset page: https://huggingface.co/datasets/AI-Foundation/chimera-training-data.
