CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Baoruixi /chimera-bench CHIMERA-Bench v1.0 A unified benchmark for epitope-specific antibody CDR sequence-structure co-design. Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop) Code: github.com/mansoorbaloch/chimera-bench Dataset Summary Property Value Complexes 2,922 PDB structures 2,721 Pre-computed features 2,941 .pt files Splits 3 (epitope-group, antigen-fold, temporal) Numbering schemes IMGT, Chothia Contact… See the full description on the dataset page: https://huggingface.co/datasets/Baoruixi/chimera-bench.tabularother1K<n<10K0 likes1.1k downloads2mo agoHugging Face02mansoorbaloch /chimera-bench CHIMERA-Bench v1.0 A unified benchmark for epitope-specific antibody CDR sequence-structure co-design. Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop) Code: github.com/mansoorbaloch/chimera-bench Dataset Summary Property Value Complexes 2,922 PDB structures 2,721 Pre-computed features 2,941 .pt files Splits 3 (epitope-group, antigen-fold, temporal) Numbering schemesIMGT, Chothia Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.tabularother1K<n<10K0 likes668 downloads4mo agoHugging Face03TianHongZXY /CHIMERA CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning CHIMERA is a compact but high-difficulty synthetic reasoning datasetwith long Chain-of-Thought (CoT) trajectories and broad STEM coverage, designed for reasoning post-training. All examples are fully LLM-generated and automatically verified without human annotation. Total: 9,225 problems Subjects: 8 Topics: 1,179 🔥 Why CHIMERA? Recent reasoning advances rely heavily on high-quality… See the full description on the dataset page: https://huggingface.co/datasets/TianHongZXY/CHIMERA.texttext-generation1K<n<10K24 likes312 downloads5mo agoHugging Face04AmanPriyanshu /reasoning-sft-CHIMERA reasoning-sft-CHIMERA Converted version of TianHongZXY/CHIMERA, filtered and reformatted for SFT/reasoning training. Both subsets (Qwen3-235B-2507 and Qwen3.5-397B) are included. No content was modified or regenerated, just reformatted the columns into a standard messages format. Filtering Kept only rows with correctness == True Randomly dropped 50% of Mathematics rows to reduce math dominance Both subsets combined into a single file Format Each row has three… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-CHIMERA.texttext-generation10K<n<100K0 likes229 downloads7mo agoHugging Face05VoidWalkercero /chimera-1000 CHIMERA-1000 The Benchmark That Refuses to Be Solved 1000 items · 10 cognitive dimensions · English + 中文 Designed when every existing benchmark is saturated or being saturated. v1.0 — August 2026 Why CHIMERA-1000 exists (EN) Every existing AI benchmark is saturated, contaminated, or structurally blind to the capabilities that actually matter for useful, safe AI: Problem with 2026 benchmarks CHIMERA's answer MMLU / GPQA-Diamond / MATH are saturated —… See the full description on the dataset page: https://huggingface.co/datasets/VoidWalkercero/chimera-1000.text1K<n<10K0 likes167 downloads2mo agoHugging Face06vectionlabs /chimera-bench-v1 CHIMERA Bench v3.0 Mega Comprehensive Hybrid Intelligence Metric for Excellence in Reasoning & Analysis 8503 articulated multi-step problems across 4 domains (larger than GSM8K). Domain Problems Focus MATH 3803 Multi-step word problems: shopping, speed/distance, geometry, combinatorics, algebra, number theory, calculus CODE 1500 Code tracing, bug finding, algorithm design, complexity analysis, OOP, recursion SCIENCE 1500 Physics (projectile, energy, circuits)… See the full description on the dataset page: https://huggingface.co/datasets/vectionlabs/chimera-bench-v1.tabulartext-generation1K<n<10K6 likes71 downloads5mo agoHugging Face07fabricioslv /omnimind-human-mouse-chimera-dopamine-inventory OmniMind — Human-mouse chimera dopamine neurons inventory Source: https://zenodo.org/records/10215356 Paper: Dawson et al., Stem Cell Reports — Interspecies chimerism with human embryonic stem cells generates functional human dopamine neurons at low efficiency This is an inventory of the 508 files in the RAR archive (702 MB), not a full extraction. Most files are TIFF/JPG microscopy images and .fcs flow cytometry files. Also extracted: injection numbers.xlsx pups death… See the full description on the dataset page: https://huggingface.co/datasets/fabricioslv/omnimind-human-mouse-chimera-dopamine-inventory.textn<1K0 likes46 downloads22d agoHugging Face08danielhanchen /chimeratext10K<n<100K0 likes40 downloads3y agoHugging Face09fzzhang /qwen3_4b_chimera_fixedtopics_questions_nofiltertext10K<n<100K1 likes40 downloads4mo agoHugging Face10PJMixers-Dev /Some-RP-v2-R1T2-Chimera-allModel turns in ToastyPigeon/some-rp-v2 regenerated using tngtech/DeepSeek-TNG-R1T2-Chimera. You should mask everything except the last turn when training. All previous model turns are the original dataset. It's setup to be trained like R1: NousResearch/Minos-v1 was used to avoid refusals. Only checked against <|user|>\n{latest_user_turn}\n<|assistant|>\n{response_without_thinking}, regenerating if not at least 80% confident it's a non-refusal. textn<1K0 likes36 downloads9mo agoHugging Face11anonymous-somebody /CHIMERA CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation. Total: 9,225 problems Subjects: 8 Topics: 1,179 Overview Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.texttext-generation1K<n<10K0 likes35 downloads5mo agoHugging Face12Umranz /Chimera-XTRM-Dataset-v1gatedtext1K<n<10K1 likes35 downloads3mo agoHugging Face13PJMixers-Dev /foundRP-R1T2-Chimera-allModel turns in BeaverAI/foundRP regenerated using tngtech/DeepSeek-TNG-R1T2-Chimera. You should mask everything except the last turn when training. All previous model turns are the original dataset. It's setup to be trained like R1: NousResearch/Minos-v1 was used to avoid refusals. Only checked against <|user|>\n{latest_user_turn}\n<|assistant|>\n{response_without_thinking}, regenerating if not at least 80% confident it's a non-refusal. textn<1K0 likes34 downloads9mo agoHugging Face14open-llm-leaderboard /mlabonne__ChimeraLlama-3-8B-v3-detailsgated Dataset Card for Evaluation run of mlabonne/ChimeraLlama-3-8B-v3 Dataset automatically created during the evaluation run of model mlabonne/ChimeraLlama-3-8B-v3 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlabonne__ChimeraLlama-3-8B-v3-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face15fzzhang /qwen3_8b_chimera_fixedtopics_highschool_75k_nofilter_instill_n8_valredundancy5_round1text100K<n<1M0 likes28 downloads2mo agoHugging Face16open-llm-leaderboard /utkmst__chimera-beta-test2-lora-merged-detailsgated Dataset Card for Evaluation run of utkmst/chimera-beta-test2-lora-merged Dataset automatically created during the evaluation run of model utkmst/chimera-beta-test2-lora-merged The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/utkmst__chimera-beta-test2-lora-merged-details.tabular10K<n<100K1 likes26 downloads2y agoHugging Face17PJMixers-Dev /Gryphe-Sonnet3.5-Charcard-Roleplay-R1T2-Chimera-alltextn<1K1 likes26 downloads9mo agoHugging Face18fzzhang /qwen3_32b_chimera_fixedtopics_questions_nofiltertext10K<n<100K0 likes26 downloads3mo agoHugging Face19fzzhang /qwen3_32b_chimera_fixedtopics_probval_instill_n8_valredundancy5_round1text10K<n<100K0 likes25 downloads3mo agoHugging Face20fzzhang /qwen3_8b_chimera_fixedtopics_highschool_75k_probval_instill_n8_valredundancy5_round1text10K<n<100K0 likes25 downloads2mo agoHugging Face21open-llm-leaderboard /mlabonne__ChimeraLlama-3-8B-v2-detailsgated Dataset Card for Evaluation run of mlabonne/ChimeraLlama-3-8B-v2 Dataset automatically created during the evaluation run of model mlabonne/ChimeraLlama-3-8B-v2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlabonne__ChimeraLlama-3-8B-v2-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face22fzzhang /qwen3_8b_chimera_fixedtopics_75k_probval_instill_n8_valredundancy5_round1text10K<n<100K0 likes22 downloads3mo agoHugging Face23PJMixers-Dev /Gryphe-Aesir-RPG-Charcards-Opus-Mixed-R1T2-Chimera-alltextn<1K0 likes21 downloads9mo agoHugging Face24fzzhang /qwen3_8b_chimera_fixedtopics_questions_probvaltext10K<n<100K0 likes18 downloads3mo agoHugging Face25fzzhang /qwen3_4b_chimera_treetopics_v1_questions_probvaltext10K<n<100K0 likes17 downloads4mo agoHugging Face26CHIzhP /Chimeraimage1K<n<10K3 likes16 downloads1y agoHugging Face27PJMixers-Dev /LMSYS-Chat-1M-R1T2-Chimera-allModel turns in natong19/lmsys-chat-1m-filtered regenerated using tngtech/DeepSeek-TNG-R1T2-Chimera. You should mask everything except the last turn when training. All previous model turns are the original dataset. It's setup to be trained like R1: NousResearch/Minos-v1 was used to avoid refusals. Only checked against <|user|>\n{latest_user_turn}\n<|assistant|>\n{response_without_thinking}, regenerating if not at least 80% confident it's a non-refusal. textn<1K0 likes15 downloads9mo agoHugging Face28PJMixers-Dev /CoSER-R1T2-Chimera-alltextn<1K0 likes15 downloads9mo agoHugging Face29rx1lora /StoryPlay_Gryphe-Sonnet3.5-Charcard-Roleplay-R1T2-Chimera-alltextn<1K0 likes14 downloads2mo agoHugging Face30AI-Foundation /chimera-training-data Chimera Training Data Training data for Project Chimera — research into widened I/O bandwidth for AI models. The Goal Train a single neural network to process multiple inputs and generate multiple outputs simultaneously. Not interleaving, not fast-switching — genuinely parallel cognitive streams from ONE brain. ┌─────────────────────────────────────────────┐ Input A ─┤ ├─ Output A │ ONE BRAIN… See the full description on the dataset page: https://huggingface.co/datasets/AI-Foundation/chimera-training-data.texttext-generation100K<n<1M0 likes13 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.