datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tmax-9b-atlas
allenai/tmax-9b Brain Atlas — The Sweet Spot of the Hybrid Family
Cross-post: I ran a brain atlas on the mid-size tmax. Sub-Zero coverage is concentrated in layers 16–30, so read the surgical headroom numbers as a late-layer snapshot.
model: allenai/tmax-9batlas type: activation census + Sub-Zero brain atlas + OV-circuit SVDcorpus: 8,965 promptslayers: 32attention layers: 3, 7, 11, 15, 19, 23, 27, 31hybrid layers: everything elsesacred (fully probed) layers:… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-9b-atlas.qwen3.5-9b-atlas
qwen3.5-9b-atlas
appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1
appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3953125
Action score: 0.446875
Valid samples: 320/320
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1-reeval1
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.41328125
Action score: 0.4359375
Valid samples: 320/320
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1
appworld-qwen35-4b-9b-s_signal_5-epoch4-iter1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.41953125
Action score: 0.4515625
Valid samples: 320/320
Ornith-1.0-9B-atlas
juiceb0xc0de/Ornith-1.0-9B-atlas
A brain atlas for deepreinforce-ai/Ornith-1.0-9B, the 9B agentic-coding model that reports SOTA results on Terminal-Bench, SWE-Bench, and other agentic coding benchmarks. This is not a chat dataset or a benchmark — it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know why this model survives surgical edits, where… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Ornith-1.0-9B-atlas.details_01-ai__Yi-9B-200K
Dataset Card for Evaluation run of 01-ai/Yi-9B-200K
Dataset automatically created during the evaluation run of model 01-ai/Yi-9B-200K.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_01-ai__Yi-9B-200K.Qwen3.5-9B-Base
juiceb0xc0de/Qwen3.5-9B-Base
A brain atlas for Qwen/Qwen3.5-9B-Base, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
This is a base model, before any instruction tuning. That makes it a useful thing to have a map of: whatever… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-9B-Base.qwen-9b-3m
qwen-9b-3m
Multi-domain SFT-target dataset: ~3,000,000 prompts, each with ONE completion
generated offline by Qwen/Qwen3.5-9B (thinking mode). Exported snapshot from an
offline queue pipeline; repartitioned into 512 parquet shards.
Columns (21)
record_index, input_sha256, prompt_sha256, dataset, split, source, upstream_id,
bucket, messages_json, prompt, prompt_token_count, generation_seed, enable_thinking,
worker, executor_worker, completion, completion_input_ids… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/qwen-9b-3m.llama-9b-bulk-npzQwythos-9B-Claude-Mythos-5-1M-atlas
juiceb0xc0de/Qwythos-9B-Claude-Mythos-5-1M-atlas
A brain atlas for empero-ai/Qwythos-9B-Claude-Mythos-5-1M, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
The interesting thing about this model is how little of it is full… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwythos-9B-Claude-Mythos-5-1M-atlas.Ornith-1.5-9B-GGUF-metricsafrica-drc-congo-dem-rep-environment-9b29113f
Congo, Dem. Rep. - Environment | Africa (DRC official open data)
4,653 rows - 1 Africa country - 1960-2025 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from DRC as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Congo, Dem. Rep. - Environment
Publisher: World Bank Group… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-drc-congo-dem-rep-environment-9b29113f.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.spite-gigaspeech-Euro9B
Spite Dataset
Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from GigaSpeech and translations from EuroLLM-9B-Instruct.
Configs
en_de
en_es
en_fr
en_it
en_ko
en_nl
en_pt
en_ru
en_zh
Usage
from datasets import load_dataset
ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt")
details_Delta-Vector__Odin-9B
Dataset Card for Evaluation run of Delta-Vector/Odin-9B
Dataset automatically created during the evaluation run of model Delta-Vector/Odin-9B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Delta-Vector__Odin-9B.gemma2_9b_it_gsm8k_mcmc_with_promptgenerations-nemotron-nano-9b-v2-simnpo-gentle-igm-10bdetails_NotAiLOL__Yi-1.5-dolphin-9B
Dataset Card for Evaluation run of NotAiLOL/Yi-1.5-dolphin-9B
Dataset automatically created during the evaluation run of model NotAiLOL/Yi-1.5-dolphin-9B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_NotAiLOL__Yi-1.5-dolphin-9B.generations-nemotron-nano-9b-v2-simnpo-gentle-baselinedetails_anthracite-org__magnum-v3-9b-chatml
Dataset Card for Evaluation run of anthracite-org/magnum-v3-9b-chatml
Dataset automatically created during the evaluation run of model anthracite-org/magnum-v3-9b-chatml.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_anthracite-org__magnum-v3-9b-chatml.generations-nemotron-nano-9b-v2-simnpo-baselinedetails_anthracite-org__magnum-v3-9b-customgemma2
Dataset Card for Evaluation run of anthracite-org/magnum-v3-9b-customgemma2
Dataset automatically created during the evaluation run of model anthracite-org/magnum-v3-9b-customgemma2.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_anthracite-org__magnum-v3-9b-customgemma2.generations-nemotron-nano-9b-v2-pre_valQwen3.5-9B-EquationalTheories-Proof-eval
Qwen3.5-9B on EquationalTheories-Proof: equational internalization evaluation
Model: the pinned base Qwen3.5-9B (no training)
Evaluation of Qwen/Qwen3.5-9B (revision c202236235762e1c871ad0ccb60c8ee5ba337b9a) on the equational internalization
protocol (recipe/equational_internalization/EVALUATION_PROTOCOL.md, environment
EquationalTheories-Proof): closed-book recall of implication/non-implication relationships between the
4,694 equations of the Equational Theories Project… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/Qwen3.5-9B-EquationalTheories-Proof-eval.details_byroneverson__Yi-1.5-9B-Chat-16K-abliterated
Dataset Card for Evaluation run of byroneverson/Yi-1.5-9B-Chat-16K-abliterated
Dataset automatically created during the evaluation run of model byroneverson/Yi-1.5-9B-Chat-16K-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_byroneverson__Yi-1.5-9B-Chat-16K-abliterated.details_byroneverson__Yi-1.5-9B-Chat-abliterated
Dataset Card for Evaluation run of byroneverson/Yi-1.5-9B-Chat-abliterated
Dataset automatically created during the evaluation run of model byroneverson/Yi-1.5-9B-Chat-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_byroneverson__Yi-1.5-9B-Chat-abliterated.details_01-ai__Yi-1.5-9B-Chat-16K
Dataset Card for Evaluation run of 01-ai/Yi-1.5-9B-Chat-16K
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-9B-Chat-16K.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_01-ai__Yi-1.5-9B-Chat-16K.
