datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tmax-9b-atlas
allenai/tmax-9b Brain Atlas — The Sweet Spot of the Hybrid Family
Cross-post: I ran a brain atlas on the mid-size tmax. Sub-Zero coverage is concentrated in layers 16–30, so read the surgical headroom numbers as a late-layer snapshot.
model: allenai/tmax-9batlas type: activation census + Sub-Zero brain atlas + OV-circuit SVDcorpus: 8,965 promptslayers: 32attention layers: 3, 7, 11, 15, 19, 23, 27, 31hybrid layers: everything elsesacred (fully probed) layers:… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-9b-atlas.qwen3.5-9b-atlas
qwen3.5-9b-atlas
Ornith-1.0-9B-atlas
juiceb0xc0de/Ornith-1.0-9B-atlas
A brain atlas for deepreinforce-ai/Ornith-1.0-9B, the 9B agentic-coding model that reports SOTA results on Terminal-Bench, SWE-Bench, and other agentic coding benchmarks. This is not a chat dataset or a benchmark — it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know why this model survives surgical edits, where… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Ornith-1.0-9B-atlas.Qwen3.5-9B-Base
juiceb0xc0de/Qwen3.5-9B-Base
A brain atlas for Qwen/Qwen3.5-9B-Base, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
This is a base model, before any instruction tuning. That makes it a useful thing to have a map of: whatever… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-9B-Base.Qwythos-9B-Claude-Mythos-5-1M-atlas
juiceb0xc0de/Qwythos-9B-Claude-Mythos-5-1M-atlas
A brain atlas for empero-ai/Qwythos-9B-Claude-Mythos-5-1M, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
The interesting thing about this model is how little of it is full… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwythos-9B-Claude-Mythos-5-1M-atlas.mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 9B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k.pack-0t8-of-my-project-9bvey3-57582a13
Industrial VLA Training Pack — Grasping & Manipulation
Training dataset for the OpenVLA vision-language-action model in industrial environments (warehouse and garage workshops). Covers grasping, pick-and-place and manipulation of graspable industrial objects such as tool boxes, drills, screwdrivers, boxes and crates. 100 renders at 224x224 with metric depth, world-space normals (OpenGL linear) and material index passes, per-frame annotations enabled, using each environment's… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/pack-0t8-of-my-project-9bvey3-57582a13.flux-2-klein-9B-schematic-dataset
FLUX.2 Klein Schematic LoRA Dataset
This is the training dataset used for the FLUX.2 Klein Schematic LoRA project.
The LoRA repository is available here:
LoRA: https://huggingface.co/nomadoor/flux-2-klein-9B-schematic-lora
For details about the experiment and dataset construction, see the blog post:
Blog: https://comfyui.nomadoor.net/en/notes/flux2-klein-schematic-lora/
Structure
data/
depth_relative/
input/
target/
caption/
surface_normal/… See the full description on the dataset page: https://huggingface.co/datasets/nomadoor/flux-2-klein-9B-schematic-dataset.mobilegym-trajectories-autoglm-phone-9b
MobileGym Trajectories — AutoGLM-Phone-9B (L1)
Agent rollout trajectories collected on the MobileGym simulated-Android
environment with AutoGLM-Phone-9B as the policy. One row per episode (rollout), in the spirit
of SWE-Gym-style instance datasets. Each row carries the task spec, the full multimodal
interaction (screenshots + model responses + parsed actions), and a deterministic reward
from MobileGym's JSON-state judge (no VLM judging — verdicts are exact).… See the full description on the dataset page: https://huggingface.co/datasets/gray311/mobilegym-trajectories-autoglm-phone-9b.humatheque-vlm-pred-ornith15-9bivf-bench-orpo-qwen9b
IVF-Bench ORPO preference pairs
550 preference pairs for training a vision-language model to write clinical
assessments of IVF embryo cases. Each row holds a day-5 blastocyst image, the
case prompt, a preferred response, and a rejected one.
Built from IVF-Bench. The underlying
embryo images and real clinical fields come from the
Kromp et al. (2023) blastocyst
dataset, released under CC BY 4.0.
How the pairs were made
Seven vision-language models answered each of… See the full description on the dataset page: https://huggingface.co/datasets/thefertilityplan/ivf-bench-orpo-qwen9b.ivf-bench-orpo-qwen9b-clipped
IVF-Bench ORPO preference pairs
550 preference pairs for training a vision-language model to write clinical
assessments of IVF embryo cases. Each row holds a day-5 blastocyst image, the
case prompt, a preferred response, and a rejected one.
Built from IVF-Bench. The underlying
embryo images and real clinical fields come from the
Kromp et al. (2023) blastocyst
dataset, released under CC BY 4.0.
How the pairs were made
Seven vision-language models answered each of… See the full description on the dataset page: https://huggingface.co/datasets/thefertilityplan/ivf-bench-orpo-qwen9b-clipped.Works_with_Flux2_Klein_9B_TrueSample images generated with "Flux2-Klein-9B-True" FP8 published by @wikeeyang (https://huggingface.co/wikeeyang), thanks!
Configuration on Forge Neo:
Steps: 4
Sampler: Euler
Schedule type: Beta
CFG scale: 1
Seed: random
Model: flux2Klein9BTrue_v10Fp8
Clip skip: 2
RNG: CPU
Module 1: flux2-vae
Module 2: qwen_3_8b
humatheque-vlm-pred-qwen35-9bMMK12-QWEN35-9B-VERIFIED
Dataset card for MMK12-QWEN35-9B-VERIFIED
Strongich/MMK12-QWEN35-9B-VERIFIED is a filtered teacher-trace dataset built from FanqingM/MMK12 using Qwen/Qwen3.5-9B. The goal is to provide high-quality multimodal reasoning traces that can be used for off-policy distillation into smaller models.
Each row keeps the original MMK12 example fields and adds a responses column. The responses column contains a list of verified Qwen responses that survived the filtering pipeline, so a… See the full description on the dataset page: https://huggingface.co/datasets/Strongich/MMK12-QWEN35-9B-VERIFIED.xdownloader.com__9bd73-part1d2ec2286-10fb-4662-9b72-c146b7d7293c_training_20251128111529
PAL FullFlow Golden - LoRA Training Dataset
Training dataset for PAL FullFlow Golden character LoRA used with WAN 2.2.
Dataset Information
Character: PAL FullFlow Golden
Trigger Word: chr_pal_fullflow_golden
ZIP Size: 7.0 MB
File: training_dataset.zip
Character Attributes
Build: curvy
Ethnicity: mixed ethnicity
Facial Features: oval face, high cheekbones, dark arched eyebrows, dark eyes, full lips, defined jawline
Hair: long, dark brown, pulled back into a… See the full description on the dataset page: https://huggingface.co/datasets/KozMi/d2ec2286-10fb-4662-9b72-c146b7d7293c_training_20251128111529.twitter-xingyinshaonv_-2025.09.20-1969263342637248562-9biMvQUnMRqbrhLM-part1ebd7817f-9bec-4ada-a48d-3c7dd3bf34eaxdownloader.com__9bd02-part1dataset_9bbc16a8-d43a-4b5f-83d0-c796ef8b2f6b69d597fd-fd0a-4df4-bb9b-705805a5481e_training_20251125030259
CLI Character - WAN 2.2 1764039778 - LoRA Training Dataset
Training dataset for CLI Character - WAN 2.2 1764039778 character LoRA used with WAN 2.2.
Dataset Information
Character: CLI Character - WAN 2.2 1764039778
Trigger Word: cli_user_wan_1764039778_1764039778_trigger
ZIP Size: 0.0 MB
File: training_dataset.zip
Character Attributes
Build: athletic
Ethnicity: latin
Facial Features: oval face, bright eyes, confident smile
Hair: dark wavy hair
Distinctive… See the full description on the dataset page: https://huggingface.co/datasets/KozMi/69d597fd-fd0a-4df4-bb9b-705805a5481e_training_20251125030259.twitter-gxmn2023-2026.03.17-2033834586883682528-R7FxbVo7ye7D9B-a-part1twitter-Meixyxtxin-2025.12.30-2005945766905254187-9bUdiV5wLkCqFhEj-part1twitter-ws_yiyi-2025.09.28-1972287606802247937-Sl_9boVo_ceSNvcE-part1xdownloader.com__9bc85-part1
