datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma-4-e2b-atlas
financial-english-source-corpus-gemma4-e2b-1280gemma-4-e4b-it-atlas
juiceb0xc0de/gemma-4-e4b-it-atlas
A brain atlas for google/gemma-4-E4B-it, the instruction-tuned E4B member of the Gemma 4 family. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know what sliding-window and full-attention layers actually do differently inside one model, how KV cache sharing splits a… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e4b-it-atlas.gemma4-e4b-rl100-hf-bf16-sdpa-topk128-overlay
Gemma 4 E4B RL100 top-k-128 target overlay
Precomputed off-policy distillation targets for the E4B-RL-step-100 to E2B experiment.
Source traces: JWei05/gemma4-e4b-rl100-topk128-traces at revision 2b6e49a0a456ee9d67b16a1dc61785562bee90c9
Direction: Gemma 4 E4B RL step 100 teacher to Gemma 4 E2B base student
Target engine: Hugging Face BF16 SDPA full forward
Width: top-k 128
Stored target token IDs: int32
Stored target log-probabilities: float16
Causal alignment: response token… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e4b-rl100-hf-bf16-sdpa-topk128-overlay.gemma4-e2b-base-topk128-hf-overlay-v128-seed42
Gemma 4 E2B base top-k-128 HF training overlay
This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into
Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from
JWei05/gemma4-e2b-base-topk128-traces,
but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training
engine.
This repository is a reproducibility artifact for the corresponding distillation run. It is not a
new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.2026_08_20_refinement_math_chess_gemma3_12b_gemma4_31b_transition_feedback_tok2026_08_05_refinement_5env_gemma3_12b_gemma4_31b_tok2026_08_12_refinement_math_chess_gemma3_12b_gemma4_31b_raw_student_tok2026_08_11_refinement_5env_gemma3_12b_gemma4_31b_raw_student_tok2026_08_09_refinement_5env_gemma3_12b_gemma4_31b_flsft_tokcot-gemma4-26b-a4b
Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus
Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE,
25.2B total / 3.8B active), in its native thinking mode, across a diverse suite
of reasoning tasks. Structure follows
ceselder/cot-oracle-corpus-v5
(CoT-only subset of the columns), built for chain-of-thought monitoring /
activation-oracle research.
2,121,354 rollouts over 212,161 unique problems (10 sampled
thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.synthweb-gemma4-26b-a4b
Gemma-4-26B-A4B FineWeb Rollouts (~580k docs)
Open-ended continuations of FineWeb
(sample-10BT) document prefixes, generated by google/gemma-4-26b-a4b (the base, non-it
Gemma-4 26B-A4B mixture-of-experts model), then mode-collapse filtered. This is the Gemma-4
analogue of cds-jb/qwen3-8b-fineweb-rollouts-100k:
a "synthweb" corpus of natural model-generated documents, intended as the substrate for
activation-oracle / interpretability probing (extract a base model's residual… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthweb-gemma4-26b-a4b.2026_07_19_collect_leandojo_gemma3_12b_gemma4_31b_flsft_tok2026_07_19_collect_leandojo_gemma3_12b_gemma4_31b_raw_student_tok2026_07_29_collect_mathnet_gemma3_12b_gemma4_31b_flsft_tokgemma-4-e4b-it-atlas-SAE
juiceb0xc0de/gemma-4-e4b-it-atlas-sae
This joins the SAE half of the E4B atlas to the census half. Built from
juiceb0xc0de/gemma-4-e4b-it-SAE-v2
(42 layers, 32x, d_sae 81,920) against
juiceb0xc0de/gemma-4-e4b-it-atlas.
Three splits, all queryable in the browser. No download needed to look around.
The join
The atlas and the SAEs describe the same model in two index spaces that could not talk to each other:
atlas features.feature_idx in [0, 10240) an MLP hidden… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e4b-it-atlas-SAE.2026_07_29_collect_mathnet_gemma3_12b_gemma4_31b_transition_feedback_tok2026_07_20_collect_lichess_gemma3_12b_gemma4_31b_raw_student_no_feedback_tok2026_07_29_collect_mathnet_gemma3_12b_gemma4_31b_raw_student_no_feedback_tok2026_07_20_collect_lichess_gemma3_12b_gemma4_31b_transition_feedback_tok2026_07_20_collect_codeforces_gemma3_12b_gemma4_31b_raw_student_tokTeleQnA-router-gemma4-e4b
TeleQnA router data, Gemma4-E4B
Training data for a quality-estimation classifier over Gemma4-E4B answers on
TeleQnA. Each row is one question on one run: the model's own answer text and a
label saying whether that answer was correct.
Companion to ymoslem/TeleQnA-router,
which holds the same thing for Qwen3-4B-Instruct. The schema is identical, so
the same training script works on either by changing the dataset name.
Split
Rows
Questions
Runs
Accept rate… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/TeleQnA-router-gemma4-e4b.2026_07_29_collect_mathnet_gemma3_12b_gemma4_31b_tok2026_08_26_omni_math_train_feedback_adherence_gemma3_12b_gemma4_31b_candidates
Omni-MATH train feedback-adherence candidates
Production candidate data for studying whether a student follows teacher feedback.
Student: google/gemma-3-12b-it
Teacher and adherence judge: google/gemma-4-31B-it
Source problems: LLParallax/Omni-MATH-filtered, train partition after a fixed 512-problem test split
Source trajectories: LLParallax/2026_07_16_collect_omni_math_gemma3_12b_gemma4_31b
Collection config:… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/2026_08_26_omni_math_train_feedback_adherence_gemma3_12b_gemma4_31b_candidates.2026_07_16_collect_omni_math_gemma3_12b_gemma4_31b_tok2026_07_16_collect_omni_math_gemma3_12b_gemma4_31b_transition_feedback_tok2026_08_17_refinement_math_chess_gemma3_12b_gemma4_31b_raw_student_no_feedback_tokocn-empty-negations-generations-main-gemma4-qwen35
OCN OSS Model Generations
This dataset contains open-source model generations for prompts designed to elicit or suppress contrastive-negation framing.
Columns
prompt metadata from the OCN prompt bank;
model_id: Hugging Face model id;
model_family: model family;
model_stage: base, instruct, or other;
decoding: decoding configuration name;
seed: generation seed;
response: generated answer;
created_at: notebook run timestamp.
experiment_id: experiment cohort… See the full description on the dataset page: https://huggingface.co/datasets/ritwikraha/ocn-empty-negations-generations-main-gemma4-qwen35.2026_07_16_collect_omni_math_gemma3_12b_gemma4_31b_raw_student_tok2026_07_16_collect_omni_math_gemma3_12b_gemma4_31b_flsft_tok
