datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma-4-e2b-atlas
gemma4-e2b-base-topk128-hf-overlay-v128-seed42
Gemma 4 E2B base top-k-128 HF training overlay
This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into
Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from
JWei05/gemma4-e2b-base-topk128-traces,
but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training
engine.
This repository is a reproducibility artifact for the corresponding distillation run. It is not a
new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.gemma-4-e2b-SAE-sqlite
Gemma 4 E2B SAE SQLite Atlas
An exact, queryable SQLite representation of all 35 residual-stream
sparse autoencoders from
juiceb0xc0de/gemma-4-e2b-it-SAE.
The database contains 1,720,320 feature rows. Encoder and decoder
vectors preserve the source checkpoints' float32 values exactly.
Files
gemma-4-e2b-sae.sqlite3 — SQLite database (20.63 GiB)
manifest.json — source revision, dimensions, SHA-256, and integrity result
Database SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e2b-SAE-sqlite.Gemma4-E2B-SFT-WebCode
Gemma4-E2B-SFT-WebCode
Synthetic frontend web development dataset. Natural language component description → production-ready code.
Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS.
Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc.
Format: ShareGPT/ChatML. Includes accessibility attributes and comments.
Use: Fine-tune models for frontend copilot tasks.
Generator: DuoNeural/TurboGemma4E2B, temperature 0.65.
nla-gemma4e2b-relabel-v1-eval
Gemma-4-E2B layer-23 evaluation set, relabeled (v1), with contamination flags
The 580-document evaluation pool on which every activation-verbalizer result in this
project is scored, with each row's evaluation text rewritten from a topic summary to
a feature-attribution label. The activations are byte-identical to the original
evaluation set; only the text column changed, and the original text is preserved.
This pool is not disjoint from the training corpus. Read this… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-eval.eval_bosch_gemma-4-E2B-it_gens_T0_wfs2nla-gemma4e2b-relabel-v1-corpus
Gemma-4-E2B layer-23 activation corpus, relabeled (v1)
1356 training rows for an activation verbalizer. Each row pairs a residual-stream
activation captured at layer 23 of google/gemma-4-E2B with a natural-language label
describing what the model must have integrated at that position to predict its next
token. This is the training set behind
Solshine/gemma-4-e2b-nla-L23-av-priordev-relabel-v1-wd3.
Why it exists
An audit of the previous version of this corpus found… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-corpus.Gemma4-E2B-SFT-WebCode
Gemma4-E2B-SFT-WebCode
Synthetic frontend web development dataset. Natural language component description → production-ready code.
Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS.
Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc.
Format: ShareGPT/ChatML. Includes accessibility attributes and comments.
Use: Fine-tune models for frontend copilot tasks.
Generator: DuoNeural/TurboGemma4E2B, temperature 0.65.
Gemma-4-E2B-SSFT
Gemma-4-E2B-SSFT
This is a dataset that has been created using SSFT (simple-sft), a synthetic data generation tool written by me.
It contains a few hundred samples for testing.
Try it out yourself!
redred-gemma-4-E2B-it-lora-summariesnla-gemma4e2b-activation-labels
Gemma-4-E2B NLA activation-label corpus
Work in progress, part of ongoing research. Released for replicability ahead of a likely future write-up. Structure may change.
A labeled dataset of language-model activations paired with short natural-language descriptions of what each activation represents. Each row is one 1536-dimensional residual-stream activation from layer 23 of google/gemma-4-E2B, plus a content-specific label. It is the training and evaluation data behind the… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-activation-labels.eval_bosch_gemma-4-E2B-it_gens_T0_wfs0gemma-4-e2b-deception-behavior-completions
Gemma-4-E2B deception & behavior completions
Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included.
The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.Gemma4-E2B-SFT-SQL
Gemma4-E2B-SFT-SQL
Synthetic text-to-SQL dataset covering real-world database schemas.
Schemas: E-commerce, healthcare, SaaS analytics.
Query types: Joins, subqueries, aggregations, window functions, CTEs.
Format: ShareGPT/ChatML. Natural language question + SQL answer + brief explanation.
Use: Fine-tune models for autonomous database querying and agentic SQL generation.
Generator: DuoNeural/TurboGemma4E2B, temperature 0.4.
gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit
Gemma-4-E2B NLA AV-SFT Training Corpus (v0.1.x, Gemini persona+audit)
The 4,734-row AV-SFT training corpus for the v0.1.x Gemma-4-E2B NLA — a 9-source-family diversified expansion over the v0.0.x OpenWebText-only corpus. Labels generated by Gemini CLI following the persona+audit pipeline (Dr. Marisol Chen labels, Dr. Riley Otsuka audits).
This is the in-progress v0.1.x labeled training set. AR-SFT companion is still being labeled (~16% complete as of this dataset publish). When the… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit.gemma-4-e2b-nla-eval-smoke
Gemma-4-E2B NLA smoke-eval (20-row held-out set)
A 20-row held-out subset of OpenWebText activations extracted from google/gemma-4-E2B at layer 23. Used as the canonical eval set for smoke-testing the v0.0.1 Gemma-4-E2B NLA pair on a fresh environment.
This dataset is a subset of the held-out rl.parquet evaluation set used for the v0.0.1 round-trip eval (n=50 attempted, 42 evaluated after 8 empty-output exclusions, cos 0.438 ± 0.054). The 20-row subset preserves the activation… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-eval-smoke.gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit
Gemma-4-E2B NLA AR-SFT Training Corpus (v0.0.x, Claude Haiku persona+audit)
The 696-row AR-SFT training corpus used for the Option B Gemma-4-E2B NLA pair. Labels generated by Claude Haiku 4.5 following the persona+audit pipeline — Dr. Marisol Chen (synthetic mech-interp expert) labels first, Dr. Riley Otsuka (synthetic senior editor) audits the labels.
This is the matched companion to the v0.0.x AV labeled corpus. The pair completes the first open-source non-Anthropic-team NLA… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit.Gemma4-E2B-SFT-CoT
Gemma4-E2B-SFT-CoT
Synthetic chain-of-thought reasoning dataset generated by DuoNeural/TurboGemma4E2B (Gemma 4 E2B abliterated).
Generation: 2-pass synthesis + self-evaluation (FineWeb-Edu style), only examples scoring ≥4/5 retained.
Topics: Math, logic, physics, probability, algorithm analysis, number theory.
Format: ShareGPT/ChatML (messages column, user+assistant turns).
Use: Fine-tune small models for step-by-step reasoning capabilities.
Generator: DuoNeural/TurboGemma4E2B with… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/Gemma4-E2B-SFT-CoT.Gemma4-E2B-SFT-JSON
Gemma4-E2B-SFT-JSON
Synthetic structured JSON entity extraction dataset. Model receives unstructured document → outputs strictly valid JSON.
Domains: Medical (clinical notes), legal (contracts), financial (earnings reports), job postings, research papers.
Format: ShareGPT/ChatML. Two-step generation: document synthesized first, then extracted.
Use: Fine-tune models for robust information extraction and structured output generation.
Generator: DuoNeural/TurboGemma4E2B, temperature… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/Gemma4-E2B-SFT-JSON.gemma4_E2B_grpo_v1
dataset original https://huggingface.co/datasets/mlabonne/orpo-dpo-mix-40k
'Chosen' column of text converted to image for GRPO training or similar
Made by https://huggingface.co/NickyNicky
gemma-4-e2b-nla-av_sft-v0_1_x-short-hybrid-labels
⚠ CORRECTED 2026-05-16 — label format mismatch flagged
This corpus's label format does NOT match Anthropic's NLA methodology. Median response length is 6 words with ≤7-word "explanation" tags (e.g., "<explanation> Galileo refuting Aristotle's gravity </explanation>"). Anthropic's NLA training data uses multi-paragraph ~80–120-word explanations with bolded topic headings (see https://huggingface.co/kitft/Llama-3.3-70B-NLA-L53-av for an example of the expected label format).
This… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-av_sft-v0_1_x-short-hybrid-labels.gemma4-e2b-generated-instructions-demo-v1
Unsloth Dataset Workflow Test
Overview
This dataset is a workflow validation dataset generated using Unsloth Studio.
It demonstrates the complete pipeline:
Source dataset
AI-generated instructions
Export to Parquet
Upload to Hugging Face
Dataset viewer validation
This repository is intended for testing the publication workflow before creating a larger production-quality dataset.
Dataset Structure
Columns
output
generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.
