CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01juiceb0xc0de /gemma-4-e2b-atlas image1M<n<10M4 likes810 downloads10d agoHugging Face02JWei05 /gemma4-e2b-base-topk128-hf-overlay-v128-seed42 Gemma 4 E2B base top-k-128 HF training overlay This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from JWei05/gemma4-e2b-base-topk128-traces, but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training engine. This repository is a reproducibility artifact for the corresponding distillation run. It is not a new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.tabulartext-generation10K<n<100K0 likes308 downloads2mo agoHugging Face03juiceb0xc0de /gemma-4-e2b-SAE-sqlite Gemma 4 E2B SAE SQLite Atlas An exact, queryable SQLite representation of all 35 residual-stream sparse autoencoders from juiceb0xc0de/gemma-4-e2b-it-SAE. The database contains 1,720,320 feature rows. Encoder and decoder vectors preserve the source checkpoints' float32 values exactly. Files gemma-4-e2b-sae.sqlite3 — SQLite database (20.63 GiB) manifest.json — source revision, dimensions, SHA-256, and integrity result Database SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e2b-SAE-sqlite.text1M<n<10M0 likes69 downloads27d agoHugging Face04DuoNeural /Gemma4-E2B-SFT-WebCode Gemma4-E2B-SFT-WebCode Synthetic frontend web development dataset. Natural language component description → production-ready code. Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS. Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc. Format: ShareGPT/ChatML. Includes accessibility attributes and comments. Use: Fine-tune models for frontend copilot tasks. Generator: DuoNeural/TurboGemma4E2B, temperature 0.65. text1K<n<10K0 likes51 downloads5mo agoHugging Face05Solshine /nla-gemma4e2b-relabel-v1-eval Gemma-4-E2B layer-23 evaluation set, relabeled (v1), with contamination flags The 580-document evaluation pool on which every activation-verbalizer result in this project is scored, with each row's evaluation text rewritten from a topic summary to a feature-attribution label. The activations are byte-identical to the original evaluation set; only the text column changed, and the original text is preserved. This pool is not disjoint from the training corpus. Read this… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-eval.texttext-generationn<1K0 likes42 downloads7d agoHugging Face06leobianco /eval_bosch_gemma-4-E2B-it_gens_T0_wfs2text1K<n<10K0 likes41 downloads11d agoHugging Face07Solshine /nla-gemma4e2b-relabel-v1-corpus Gemma-4-E2B layer-23 activation corpus, relabeled (v1) 1356 training rows for an activation verbalizer. Each row pairs a residual-stream activation captured at layer 23 of google/gemma-4-E2B with a natural-language label describing what the model must have integrated at that position to predict its next token. This is the training set behind Solshine/gemma-4-e2b-nla-L23-av-priordev-relabel-v1-wd3. Why it exists An audit of the previous version of this corpus found… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-corpus.tabulartext-generation1K<n<10K0 likes37 downloads7d agoHugging Face08Mauricette /Gemma4-E2B-SFT-WebCode Gemma4-E2B-SFT-WebCode Synthetic frontend web development dataset. Natural language component description → production-ready code. Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS. Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc. Format: ShareGPT/ChatML. Includes accessibility attributes and comments. Use: Fine-tune models for frontend copilot tasks. Generator: DuoNeural/TurboGemma4E2B, temperature 0.65. text1K<n<10K0 likes36 downloads6d agoHugging Face09mags0ft /Gemma-4-E2B-SSFT Gemma-4-E2B-SSFT This is a dataset that has been created using SSFT (simple-sft), a synthetic data generation tool written by me. It contains a few hundred samples for testing. Try it out yourself! texttext-generationn<1K1 likes35 downloads2mo agoHugging Face10pameydorke /redred-gemma-4-E2B-it-lora-summariestextn<1K0 likes34 downloads4d agoHugging Face11Solshine /nla-gemma4e2b-activation-labels Gemma-4-E2B NLA activation-label corpus Work in progress, part of ongoing research. Released for replicability ahead of a likely future write-up. Structure may change. A labeled dataset of language-model activations paired with short natural-language descriptions of what each activation represents. Each row is one 1536-dimensional residual-stream activation from layer 23 of google/gemma-4-E2B, plus a content-specific label. It is the training and evaluation data behind the… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-activation-labels.text1K<n<10K0 likes28 downloads3mo agoHugging Face12leobianco /eval_bosch_gemma-4-E2B-it_gens_T0_wfs0text1K<n<10K0 likes28 downloads11d agoHugging Face13Solshine /gemma-4-e2b-deception-behavior-completions Gemma-4-E2B deception & behavior completions Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included. The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.tabulartext-generationn<1K0 likes24 downloads5mo agoHugging Face14DuoNeural /Gemma4-E2B-SFT-SQL Gemma4-E2B-SFT-SQL Synthetic text-to-SQL dataset covering real-world database schemas. Schemas: E-commerce, healthcare, SaaS analytics. Query types: Joins, subqueries, aggregations, window functions, CTEs. Format: ShareGPT/ChatML. Natural language question + SQL answer + brief explanation. Use: Fine-tune models for autonomous database querying and agentic SQL generation. Generator: DuoNeural/TurboGemma4E2B, temperature 0.4. text1K<n<10K0 likes23 downloads5mo agoHugging Face15Solshine /gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit Gemma-4-E2B NLA AV-SFT Training Corpus (v0.1.x, Gemini persona+audit) The 4,734-row AV-SFT training corpus for the v0.1.x Gemma-4-E2B NLA — a 9-source-family diversified expansion over the v0.0.x OpenWebText-only corpus. Labels generated by Gemini CLI following the persona+audit pipeline (Dr. Marisol Chen labels, Dr. Riley Otsuka audits). This is the in-progress v0.1.x labeled training set. AR-SFT companion is still being labeled (~16% complete as of this dataset publish). When the… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit.tabulartext-generation1K<n<10K0 likes23 downloads5mo agoHugging Face16Solshine /gemma-4-e2b-nla-eval-smoke Gemma-4-E2B NLA smoke-eval (20-row held-out set) A 20-row held-out subset of OpenWebText activations extracted from google/gemma-4-E2B at layer 23. Used as the canonical eval set for smoke-testing the v0.0.1 Gemma-4-E2B NLA pair on a fresh environment. This dataset is a subset of the held-out rl.parquet evaluation set used for the v0.0.1 round-trip eval (n=50 attempted, 42 evaluated after 8 empty-output exclusions, cos 0.438 ± 0.054). The 20-row subset preserves the activation… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-eval-smoke.tabulartext-generationn<1K0 likes22 downloads5mo agoHugging Face17Solshine /gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit Gemma-4-E2B NLA AR-SFT Training Corpus (v0.0.x, Claude Haiku persona+audit) The 696-row AR-SFT training corpus used for the Option B Gemma-4-E2B NLA pair. Labels generated by Claude Haiku 4.5 following the persona+audit pipeline — Dr. Marisol Chen (synthetic mech-interp expert) labels first, Dr. Riley Otsuka (synthetic senior editor) audits the labels. This is the matched companion to the v0.0.x AV labeled corpus. The pair completes the first open-source non-Anthropic-team NLA… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit.tabulartext-generationn<1K0 likes19 downloads5mo agoHugging Face18DuoNeural /Gemma4-E2B-SFT-CoT Gemma4-E2B-SFT-CoT Synthetic chain-of-thought reasoning dataset generated by DuoNeural/TurboGemma4E2B (Gemma 4 E2B abliterated). Generation: 2-pass synthesis + self-evaluation (FineWeb-Edu style), only examples scoring ≥4/5 retained. Topics: Math, logic, physics, probability, algorithm analysis, number theory. Format: ShareGPT/ChatML (messages column, user+assistant turns). Use: Fine-tune small models for step-by-step reasoning capabilities. Generator: DuoNeural/TurboGemma4E2B with… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/Gemma4-E2B-SFT-CoT.text1K<n<10K0 likes18 downloads5mo agoHugging Face19DuoNeural /Gemma4-E2B-SFT-JSON Gemma4-E2B-SFT-JSON Synthetic structured JSON entity extraction dataset. Model receives unstructured document → outputs strictly valid JSON. Domains: Medical (clinical notes), legal (contracts), financial (earnings reports), job postings, research papers. Format: ShareGPT/ChatML. Two-step generation: document synthesized first, then extracted. Use: Fine-tune models for robust information extraction and structured output generation. Generator: DuoNeural/TurboGemma4E2B, temperature… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/Gemma4-E2B-SFT-JSON.text1K<n<10K0 likes18 downloads5mo agoHugging Face20tepirale /gemma4_E2B_grpo_v1 dataset original https://huggingface.co/datasets/mlabonne/orpo-dpo-mix-40k 'Chosen' column of text converted to image for GRPO training or similar Made by https://huggingface.co/NickyNicky imageimage-text-to-text10K<n<100K0 likes15 downloads3mo agoHugging Face21Solshine /gemma-4-e2b-nla-av_sft-v0_1_x-short-hybrid-labels ⚠ CORRECTED 2026-05-16 — label format mismatch flagged This corpus's label format does NOT match Anthropic's NLA methodology. Median response length is 6 words with ≤7-word "explanation" tags (e.g., "<explanation> Galileo refuting Aristotle's gravity </explanation>"). Anthropic's NLA training data uses multi-paragraph ~80–120-word explanations with bolded topic headings (see https://huggingface.co/kitft/Llama-3.3-70B-NLA-L53-av for an example of the expected label format). This… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-av_sft-v0_1_x-short-hybrid-labels.tabular1K<n<10K0 likes12 downloads4mo agoHugging Face22cloudcastnepal-ai-labs /gemma4-e2b-generated-instructions-demo-v1 Unsloth Dataset Workflow Test Overview This dataset is a workflow validation dataset generated using Unsloth Studio. It demonstrates the complete pipeline: Source dataset AI-generated instructions Export to Parquet Upload to Hugging Face Dataset viewer validation This repository is intended for testing the publication workflow before creating a larger production-quality dataset. Dataset Structure Columns output generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.texttext-generationn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.