CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JWei05 /gemma4-e2b-base-topk128-hf-overlay-v128-seed42 Gemma 4 E2B base top-k-128 HF training overlay This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from JWei05/gemma4-e2b-base-topk128-traces, but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training engine. This repository is a reproducibility artifact for the corresponding distillation run. It is not a new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.tabulartext-generation10K<n<100K0 likes308 downloads2mo agoHugging Face02Solshine /nla-gemma4e2b-relabel-v1-eval Gemma-4-E2B layer-23 evaluation set, relabeled (v1), with contamination flags The 580-document evaluation pool on which every activation-verbalizer result in this project is scored, with each row's evaluation text rewritten from a topic summary to a feature-attribution label. The activations are byte-identical to the original evaluation set; only the text column changed, and the original text is preserved. This pool is not disjoint from the training corpus. Read this… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-eval.texttext-generationn<1K0 likes42 downloads6d agoHugging Face03esherialabs /saferide-gemma-4-e2b-v058-original-419806-training-data SafeRide Synthetic Bilingual Safety Guidance Dataset v0.5.8 This research and development dataset contains synthetic English and Kiswahili chat conversations. It was designed to help a language model practice cautious, agency-preserving safety guidance, useful refusal behavior, and responses that avoid inventing facts. It contains no real survivor reports or production records. The frozen dataset is publicly available under Creative Commons Attribution 4.0 International (CC BY… See the full description on the dataset page: https://huggingface.co/datasets/esherialabs/saferide-gemma-4-e2b-v058-original-419806-training-data.texttext-generation1K<n<10K0 likes41 downloads1mo agoHugging Face04Solshine /nla-gemma4e2b-relabel-v1-corpus Gemma-4-E2B layer-23 activation corpus, relabeled (v1) 1356 training rows for an activation verbalizer. Each row pairs a residual-stream activation captured at layer 23 of google/gemma-4-E2B with a natural-language label describing what the model must have integrated at that position to predict its next token. This is the training set behind Solshine/gemma-4-e2b-nla-L23-av-priordev-relabel-v1-wd3. Why it exists An audit of the previous version of this corpus found… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-corpus.tabulartext-generation1K<n<10K0 likes37 downloads6d agoHugging Face05mags0ft /Gemma-4-E2B-SSFT Gemma-4-E2B-SSFT This is a dataset that has been created using SSFT (simple-sft), a synthetic data generation tool written by me. It contains a few hundred samples for testing. Try it out yourself! texttext-generationn<1K1 likes35 downloads2mo agoHugging Face06Solshine /gemma-4-e2b-deception-behavior-completions Gemma-4-E2B deception & behavior completions Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included. The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.tabulartext-generationn<1K0 likes24 downloads5mo agoHugging Face07Solshine /gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit Gemma-4-E2B NLA AV-SFT Training Corpus (v0.1.x, Gemini persona+audit) The 4,734-row AV-SFT training corpus for the v0.1.x Gemma-4-E2B NLA — a 9-source-family diversified expansion over the v0.0.x OpenWebText-only corpus. Labels generated by Gemini CLI following the persona+audit pipeline (Dr. Marisol Chen labels, Dr. Riley Otsuka audits). This is the in-progress v0.1.x labeled training set. AR-SFT companion is still being labeled (~16% complete as of this dataset publish). When the… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit.tabulartext-generation1K<n<10K0 likes23 downloads5mo agoHugging Face08Solshine /gemma-4-e2b-nla-eval-smoke Gemma-4-E2B NLA smoke-eval (20-row held-out set) A 20-row held-out subset of OpenWebText activations extracted from google/gemma-4-E2B at layer 23. Used as the canonical eval set for smoke-testing the v0.0.1 Gemma-4-E2B NLA pair on a fresh environment. This dataset is a subset of the held-out rl.parquet evaluation set used for the v0.0.1 round-trip eval (n=50 attempted, 42 evaluated after 8 empty-output exclusions, cos 0.438 ± 0.054). The 20-row subset preserves the activation… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-eval-smoke.tabulartext-generationn<1K0 likes22 downloads5mo agoHugging Face09Solshine /gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit Gemma-4-E2B NLA AR-SFT Training Corpus (v0.0.x, Claude Haiku persona+audit) The 696-row AR-SFT training corpus used for the Option B Gemma-4-E2B NLA pair. Labels generated by Claude Haiku 4.5 following the persona+audit pipeline — Dr. Marisol Chen (synthetic mech-interp expert) labels first, Dr. Riley Otsuka (synthetic senior editor) audits the labels. This is the matched companion to the v0.0.x AV labeled corpus. The pair completes the first open-source non-Anthropic-team NLA… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit.tabulartext-generationn<1K0 likes19 downloads5mo agoHugging Face10cloudcastnepal-ai-labs /gemma4-e2b-generated-instructions-demo-v1 Unsloth Dataset Workflow Test Overview This dataset is a workflow validation dataset generated using Unsloth Studio. It demonstrates the complete pipeline: Source dataset AI-generated instructions Export to Parquet Upload to Hugging Face Dataset viewer validation This repository is intended for testing the publication workflow before creating a larger production-quality dataset. Dataset Structure Columns output generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.texttext-generationn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.