CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mechanistic-Anomaly-Detection /gemma2-jailbreakstext10K<n<100K2 likes6.5k downloads2y agoHugging Face02haoranli-ml /prolong-data-64K-gemma0 likes6.4k downloads5mo agoHugging Face03abotresol /emotion-vectors-gemma-4-31b-it-postfix Emotion vectors, google/gemma-4-31b-it (corrected extraction) Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.feature-extraction0 likes5.6k downloads2mo agoHugging Face04vaghawan /hausa_response_gemmatext1K<n<10K0 likes2.6k downloads27d agoHugging Face05AbhinandT /classified_images_gemmaimagen<1K0 likes2.5k downloads1y agoHugging Face06abotresol /emotion-vectors-gemma-4-31b-it Emotion vectors — gemma-4-31b-it (instruct) probed on the external gemma-4-4B story corpus Data provenance (what made these activations) Probed model: google/gemma-4-31b-it (instruct) Input corpus: snae/emotion_stories_gemma_4_4B — stories written by gemma-4-4B, a smaller EXTERNAL model (generator is NOT the probed model) Per-story pooled residual-stream activations and per-emotion mean vectors, extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it.feature-extraction0 likes2k downloads2mo agoHugging Face07ReasoningMila /llama3.1_8b_inst_as_ver_gemma27b_it_math158_32gen_async0 likes2k downloads2y agoHugging Face08abotresol /emotion-vectors-gemma-4-31b Emotion vectors — gemma-4-31b (base) probed on the external gemma-4-4B story corpus Data provenance (what made these activations) Probed model (whose activations these are): google/gemma-4-31b (base) Input corpus: snae/emotion_stories_gemma_4_4B — third-person emotion stories written by gemma-4-4B, a smaller EXTERNAL model (the open replication's published corpus; generator is NOT the probed model) Per-story pooled residual-stream activations and per-emotion… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b.feature-extraction0 likes2k downloads2mo agoHugging Face09sida /1M_activations_pile_10k_GPT_Gemma_Qwen0 likes2k downloads3mo agoHugging Face10kisate-team /gemma-2b-suite-explanationstext1M<n<10M0 likes2k downloads2y agoHugging Face11abotresol /emotion-vectors-gemma-4-31b-postfix Emotion vectors, google/gemma-4-31b (corrected extraction) Residual-stream activations for google/gemma-4-31b, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-vectors-gemma-4-31b. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set re-extracts… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-postfix.feature-extraction0 likes1.8k downloads2mo agoHugging Face12nileshsarkar-ai /gemma-crafter-five-experiments-202609150 likes1.7k downloads4d agoHugging Face13SEBK4C /gemma4-serving-bench-data Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data Test data, charts, and the running research log from an autonomous research loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop summarizes findings, proposes a goal, tests it end-to-end, documents success or failure, and publishes here + to GitHub. Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.imagen<1K0 likes1.3k downloads3mo agoHugging Face14roo5150 /eagle3-hidden-states-gemma40 likes1.2k downloads5mo agoHugging Face15abotresol /emotion-selfstory-vectors-gemma-4-31b-it-postfix Emotion vectors, google/gemma-4-31b-it (corrected extraction) Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-selfstory-vectors-gemma-4-31b-it. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it-postfix.feature-extraction0 likes1.2k downloads2mo agoHugging Face16abotresol /emotion-combined-trajectories-gemma-4-31b-it-v2 Per-token emotion trajectories, instruction-tuned model (primary set) google/gemma-4-31b-it read over stories written to move through three emotions in sequence, scored against three different emotion-vector sets (corpus-built, self-generated and DeepSeek-written). This is the primary trajectory set behind the project's story-following results. Each story is stored as one .npz. The arrays are per token, so a trajectory can be replayed word by word rather than only summarised.… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b-it-v2.feature-extraction0 likes1.2k downloads2mo agoHugging Face17mlnomad /fineweb-edu-gemma4-1024 FineWeb-Edu — pre-tokenized for fast LM pretraining (Gemma tokenizer, ArrayRecord/Grain) Pre-tokenized FineWeb-Edu (sample/100BT), packed into fixed-length sequences and stored as ArrayRecord shards for zero-overhead streaming with Grain. No on-the-fly tokenization at train time — you read int32 tokens straight off disk. Format Tokenizer: google/gemma-4-12B-it (vocab size 262144). Documents are separated by the EOS token id 1. Packing: the token stream is… See the full description on the dataset page: https://huggingface.co/datasets/mlnomad/fineweb-edu-gemma4-1024.text-generation10B<n<100B0 likes1.2k downloads3mo agoHugging Face18credi-net /CDB_DEC2024-CochranSampled_Gemma-300m_Embtext100M<n<1B1 likes1.1k downloads2mo agoHugging Face19abotresol /emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it Per-token emotion trajectories, DeepSeek-written stories google/gemma-4-31b-it read over the DeepSeek-written three-emotion stories. Pairs with the Gemma-written set to separate what the model does from what the story writer does. Each story is stored as one .npz. The arrays are per token, so a trajectory can be replayed word by word rather than only summarised. Contents Path Contents shards/<story_id>.npz one story, arrays below manifest.jsonl one… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it.feature-extraction0 likes1.1k downloads2mo agoHugging Face20veriga /openwebtext-gemma3-tokenized-1024-activations-layer23 OpenWebText — Gemma-3-1B Hidden State Activations (Layer 23) Precomputed hidden state activations before layer 23 of Gemma-3-1B-IT for the OpenWebText dataset, tokenized with sequence length 1024. Designed for training a Titans memory layer that replaces layer 23 of Gemma 3. Dataset Structure Each example contains the inputs to layer 23: Field Shape Dtype Description activations (1024, 1152) float32 Hidden state activations (cast from bfloat16) mask(1024… See the full description on the dataset page: https://huggingface.co/datasets/veriga/openwebtext-gemma3-tokenized-1024-activations-layer23.timeseriesother10K<n<100K0 likes1.1k downloads4mo agoHugging Face21ketanpatil03 /rlhf-gemma3-indfood-1kimage1K<n<10K0 likes967 downloads2mo agoHugging Face22juiceb0xc0de /gemma-4-e2b-atlas image1M<n<10M4 likes839 downloads7d agoHugging Face23lamm-mit /gemma4-interpretability Gemma materials-science interpretability research archive Research records supporting Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model, Markus J. Buehler. Release identifier: paper-revision-2026-09-06. This archive supplies the original observations, supporting state arrays, exact prompts, protocols, intervention records, statistics, analysis source, and generated research figures. It includes the original 4B readout and geometry… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-interpretability.1 likes831 downloads15d agoHugging Face24abotresol /emotion-selfstory-vectors-gemma-4-31b-it Emotion vectors — gemma-4-31b-it probed on its OWN self-generated stories Data provenance (what made these activations) Probed model: google/gemma-4-31b-it (instruct) Input corpus: abotresol/emotion-stories-gemma-4-31b-it — stories written by the probed model itself (generator = probed model, the reference's convention; 12 the twelve emotions, up to 256 stories each — the E6 scale corpus) Per-story pooled residual-stream activations and per-emotion mean vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it.feature-extraction0 likes804 downloads2mo agoHugging Face25kalomaze /glm52-usersim-two-pass-gemma-audit-v1 GLM-5.2 Usersim Two-Pass Gemma Audit v1 This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k. The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length. Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.tabulartext-generation100K<n<1M4 likes781 downloads1mo agoHugging Face26lightonai /ms-marco-en-bge-gemma ms-marco-en-bge This dataset contains the MS MARCO dataset with negatives mined using ColBERT and then scored by bge-reranker-v2-gemma. It can be used to train a retrieval model using knowledge distillation, for example using PyLate. knowledge distillation To fine-tune a model using knowledge distillation loss we will need three distinct file: Datasetsfrom datasets import load_dataset train = load_dataset( "lightonai/ms-marco-en-gemma", "train"… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/ms-marco-en-bge-gemma.textfeature-extraction10M<n<100M13 likes776 downloads1y agoHugging Face27abotresol /emotion-dialogue-vectors-gemma-4-31b Emotion vectors — gemma-4-31b (base) probed on base-generated dialogues Data provenance (what made these activations) Probed model: google/gemma-4-31b (base) Input corpus: abotresol/emotion-dialogues-gemma-4-31b — two-person dialogues written by the base model (generator = probed model; 44% emotion-word leakage, documented) Per-story pooled residual-stream activations and per-emotion mean vectors, extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-dialogue-vectors-gemma-4-31b.feature-extraction0 likes768 downloads2mo agoHugging Face28model-organisms-for-real /kd-dataset-gemma-milsub-benignmix-hs3 Benign mixing completions — gemma milsub teachers on hs3-filtered The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students. One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's completions on a seeded 6,584-prompt subset of model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0, max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.texttext-generation1K<n<10K0 likes745 downloads20d agoHugging Face29model-organisms-for-real /gemma2_9b_it_taboo_wave_oracle_v1-training-data0 likes741 downloads3mo agoHugging Face30abotresol /emotion-combined-trajectories-gemma-4-31b emotion-combined-trajectories-gemma-4-31b BASE-model condition of the per-token trajectory stories (companion to emotion-combined-trajectories-gemma-4-31b-it, same corpus and layout). One shard per story from a teacher-forced forward pass through google/gemma-4-31b (base, bf16, right padding), residual captured at layers [6, 15, 24, 33, 42, 51]. Provenance caveats specific to this condition The stories were generated by the INSTRUCT model (base cannot chat-write… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b.0 likes728 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.