datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma-4-e4b-kinetics_54K
Gemma-4 Kinetics 54K Video Caption Data
What: 54,618 cleaned Kinetics-600 video-caption records (75 action labels) in multimodal chat JSON, for video-VLM supervised fine-tuning.
Splits: train 43,696 / validation 5,461 / test 5,461 (80/10/10, stratified per label, seed 42, zero video overlap across splits).
Two prompt variants: annotations/splits-MQ/ (recommended) randomly combines 3 system × 5 user prompts per record to prevent prompt overfitting and format collapse;… See the full description on the dataset page: https://huggingface.co/datasets/bear7011/gemma-4-e4b-kinetics_54K.ConvAI2-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.ConvAI2-Qwen-original-gemma-4-E4B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-E4B-it.ConvAI2-ERNIE-original-gemma-4-E4B-it
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-gemma-4-E4B-it.ConvAI2-FLUX-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E4B-it.PersonaChat-FLUX-enhanced-gemma-4-E4B-it
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E4B-it.ConvAI2-Qwen-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E4B-it.PersonaChat-Qwen-original-gemma-4-E4B-it
Visual Memory Results: personachat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-original-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-Qwen-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-original-gemma-4-E4B-it.PersonaChat-ERNIE-original-gemma-4-E4B-it
Visual Memory Results: personachat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-original-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-ERNIE-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-original-gemma-4-E4B-it.ConvAI2-FLUX-original-gemma-4-E4B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-E4B-it.Synthetic-Persona-Chat-FLUX-original-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E4B-it.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-ERNIE-original-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E4B-it.PersonaChat-Qwen-enhanced-gemma-4-E4B-it
Visual Memory Results: personachat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-Qwen-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-Qwen-original-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-original-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-original-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-original-gemma-4-E4B-it.PersonaChat-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: personachat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E4B-it.PersonaChat-FLUX-original-gemma-4-E4B-it
Visual Memory Results: personachat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-original-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-FLUX-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-original-gemma-4-E4B-it.eh-gemma4-e4b-kv-seam-quarantine
gemma4-e4b-kv-seam-quarantine -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-gemma4-e4b-kv-seam-quarantine
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-gemma4-e4b-kv-seam-quarantine.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.gemma-4-e4b-kinetics_330K
This datset compose of 295,612 training and 32,845 validation Kinetics-600 video-caption pairs across 479 action labels.
Please unzip the file first
gemma-4-e4b-it-ask-dataset
ask training and evaluation dataset
Conversational SFT data for ask. The explicit train split has 150 examples and evaluate has 27. Each record contains messages and tools in TRL tool-calling format.
gemma-4-E4B-it-MathVision-benchmarkBenchmark of google/gemma-4-E4B-it against MathLLMs/MathVision dataset.
Accuracy: 49.2% with Python tool.
Metric
Value
Correct
754
Incorrect
776
Errors
2
Total samples
1532
Python tool calls
7
Python tool errors
0
Total completion tokens
4,188,239
Raw stats:
{
"accuracy": 0.492,
"correct": 754,
"incorrect": 776,
"error": 2,
"total": 1532,
"python_tool_calls": 7,
"python_tool_errors":0,
"completion_tokens": 4188239
}
gemma-4-E4B-it-MedXpertQA-benchmarkBenchmark of google/gemma-4-E4B-it against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split.
Accuracy: 19.0%.
Metric
Value
Correct
465
Incorrect
1985
Errors
0
Total samples
2450
Total completion tokens
3,044,553
Raw stats:
{
"accuracy": 0.19,"correct": 465,
"incorrect": 1985,
"error": 0,
"total": 2450,
"completion_tokens": 3044553
}
gemma-4-E4B-it-imo-answerbench-benchmarkBenchmark of google/gemma-4-E4B-it against Hwilner/imo-answerbench dataset.
Accuracy: 32.5% with Python tool.
Metric
Value
Correct
130
Incorrect
270
Errors
0
Total samples
400
Python tool calls
447
Python tool errors
21
Total completion tokens
2,429,217
Raw stats:
{
"accuracy": 0.325,
"correct": 130,
"incorrect": 270,
"error": 0,
"total": 400,
"python_tool_calls": 447,
"python_tool_errors":21,
"completion_tokens": 2429217
}
gemma-4-E4B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E4B-it against m-a-p/SuperGPQA dataset.
Accuracy: 38.1% with Python tool.
Metric
Value
Correct
761
Incorrect
1239
Errors
0
Total samples
2000
Python tool calls
200
Python tool errors
10
Total completion tokens
4,253,773
Raw stats:
{
"accuracy": 0.381,
"correct": 761,
"incorrect": 1239,
"error": 0,
"total": 2000,
"python_tool_calls": 200,
"python_tool_errors": 10,
"completion_tokens": 4253773
}
gemma-4-e4b-kinetics-qa-subset
QA question
Does anyone fall in the video?
Requirement
Please download the corresponded videos at bear7011/gemma-4-e4b-kinetics_54K.
Dataset Structure
Split
File
Records
Share
Train
train.json
13,107
80%
Validation
val.json
1,637
10%
Test
test.json
1,637
10%
Summary
summary.json
-
-
Gemma-4-E4B-it-SSD
SSD Dataset Replication (Gemma-4-E4B-it)
This dataset is a replication of the "Embarrassingly Simple Self-Distillation Improves Code Generation" (SSD) paper (arXiv:2604.01193).
Overview
The dataset contains coding problems and their corresponding solutions generated by Gemma-4-E4B-it using high-temperature sampling (T=1.1) to explore the model's latent capabilities. This approach, known as SSD, focuses on "self-distillation" where a model's own correct but non-greedy… See the full description on the dataset page: https://huggingface.co/datasets/wrmedford/Gemma-4-E4B-it-SSD.gemma-4-E4B-it-MMLU-Pro-benchmarkBenchmark of google/gemma-4-E4B-it against TIGER-Lab/MMLU-Pro dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
1383
Incorrect
617
Errors
0
Total samples
2000
Python tool calls
235
Python tool errors
11
Total completion tokens
3,328,419
Raw stats:
{
"accuracy": 0.692,
"correct": 1383,
"incorrect": 617,
"error": 0,
"total": 2000,
"python_tool_calls": 235,
"python_tool_errors":11,
"completion_tokens": 3328419
}
