datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ConvAI2-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.ConvAI2-Qwen-original-gemma-4-E4B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-E4B-it.ConvAI2-ERNIE-original-gemma-4-E4B-it
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-gemma-4-E4B-it.ConvAI2-FLUX-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E4B-it.PersonaChat-FLUX-enhanced-gemma-4-E4B-it
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E4B-it.ConvAI2-Qwen-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E4B-it.PersonaChat-Qwen-original-gemma-4-E4B-it
Visual Memory Results: personachat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-original-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-Qwen-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-original-gemma-4-E4B-it.PersonaChat-ERNIE-original-gemma-4-E4B-it
Visual Memory Results: personachat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-original-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-ERNIE-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-original-gemma-4-E4B-it.ConvAI2-FLUX-original-gemma-4-E4B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-E4B-it.Synthetic-Persona-Chat-FLUX-original-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E4B-it.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-ERNIE-original-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E4B-it.PersonaChat-Qwen-enhanced-gemma-4-E4B-it
Visual Memory Results: personachat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-Qwen-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-Qwen-original-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-original-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-original-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-original-gemma-4-E4B-it.PersonaChat-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: personachat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E4B-it
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E4B-it.PersonaChat-FLUX-original-gemma-4-E4B-it
Visual Memory Results: personachat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-original-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-FLUX-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-original-gemma-4-E4B-it.eh-gemma4-e4b-kv-seam-quarantine
gemma4-e4b-kv-seam-quarantine -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-gemma4-e4b-kv-seam-quarantine
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-gemma4-e4b-kv-seam-quarantine.gemma-4-E4B-it-MathVision-benchmarkBenchmark of google/gemma-4-E4B-it against MathLLMs/MathVision dataset.
Accuracy: 49.2% with Python tool.
Metric
Value
Correct
754
Incorrect
776
Errors
2
Total samples
1532
Python tool calls
7
Python tool errors
0
Total completion tokens
4,188,239
Raw stats:
{
"accuracy": 0.492,
"correct": 754,
"incorrect": 776,
"error": 2,
"total": 1532,
"python_tool_calls": 7,
"python_tool_errors":0,
"completion_tokens": 4188239
}
gemma-4-E4B-it-MedXpertQA-benchmarkBenchmark of google/gemma-4-E4B-it against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split.
Accuracy: 19.0%.
Metric
Value
Correct
465
Incorrect
1985
Errors
0
Total samples
2450
Total completion tokens
3,044,553
Raw stats:
{
"accuracy": 0.19,"correct": 465,
"incorrect": 1985,
"error": 0,
"total": 2450,
"completion_tokens": 3044553
}
gemma-4-E4B-it-imo-answerbench-benchmarkBenchmark of google/gemma-4-E4B-it against Hwilner/imo-answerbench dataset.
Accuracy: 32.5% with Python tool.
Metric
Value
Correct
130
Incorrect
270
Errors
0
Total samples
400
Python tool calls
447
Python tool errors
21
Total completion tokens
2,429,217
Raw stats:
{
"accuracy": 0.325,
"correct": 130,
"incorrect": 270,
"error": 0,
"total": 400,
"python_tool_calls": 447,
"python_tool_errors":21,
"completion_tokens": 2429217
}
gemma-4-E4B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E4B-it against m-a-p/SuperGPQA dataset.
Accuracy: 38.1% with Python tool.
Metric
Value
Correct
761
Incorrect
1239
Errors
0
Total samples
2000
Python tool calls
200
Python tool errors
10
Total completion tokens
4,253,773
Raw stats:
{
"accuracy": 0.381,
"correct": 761,
"incorrect": 1239,
"error": 0,
"total": 2000,
"python_tool_calls": 200,
"python_tool_errors": 10,
"completion_tokens": 4253773
}
gemma-4-E4B-it-MMLU-Pro-benchmarkBenchmark of google/gemma-4-E4B-it against TIGER-Lab/MMLU-Pro dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
1383
Incorrect
617
Errors
0
Total samples
2000
Python tool calls
235
Python tool errors
11
Total completion tokens
3,328,419
Raw stats:
{
"accuracy": 0.692,
"correct": 1383,
"incorrect": 617,
"error": 0,
"total": 2000,
"python_tool_calls": 235,
"python_tool_errors":11,
"completion_tokens": 3328419
}
gemma-4-E4B-it-Health_Benchmarks-benchmarkBenchmark of google/gemma-4-E4B-it against yesilhealth/Health_Benchmarks dataset.
Accuracy: 77.8%.
Metric
Value
Correct
5864
Incorrect
1669
Errors
2
Total samples
7535
Total completion tokens
8,144,545
Raw stats:
{
"accuracy": 0.778,
"correct": 5864,
"incorrect": 1669,
"error": 2,
"total": 7535,
"completion_tokens": 8144545
}
gemma-4-E4B-it-GPQA-Diamond-benchmarkBenchmark of google/gemma-4-E4B-it against fingertap/GPQA-Diamond dataset. Results are averaged over 4 runs to reduce variance.
Accuracy: 54.0% with Python tool.
Metric
Value
Correct
428
Incorrect
364
Errors
0
Total samples
792
Python tool calls
54
Python tool errors
2
Total completion tokens
1,951,097
Raw stats:
{
"accuracy": 0.54,
"correct": 428,
"incorrect": 364,
"error": 0,
"total": 792,
"python_tool_calls": 54,
"python_tool_errors": 2… See the full description on the dataset page: https://huggingface.co/datasets/kth8/gemma-4-E4B-it-GPQA-Diamond-benchmark.gemma-4-E4B-it-ValleyBench-benchmarkBenchmark of google/gemma-4-E4B-it against kth8/ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 80.3% with Python tool.
Metric
Value
Correct
4014
Incorrect
958
Errors
28
Total samples
5000
Python tool calls
4843
Total completion tokens
4,595,312
Raw stats:
{
"accuracy": 0.803,
"correct": 4014,
"incorrect": 958,
"error": 28,
"total": 5000,
"python_tool_calls": 4843,
"completion_tokens": 4595312
}
