datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma-2b-dictionary-embeddings-all-layers
Gemma-2B Dictionary Embeddings - All Layers
This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers.
Dataset Structure
metadata.json: Contains dataset metadata (model info, dimensions, word count)
embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26)
Usage
import pickle
from huggingface_hub import hf_hub_download
# Download a specific layer
layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.metrics-outputs-gemma-2-2b-layer-06-token-cacheSynthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it.Synthetic-Persona-Chat-Qwen-original-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-original-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-original-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-original-gemma-4-E2B-it.ConvAI2-Qwen-original-gemma-4-E2B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-E2B-it.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E2B-it.anakin87__gemma-2b-orpo-details
Dataset Card for Evaluation run of anakin87/gemma-2b-orpo
Dataset automatically created during the evaluation run of model anakin87/gemma-2b-orpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/anakin87__gemma-2b-orpo-details.PersonaChat-FLUX-enhanced-gemma-4-E2B-it
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E2B-it.ConvAI2-ERNIE-original-gemma-4-E2B-it
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-ERNIE-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-gemma-4-E2B-it.metrics-outputs-gemma-2-2b-layer-24-token-cacheConvAI2-Qwen-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E2B-it.ConvAI2-FLUX-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E2B-it.ConvAI2-ERNIE-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it.Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E2B-it.Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E2B-it.jebish7__gemma-2-2b-it-details
Dataset Card for Evaluation run of jebish7/gemma-2-2b-it
Dataset automatically created during the evaluation run of model jebish7/gemma-2-2b-it
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jebish7__gemma-2-2b-it-details.PersonaChat-Qwen-enhanced-gemma-4-E2B-it
Visual Memory Results: personachat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/PersonaChat-Qwen-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-enhanced-gemma-4-E2B-it.PersonaChat-ERNIE-enhanced-gemma-4-E2B-it
Visual Memory Results: personachat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/PersonaChat-ERNIE-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-enhanced-gemma-4-E2B-it.Synthetic-Persona-Chat-ERNIE-original-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E2B-it.PersonaChat-Qwen-original-gemma-4-E2B-it
Visual Memory Results: personachat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-original-gemma-4-E2B-it",
"results_jsonl": "results/PersonaChat-Qwen-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-original-gemma-4-E2B-it.PersonaChat-FLUX-original-gemma-4-E2B-it
Visual Memory Results: personachat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-original-gemma-4-E2B-it",
"results_jsonl": "results/PersonaChat-FLUX-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-original-gemma-4-E2B-it.PersonaChat-ERNIE-original-gemma-4-E2B-it
Visual Memory Results: personachat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-original-gemma-4-E2B-it",
"results_jsonl": "results/PersonaChat-ERNIE-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-original-gemma-4-E2B-it.ConvAI2-FLUX-original-gemma-4-E2B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-E2B-it.metrics-outputs-gemma-2-2b-layer-12-token-cachemetrics-outputs-gemma-2-2b-layer-03-token-cachemetrics-outputs-gemma-2-2b-layer-18-token-cachegemma-4-E2B-it-ValleyBench-benchmarkBenchmark of google/gemma-4-E2B-it against ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 72.2% with Python tool.
Metric
Value
Correct
722
Incorrect
261
Errors
17
Total samples
1000
Python tool calls
915
Python tool errors
0
Total completion tokens
872,102
gemma-4-E2B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E2B-it against SuperGPQA dataset. None
Accuracy: 32.7% with Python tool.
Metric
Value
Correct
328
Incorrect
666
Errors
8
Total samples
1002
Python tool calls
274
Python tool errors
22
Total completion tokens
1,999,635
google__gemma-2-2b-details
Dataset Card for Evaluation run of google/gemma-2-2b
Dataset automatically created during the evaluation run of model google/gemma-2-2b
The dataset is composed of 77 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__gemma-2-2b-details.google__gemma-2b-details
Dataset Card for Evaluation run of google/gemma-2b
Dataset automatically created during the evaluation run of model google/gemma-2b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__gemma-2b-details.
