datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma-2b-dictionary-embeddings-all-layers
Gemma-2B Dictionary Embeddings - All Layers
This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers.
Dataset Structure
metadata.json: Contains dataset metadata (model info, dimensions, word count)
embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26)
Usage
import pickle
from huggingface_hub import hf_hub_download
# Download a specific layer
layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.Ultrafeedback-llama3-8b-Instruct-optimal-selection-grm-gemma2bgemma2b-classification-eval-by-claude3sonnetgemma_2b_outputs
Gemma 2B Green LLM Experiment Outputs
This dataset repository contains experiment artifacts for Gemma 2B green-LLM runs, including LoRA adapter checkpoints, metrics, predictions, carbon logs, and figures.
Contents
checkpoints/: LoRA adapter checkpoints for CE baseline and joint-loss variants.
metrics/: training histories, SQuAD and MMLU summaries, prediction CSVs, calibration tables, and surrogate weights.
logs/: run histories and carbon summary JSON files.
carbon/:… See the full description on the dataset page: https://huggingface.co/datasets/PhotonTJ/gemma_2b_outputs.gemma2b-closedqa-eval-by-claude3sonnetgemma2b-coding-eval-by-claude3sonnetgemma2b-classification-eval-by-gemini15flashgemma2b-closedqa-eval-by-gemini15flashgemma2b-coding-eval-by-gemini15flashgemma2b_sft_gen_intern20b_label_rmUltrafeedback-llama3-8b-Instruct-kmeans-selection-grm-gemma2bUltrafeedback-llama3-8b-Instruct-weighted-1vsbottomk-selection-grm-gemma2bgemma-2b-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
gemma-2b-finetune-soliditygemma2b_sft_gen_intern20b_label_rm_0_25noisegemma-2b-finetune-javagemma2b-it-1.1-summarize-eval-by-gpt4ogemma-2b-baseline-soliditygemma2b-closedqa-eval-by-gpt4ogemma_2b_alignscoregemma2b-classification-eval-by-gpt4ogemma2b-coding-eval-by-gpt4oalignscore_gemma2b_correct_answergemma_2balignscore_gemma2b_incorrect_answer
