datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md
base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb
chat_examples.pt is the same but for lmsys chat data
chat_base_examples.pt is a merge of the two above files.
All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.details_google__gemma-2-2b-it_v2
Dataset Card for Evaluation run of google/gemma-2-2b-it
Dataset automatically created during the evaluation run of model google/gemma-2-2b-it.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_google__gemma-2-2b-it_v2.gemma-2b-jailbreak-behavior-dataset-v2reward-bench-gemma-2-2b-it-yes-nogemma-2b-jailbreak-behavior-dataset-v2collapse_gemma-2-2b_hs2_replace_iter7_sftsd0_temp1_max_seq_len512ConvAI2-Qwen-original-gemma-4-E2B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-E2B-it.ConvAI2-ERNIE-original-gemma-4-E2B-it
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-ERNIE-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-gemma-4-E2B-it.pair-preference-dataset-700K_subset-2-of-2_gemma-2b-it_1-of-2ConvAI2-Qwen-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E2B-it.ConvAI2-FLUX-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E2B-it.ConvAI2-ERNIE-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it.jebish7__gemma-2-2b-it-details
Dataset Card for Evaluation run of jebish7/gemma-2-2b-it
Dataset automatically created during the evaluation run of model jebish7/gemma-2-2b-it
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jebish7__gemma-2-2b-it-details.eval_bosch_gemma-4-E2B-it_gens_T0_wfs2ConvAI2-FLUX-original-gemma-4-E2B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-E2B-it.triviaqa_all_gemma-2-2b-itgemma-2-2b-it-sciq-train-deploy-prefixespair-preference-dataset-700K_subset-2-of-4_gemma-2b_1of4_iter3_conf-0.8_bs128_lr1e-5_conf-0.8baseline_iter1_lora_gemma-2-2b_real8000_syn2000_seed0_baseline_hs2_massive_iter1_1ca6716fcollapse_gemma-2-2b_hs2_replace_iter3_sftsd0_temp1_max_seq_len512collapse_gemma-2-2b_hs2_accumulate_iter8_sftsd1_temp1_max_seq_len512dialogsum-gemma-2-2b-itgoogle__gemma-2-2b-details
Dataset Card for Evaluation run of google/gemma-2-2b
Dataset automatically created during the evaluation run of model google/gemma-2-2b
The dataset is composed of 77 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__gemma-2-2b-details.ai-vs-human-google-gemma-2-2b-it
AI vs Human dataset on the CNN Daily mails
Dataset Description
This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model.
Each article was randomly truncated between 25% and 50% of its length.
The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation.
Data Fields
'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-google-gemma-2-2b-it.google__gemma-2-2b-it-details
Dataset Card for Evaluation run of gg-hf/gemma-2-2b-it
Dataset automatically created during the evaluation run of model gg-hf/gemma-2-2b-it
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__gemma-2-2b-it-details.collapse_gemma-2-2b_hs2_accumulate_iter1_sftsd0_temp1_max_seq_len512gemma2-2b-it-jailbreaksreward-bench-gemma-2-2b-it-set3-scoresD_improved_v2_gemma-2-2b_real5000_syn5000_seed0_baseline_hs2_massive_iter1_sftsd_8f55bc93
