datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengpt-x_hellaswagxThis is a copy of the translations from openGPT-X/hellaswagx, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the HellaSwag dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_hellaswagx.hello-ot-imagenet-pca
HELLO representative ImageNet VA-VAE PCA features
This dataset contains four independently generated float32 NumPy matrices,
with shapes (262144, d) for d = 4, 32, 256, 2048, totaling about 2.29 GiB.
All use seed=42 and serve the public main-scaling and parameter-sensitivity
experiments. Larger sample counts are outside the published data scope.
The artifact contains numeric PCA-projected features only. It contains no
images, labels, captions, filenames, or ImageNet identifiers.… See the full description on the dataset page: https://huggingface.co/datasets/WenzhouXia/hello-ot-imagenet-pca.repro-sharp-inequalities-between-total-variation-and-hellinger-distances-for-gaussian-traces
Agent traces
Agent sessions published from a Trackio Logbook.
hellaswaggAraDICE-HellaSwag
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the HellaSwag split of the data.
Evaluation
We have used lm-harness eval framework to for… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-HellaSwag.Qwen3-32B_16episodes_comparisons_full
Qwen3-32B_16episodes_comparisons_full
This is a pairwise comparison dataset created from SWE-bench evaluation results.
Files
Qwen3-32B_16episodes_comparisons_full_comparison_pairs.jsonl: JSONL file containing comparison pairs
Metadata
{
"dataset_name": "Qwen3-32B_16episodes_comparisons_full",
"model_name": "Qwen3-32B",
"num_episodes": 16,
"episodes_used": [
1,
2,
3,
4,
5,
6,
7,
8,
9,
10,
11,
12,
13… See the full description on the dataset page: https://huggingface.co/datasets/helloelwin/Qwen3-32B_16episodes_comparisons_full.hellaswag_mr_results
Dataset Card for Evaluation run of google/gemma-2-2b
Dataset automatically created during the evaluation run of model google/gemma-2-2b
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/hellaswag_mr_results.hellaswag_ml_results
Dataset Card for Evaluation run of google/gemma-2-2b
Dataset automatically created during the evaluation run of model google/gemma-2-2b
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/hellaswag_ml_results.Turkish_HellaSwag10_Wrongshellaswaghellaswag_id_results
Dataset Card for Evaluation run of google/gemma-2-2b
Dataset automatically created during the evaluation run of model google/gemma-2-2b
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/hellaswag_id_results.finetuned_hellaswag_en_output_layer_20_results
Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0
Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_hellaswag_en_output_layer_20_results.forgetting-contamination-hellaswagThis dataset is a deduplicated subset of the validation split of hellaswag, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in https://huggingface.co/datasets/Rowan/hellaswag, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement for hellaswag if you want to work with the… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/forgetting-contamination-hellaswag.HellaSwag_English_Choices_1HellaSwag_ENG_Cosine_Afinetuned_hellaswag_ml_output_layer_25_results
Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-25-16k-4
Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-25-16k-4
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_hellaswag_ml_output_layer_25_results.Turkish_HellaswagHellaSwag10_Chunk_ThirdHellaSwag_TR_Cosine_AHellaSwag_TR_Cosine_BHellaswag_250_enHellaSwag_Turkish_Choices_0HellaSwag_Turkish_Choices_2Qwen3-32B_10episodes_comparisons_full
Qwen3-32B_10episodes_comparisons_full
This is a pairwise comparison dataset created from SWE-bench evaluation results.
Files
Qwen3-32B_10episodes_comparisons_full_comparison_pairs.jsonl: JSONL file containing comparison pairs
Metadata
{
"dataset_name": "Qwen3-32B_10episodes_comparisons_full",
"model_name": "Qwen3-32B",
"num_episodes": 10,
"episodes_used": [
1,
2,
3,
4,
5,
6,
7,
8,
9,
10
],
"statistics": {… See the full description on the dataset page: https://huggingface.co/datasets/helloelwin/Qwen3-32B_10episodes_comparisons_full.HellaSwag10_Chunk_FirstHellaSwag10_Chunk_SecondTurkish_HellaSwag10_Shuffle_Firsthellaswag_en_250_10shuffle_1hellaswag_en_250_10shuffle_2hellaswag_en_250_10shuffle_3
