datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-oss120b-generated-perfectblendgpt-oss120b-generated-magpie-1m-v0.1hh-rlhf-harmless-base-rollouts-gpt-oss-20b-diverse-openrouterhh-rlhf-helpful-base-rollouts-gpt-oss-20b-diverse-openroutertogethercomputer__GPT-NeoXT-Chat-Base-20B-details
Dataset Card for Evaluation run of togethercomputer/GPT-NeoXT-Chat-Base-20B
Dataset automatically created during the evaluation run of model togethercomputer/GPT-NeoXT-Chat-Base-20B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/togethercomputer__GPT-NeoXT-Chat-Base-20B-details.game-eval-qwen-Qwen3-8B-Base-vs-openai-gpt-4.1-mini-20250715-094130youtube_baseline_gpt-5-mini-2025-08-07olmo-3-preference-mix-deltas-gpt_and_100k_base_complement-DECONhh-rlhf-harmless-base-rollouts-gpt-5.1-adult
Gemma Reward-Scored Rollouts Dataset
Generation Parameters
{
"input": {
"source_type": "hf",
"repo_id": "MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-adult",
"split": "test",
"path": ""
},"base_dataset": {
"id": "MWilinski/hh-rlhf-harmless-base",
"split": "test",
"prompt_field": "prompt"
},
"selection": {
"prompt_indices": []
},
"scoring": {
"backend": "openrouter",
"model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-adult.hh-rlhf-helpful-base-rollouts-gpt-5.1-child
Gemma Reward-Scored Rollouts Dataset
Generation Parameters
{
"input": {
"source_type": "hf",
"repo_id": "MWilinski/hh-rlhf-helpful-base-rollouts-gpt-5.1-child",
"split": "test",
"path": ""
},"base_dataset": {
"id": "MWilinski/hh-rlhf-helpful-base",
"split": "test",
"prompt_field": "prompt"
},
"selection": {
"prompt_indices": []
},
"scoring": {
"backend": "openrouter",
"model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-helpful-base-rollouts-gpt-5.1-child.hh-rlhf-helpful-base-rollouts-gpt-5.1-adult
Gemma Reward-Scored Rollouts Dataset
Generation Parameters
{
"input": {
"source_type": "hf",
"repo_id": "MWilinski/hh-rlhf-helpful-base-rollouts-gpt-5.1-adult",
"split": "test",
"path": ""
},"base_dataset": {
"id": "MWilinski/hh-rlhf-helpful-base",
"split": "test",
"prompt_field": "prompt"
},
"selection": {
"prompt_indices": []
},
"scoring": {
"backend": "openrouter",
"model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-helpful-base-rollouts-gpt-5.1-adult.game-eval-qwen-Qwen3-4B-Base-vs-openai-gpt-4.1-mini-20250715-094045game-eval-qwen-Qwen3-0.6B-Base-vs-openai-gpt-4.1-mini-20250715-094001hh-rlhf-harmless-base-rollouts-gpt-5.1-child
Gemma Reward-Scored Rollouts Dataset
Generation Parameters
{
"input": {
"source_type": "hf",
"repo_id": "MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-child",
"split": "test",
"path": ""
},"base_dataset": {
"id": "MWilinski/hh-rlhf-harmless-base",
"split": "test",
"prompt_field": "prompt"
},
"selection": {
"prompt_indices": []
},
"scoring": {
"backend": "openrouter",
"model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-child.hh-rlhf-helpful-base-rollouts-gpt-5.1-policy
Gemma Reward-Scored Rollouts Dataset
Generation Parameters
{
"input": {
"source_type": "hf",
"repo_id": "MWilinski/hh-rlhf-helpful-base-rollouts-gpt-5.1-policy",
"split": "train",
"path": ""
},"base_dataset": {
"id": "MWilinski/hh-rlhf-helpful-base",
"split": "train",
"prompt_field": "prompt"
},
"selection": {
"prompt_indices": []
},
"scoring": {
"backend": "openrouter",
"model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-helpful-base-rollouts-gpt-5.1-policy.hh-rlhf-harmless-base-rollouts-gpt-oss-20b
Gemma Reward-Scored Rollouts Dataset
Generation Parameters
{
"input": {
"source_type": "hf",
"repo_id": "MWilinski/hh-rlhf-harmless-base-rollouts-gpt-oss-20b",
"split": "test",
"path": ""
},"base_dataset": {
"id": "MWilinski/hh-rlhf-harmless-base",
"split": "test",
"prompt_field": "prompt"
},
"selection": {
"prompt_indices": []
},
"scoring": {
"backend": "openrouter",
"model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-harmless-base-rollouts-gpt-oss-20b.hh-rlhf-helpful-base-rollouts-gpt-oss-20b
Gemma Reward-Scored Rollouts Dataset
Generation Parameters
{
"input": {
"source_type": "hf",
"repo_id": "MWilinski/hh-rlhf-helpful-base-rollouts-gpt-oss-20b",
"split": "test",
"path": ""
},
"base_dataset": {
"id": "MWilinski/hh-rlhf-helpful-base",
"split": "test",
"prompt_field": "prompt"
},
"selection": {
"prompt_indices": []
},
"scoring": {
"backend": "openrouter",
"model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-helpful-base-rollouts-gpt-oss-20b.game-eval-qwen-Qwen3-0.6B-Base-vs-openai-gpt-4.1-mini-20250714-232425hh-rlhf-harmless-base-rollouts-gpt-5.1-policy
Gemma Reward-Scored Rollouts Dataset
Generation Parameters
{
"input": {
"source_type": "hf",
"repo_id": "MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-policy",
"split": "train",
"path": ""
},"base_dataset": {
"id": "MWilinski/hh-rlhf-harmless-base",
"split": "train",
"prompt_field": "prompt"
},
"selection": {
"prompt_indices": []
},
"scoring": {
"backend": "openrouter",
"model": "google/gemma-3-27b-it"… See the full description on the dataset page: https://huggingface.co/datasets/MWilinski/hh-rlhf-harmless-base-rollouts-gpt-5.1-policy.game-eval-qwen-Qwen3-14B-Base-vs-openai-gpt-4.1-mini-20250715-002156game-eval-qwen-Qwen3-1.7B-Base-vs-openai-gpt-4.1-mini-20250715-094027game-eval-qwen-Qwen3-4B-Base-vs-openai-gpt-4.1-mini-20250714-235053reddit_baseline_gpt-5-mini-2025-08-07_persona_valuesgame-eval-qwen-Qwen3-8B-Base-vs-openai-gpt-4.1-mini-20250715-000425bigcodebench-complete_qwen7b_gpt-4o-mini_base_att20_sol5_fixed_consistencyGPT_judgement_base_vs_sft_dpoarena_baseline_gpt-large_zdraft_gpt-large_ep4_sep17olmo-3-preference-mix-deltas-gpt_and_100k_base_complementarena_baseline_gpt-small_zdraft_gpt-small_ep4_sep17amazon_baseline_gpt-5-mini-2025-08-07
