datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glm53-flash-fidelity-exl3-tr3-6bpw-v1
fidelity--glm53flash.malaiwah.quant.tr3-6bpw
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/GLM-5.3-Flash-TR3-6bpw.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-exl3-tr3-6bpw-v1.EleutherAI__gpt-j-6b-details
Dataset Card for Evaluation run of EleutherAI/gpt-j-6b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b
The dataset is composed of 111 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-j-6b-details.01-ai__Yi-1.5-6B-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-6B
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-6B
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-6B-details.01-ai__Yi-1.5-6B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-6B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-6B-Chat
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-6B-Chat-details.01-ai__Yi-6B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-6B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-6B-Chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-6B-Chat-details.stabilityai__stablelm-2-1_6b-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-1_6b
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-1_6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-1_6b-details.togethercomputer__GPT-JT-6B-v1-details
Dataset Card for Evaluation run of togethercomputer/GPT-JT-6B-v1
Dataset automatically created during the evaluation run of model togethercomputer/GPT-JT-6B-v1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/togethercomputer__GPT-JT-6B-v1-details.PygmalionAI__pygmalion-6b-details
Dataset Card for Evaluation run of PygmalionAI/pygmalion-6b
Dataset automatically created during the evaluation run of model PygmalionAI/pygmalion-6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PygmalionAI__pygmalion-6b-details.01-ai__Yi-6B-details
Dataset Card for Evaluation run of 01-ai/Yi-6B
Dataset automatically created during the evaluation run of model 01-ai/Yi-6B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration "results"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-6B-details.Youlln__2PRYMMAL-Yi1.5-6B-SLERP-details
Dataset Card for Evaluation run of Youlln/2PRYMMAL-Yi1.5-6B-SLERP
Dataset automatically created during the evaluation run of model Youlln/2PRYMMAL-Yi1.5-6B-SLERP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Youlln__2PRYMMAL-Yi1.5-6B-SLERP-details.01-ai__Yi-6B-200K-details
Dataset Card for Evaluation run of 01-ai/Yi-6B-200K
Dataset automatically created during the evaluation run of model 01-ai/Yi-6B-200K
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-6B-200K-details.LilRg__PRYMMAL-6B-slerp-details
Dataset Card for Evaluation run of LilRg/PRYMMAL-6B-slerp
Dataset automatically created during the evaluation run of model LilRg/PRYMMAL-6B-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__PRYMMAL-6B-slerp-details.stabilityai__stablelm-2-1_6b-chat-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-1_6b-chat
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-1_6b-chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-1_6b-chat-details.databricks__dolly-v1-6b-details
Dataset Card for Evaluation run of databricks/dolly-v1-6b
Dataset automatically created during the evaluation run of model databricks/dolly-v1-6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v1-6b-details.stabilityai__stablelm-2-zephyr-1_6b-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-zephyr-1_6b
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-zephyr-1_6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-zephyr-1_6b-details.ReST-MCTS_SciGLM-6B_ReST-MCTS_Policy_2ndicml2026-6bXLKr5NMT-repro-traces
Agent traces
Agent sessions published from a Trackio Logbook.
round-v6bpszemraj__Mistral-v0.3-6B-details
Dataset Card for Evaluation run of pszemraj/Mistral-v0.3-6B
Dataset automatically created during the evaluation run of model pszemraj/Mistral-v0.3-6B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/pszemraj__Mistral-v0.3-6B-details.SlimPajama-6B-JSONLDKYoon/SlimPajama-6B but in JSONL
ReST-MCTS_SciGLM-6B_ReST-EM-CoT_2ndReST-MCTS_SciGLM-6B_Self-Rewarding-DPO_2ndlalainy__ECE-PRYMMAL-YL-6B-SLERP-V1-details
Dataset Card for Evaluation run of lalainy/ECE-PRYMMAL-YL-6B-SLERP-V1
Dataset automatically created during the evaluation run of model lalainy/ECE-PRYMMAL-YL-6B-SLERP-V1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lalainy__ECE-PRYMMAL-YL-6B-SLERP-V1-details.lalainy__ECE-PRYMMAL-YL-6B-SLERP-V2-details
Dataset Card for Evaluation run of lalainy/ECE-PRYMMAL-YL-6B-SLERP-V2
Dataset automatically created during the evaluation run of model lalainy/ECE-PRYMMAL-YL-6B-SLERP-V2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lalainy__ECE-PRYMMAL-YL-6B-SLERP-V2-details.huggingface_6569_6bfsrh6d-pipeline-registryReST-MCTS_SciGLM-6B_ReST-EM-CoT_1stgamma-g1-121-gamma-30b-a6b-g1117-rc-card-20260617
Gamma-30B-A6B-G1.117 Release Candidate Card
Identity
Name: Gamma-30B-A6B-G1.117-RC
Architecture label: gamma_genesis_nemotron_h_scaffold_evostream_adapter
Scaffold/chassis: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Current accepted adapter chromosome: G1_117
What Is Real Right Now
G1_117 improved held-out v2 eval loss by 42.9602%.
Train rows used: 192.
Eval rows used: 32.
Target modules: ['q_proj', 'k_proj', 'v_proj', 'o_proj'].
Trainable… See the full description on the dataset page: https://huggingface.co/datasets/SemantaAI/gamma-g1-121-gamma-30b-a6b-g1117-rc-card-20260617.prithivMLmods__Llama-3.2-6B-AlgoCode-details
Dataset Card for Evaluation run of prithivMLmods/Llama-3.2-6B-AlgoCode
Dataset automatically created during the evaluation run of model prithivMLmods/Llama-3.2-6B-AlgoCode
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Llama-3.2-6B-AlgoCode-details.Flock_6bc0b9necf88c888d6b
