datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Caduceus-Dataset
Caduceus Project Dataset
Creator: Kquant03
About the Dataset: The Caduceus Project Dataset is a curated collection of scientific and medical protocols sourced from protocols.io and converted from PDF to markdown format. This dataset aims to help models learn to read complicated PDFs by either using computer vision on the PDF file, or through processing the raw text directly. You can find the… See the full description on the dataset page: https://huggingface.co/datasets/Kquant03/Caduceus-Dataset.qwen36-kquant-offload-mtp-swebench-lite100-results
Qwen3.6 K-Quant Offload MTP SWE-bench Lite 100 Results
This dataset contains the complete 5-model x 100-prompt runtime benchmark artifacts plus a detailed statistical analysis layer.
Primary conclusion: hot30/cold30 was the best decode-throughput run, while Q4_K_M had the best total wall clock. The ATX hot30/cold30 quantization significantly outperformed both Q4_K_M and Q3_K_XL on paired decode throughput, but Q4_K_M remains the elapsed-time control.
The ATX/K3 hot10, hot20, and… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-kquant-offload-mtp-swebench-lite100-results.details_Kquant03__Samlagast-7B-laser-bf16
Dataset Card for Evaluation run of Kquant03/Samlagast-7B-laser-bf16
Dataset automatically created during the evaluation run of model Kquant03/Samlagast-7B-laser-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Samlagast-7B-laser-bf16.details_Kquant03__Buttercup-V2-laser
Dataset Card for Evaluation run of Kquant03/Buttercup-V2-laser
Dataset automatically created during the evaluation run of model Kquant03/Buttercup-V2-laser on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Buttercup-V2-laser.details_Kquant03__CognitiveFusion2-4x7B-BF16
Dataset Card for Evaluation run of Kquant03/CognitiveFusion2-4x7B-BF16
Dataset automatically created during the evaluation run of model Kquant03/CognitiveFusion2-4x7B-BF16.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Kquant03__CognitiveFusion2-4x7B-BF16.lm-eval-results-Kquant03-NeuralTrix-7B-dpo-relaser-private
Dataset Card for Evaluation run of Kquant03/NeuralTrix-7B-dpo-relaser
Dataset automatically created during the evaluation run of model Kquant03/NeuralTrix-7B-dpo-relaser
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-NeuralTrix-7B-dpo-relaser-private.lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Nanashi-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Nanashi-2x7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private.details_Kquant03__NeuralTrix-7B-dpo-laser
Dataset Card for Evaluation run of Kquant03/NeuralTrix-7B-dpo-laser
Dataset automatically created during the evaluation run of model Kquant03/NeuralTrix-7B-dpo-laser on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__NeuralTrix-7B-dpo-laser.lm-eval-results-Kquant03-Cognito-2x7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Cognito-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Cognito-2x7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Cognito-2x7B-bf16-private.aurora-m-biden-harris-redteamed-ungatedThis is just an ungated version of aurora-m/biden-harris-redteam dataset. Makes it easier to work with when it's not gated.
BY ACCESSING THIS DATASET YOU AGREE YOU ARE 18 YEARS OLD OR OLDER AND UNDERSTAND THE RISKS OF USING THIS DATASET.
@article{tedeschi2024redteam,
author = {Simone Tedeschi, Felix Friedrich, Dung Nguyen, Nam Pham, Tanmay Laud, Chien Vu, Terry Yue Zhuo, Ziyang Luo, Ben Bogin, Tien-Tung Bui, Xuan-Son Vu, Paulo Villegas, Victor May, Huu Nguyen},
title = {Biden-Harris… See the full description on the dataset page: https://huggingface.co/datasets/Kquant03/aurora-m-biden-harris-redteamed-ungated.details_Kquant03__NeuralTrix-7B-dpo-relaser
Dataset Card for Evaluation run of Kquant03/NeuralTrix-7B-dpo-relaser
Dataset automatically created during the evaluation run of model Kquant03/NeuralTrix-7B-dpo-relaser on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__NeuralTrix-7B-dpo-relaser.details_Kquant03__Kaltsit-16x7B-bf16
Dataset Card for Evaluation run of Kquant03/Kaltsit-16x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Kaltsit-16x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Kaltsit-16x7B-bf16.details_Kquant03__Buttercup-V2-bf16
Dataset Card for Evaluation run of Kquant03/Buttercup-V2-bf16
Dataset automatically created during the evaluation run of model Kquant03/Buttercup-V2-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Buttercup-V2-bf16.details_Kquant03__CognitiveFusion-4x7B-bf16-MoE
Dataset Card for Evaluation run of Kquant03/CognitiveFusion-4x7B-bf16-MoE
Dataset automatically created during the evaluation run of model Kquant03/CognitiveFusion-4x7B-bf16-MoE on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__CognitiveFusion-4x7B-bf16-MoE.details_Kquant03__Prokaryote-8x7B-bf16
Dataset Card for Evaluation run of Kquant03/Prokaryote-8x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Prokaryote-8x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Prokaryote-8x7B-bf16.details_Kquant03__DolphinHermesPro-ModelStock
Dataset Card for Evaluation run of Kquant03/DolphinHermesPro-ModelStock
Dataset automatically created during the evaluation run of model Kquant03/DolphinHermesPro-ModelStock on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__DolphinHermesPro-ModelStock.Kquant03__L3-Pneuma-8B-details
Dataset Card for Evaluation run of Kquant03/L3-Pneuma-8B
Dataset automatically created during the evaluation run of model Kquant03/L3-Pneuma-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Kquant03__L3-Pneuma-8B-details.details_Kquant03__Samlagast-7B-bf16
Dataset Card for Evaluation run of Kquant03/Samlagast-7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Samlagast-7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Samlagast-7B-bf16.details_Kquant03__CognitiveFusion2-4x7B-BF16
Dataset Card for Evaluation run of Kquant03/CognitiveFusion2-4x7B-BF16
Dataset automatically created during the evaluation run of model Kquant03/CognitiveFusion2-4x7B-BF16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__CognitiveFusion2-4x7B-BF16.details_Kquant03__Buttercup-4x7B-bf16
Dataset Card for Evaluation run of Kquant03/Buttercup-4x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Buttercup-4x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Buttercup-4x7B-bf16.details_Kquant03__Nanashi-2x7B-bf16
Dataset Card for Evaluation run of Kquant03/Nanashi-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Nanashi-2x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Nanashi-2x7B-bf16.details_Kquant03__Azathoth-16x7B-bf16
Dataset Card for Evaluation run of Kquant03/Azathoth-16x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Azathoth-16x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Azathoth-16x7B-bf16.details_Kquant03__Hippolyta-7B-bf16
Dataset Card for Evaluation run of Kquant03/Hippolyta-7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Hippolyta-7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Hippolyta-7B-bf16.details_Kquant03__BurningBruce-SOLAR-8x10.7B-bf16
Dataset Card for Evaluation run of Kquant03/BurningBruce-SOLAR-8x10.7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/BurningBruce-SOLAR-8x10.7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__BurningBruce-SOLAR-8x10.7B-bf16.details_Kquant03__Cognito-2x7B-bf16
Dataset Card for Evaluation run of Kquant03/Cognito-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Cognito-2x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Cognito-2x7B-bf16.details_Kquant03__Medusa-7B-bf16
Dataset Card for Evaluation run of Kquant03/Medusa-7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Medusa-7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Medusa-7B-bf16.lm-eval-results-Kquant03-NeuralTrix-7B-dpo-laser-private
Dataset Card for Evaluation run of Kquant03/NeuralTrix-7B-dpo-laser
Dataset automatically created during the evaluation run of model Kquant03/NeuralTrix-7B-dpo-laser
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-NeuralTrix-7B-dpo-laser-private.details_Kquant03__FrankenDPO-4x7B-bf16
Dataset Card for Evaluation run of Kquant03/FrankenDPO-4x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/FrankenDPO-4x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__FrankenDPO-4x7B-bf16.details_Kquant03__Raiden-16x3.43B
Dataset Card for Evaluation run of Kquant03/Raiden-16x3.43B
Dataset automatically created during the evaluation run of model Kquant03/Raiden-16x3.43B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Raiden-16x3.43B.details_Kquant03__Eukaryote-8x7B-bf16
Dataset Card for Evaluation run of Kquant03/Eukaryote-8x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Eukaryote-8x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Eukaryote-8x7B-bf16.
