datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MATH-500
Dataset Card for MATH-500
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
instruction-datasetThis is the blind eval dataset of high-quality, diverse, human-written instructions with demonstrations. We will be using this for step 3 evaluations in our RLHF pipeline.
HuggingFaceH4__zephyr-7b-beta-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-beta-details.HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details.instruction-pilot-outputs-filteredHuggingFaceH4__zephyr-7b-alpha-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-alpha
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-alpha
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-alpha-details.HuggingFaceH4__zephyr-7b-gemma-v0.1-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-gemma-v0.1
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-gemma-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-gemma-v0.1-details.self-instruct-seedManually created seed dataset used in bootstrapping in the Self-instruct paper https://arxiv.org/abs/2212.10560. This is part of the instruction fine-tuning datasets.
lm-eval-results-HuggingFaceH4-mistral-7b-sft-beta-private
Dataset Card for Evaluation run of HuggingFaceH4/mistral-7b-sft-beta
Dataset automatically created during the evaluation run of model HuggingFaceH4/mistral-7b-sft-beta
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-HuggingFaceH4-mistral-7b-sft-beta-private.self-instruct-evalHuggingFaceH4-MATH-R1-dpo
Dataset Card for HuggingFaceH4-MATH-R1-dpo
本資料集為 HuggingFaceH4/MATH-500 等 MATH 評測資料延伸的 R1-style DPO 偏好集,並翻譯為繁體中文:每筆樣本針對一道數學題,提供「具完整 R1 風格 <think> 推理過程的較佳解(chosen)」與「推理品質較差的解(rejected)」。
Dataset Details
Dataset Description
資料以 MATH 風格題目為基礎(代數、幾何、組合、機率等),題幹翻譯為繁中,並使用 DeepSeek R1 等具 R1 風格之模型產出 chosen,配合較弱模型/簡化 prompt 的版本作為 rejected,以建構偏好對。
格式採 LLaMA-Factory 之 conversations + chosen + rejected 風格,gpt 回答前段為 <think>...</think> 思考段落,後段為最終解答(多以 \boxed{...} 收尾)。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/HuggingFaceH4-MATH-R1-dpo.aws-pm-pilotPilot annotations for PM dataset that will be used for RLHF. The dataset used outputs from opensource models (https://huggingface.co/spaces/HuggingFaceH4/instruction-models-outputs) on a mix on Anthropic hh-rlhf (https://huggingface.co/datasets/HuggingFaceH4/hh-rlhf) dataset and Self-Instruct's seed (https://huggingface.co/datasets/HuggingFaceH4/self-instruct-seed) dataset.
scale-pm-pilotPilot annotations for PM dataset that will be used for RLHF. The dataset used outputs from opensource models (https://huggingface.co/spaces/HuggingFaceH4/instruction-models-outputs) on a mix on Anthropic hh-rlhf (https://huggingface.co/datasets/HuggingFaceH4/hh-rlhf) dataset and Self-Instruct's seed (https://huggingface.co/datasets/HuggingFaceH4/self-instruct-seed) dataset.
HuggingFaceH4-ultrachat_200k
HuggingFaceH4/ultrachat_200k
This is a reprocessed version of HuggingFaceH4/ultrachat_200k.
Each row has been processed to fit ShareGPT-like format with the correct turn order. Duplicate rows have been removed.
The splits remain the same as the original.
Koala-test-setThis dataset is taken from https://github.com/arnav-gudibande/koala-test-set
surge-pm-pilotPilot annotations for PM dataset that will be used for RLHF. The dataset used outputs from opensource models (https://huggingface.co/spaces/HuggingFaceH4/instruction-models-outputs) on a mix on Anthropic hh-rlhf (https://huggingface.co/datasets/HuggingFaceH4/hh-rlhf) dataset and Self-Instruct's seed (https://huggingface.co/datasets/HuggingFaceH4/self-instruct-seed) dataset.
