datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-ntnhan-Llama3-8B-MetaMath-private
Dataset Card for Evaluation run of ntnhan/Llama3-8B-MetaMath
Dataset automatically created during the evaluation run of model ntnhan/Llama3-8B-MetaMath
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-ntnhan-Llama3-8B-MetaMath-private.a1_math_metamath_eval_1331
mlfoundations-dev/a1_math_metamath_eval_1331
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
GPQADiamond
MMLUPro
LiveCodeBench
CodeElo
JEEBench
Accuracy
13.7
58.0
74.2
39.6
28.8
9.8
2.5
34.4
AIME24
Average Accuracy: 13.67% ± 1.29%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
16.67%
5
30
3
13.33%
4
30
4
13.33%
4
30
5
13.33%
4
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_math_metamath_eval_1331.danish-metamath-gsmmath-rag-metamathqaesperanto-metamath-gsmlm-eval-results-meta-math-MetaMath-Mistral-7B-private
Dataset Card for Evaluation run of meta-math/MetaMath-Mistral-7B
Dataset automatically created during the evaluation run of model meta-math/MetaMath-Mistral-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-meta-math-MetaMath-Mistral-7B-private.cleand_meta-math_MetaMathQA元データ: https://huggingface.co/datasets/meta-math/MetaMathQA
データ件数: 394,369
平均トークン数: 233
最大トークン数: 2,874
合計トークン数: 91,798,611
ファイル形式: JSONL
ファイルサイズ: 297.9 MB
=================== 以下、加工内容をclaudeでまとめ。
MetaMathQAデータセット加工内容
データ読み込み・準備
HuggingFace Datasetsからmeta-math/MetaMathQAの訓練データ(395,000件)を読み込み
DeepSeek-R1-Distill-Qwen-32Bトークナイザーを使用してトークン数を計算
データ構造の理解・分析
全てのresponseが"The answer is:"で終わる統一フォーマットであることを確認
original_questionとresponseを結合してトークン数計算用テキストを作成… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_meta-math_MetaMathQA.TencentARC__MetaMath-Mistral-Pro-details
Dataset Card for Evaluation run of TencentARC/MetaMath-Mistral-Pro
Dataset automatically created during the evaluation run of model TencentARC/MetaMath-Mistral-Pro
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TencentARC__MetaMath-Mistral-Pro-details.metamath-grouped-diverse-majoritymetamath-hint-v5-qwen-base-gen__19250_21000metamath-hint-v5-qwen-32B-base-gen__2250_4500metamath-filtered-qwen32B-genmetamath-hint-v4-qwen-base-gen__17500_20000metamath-hint-v5-qwen-32B-base-gen__9000_11250oh_v1.2_sin_metamath_diversitymetamath-hint-v5-qwen-32B-base-gen__0_2250metamath-hint-v5-qwen-32B-concatmetamath-hint-v5-qwen-base-gen__14000_15750metamath-hint-v5-qwen-32B-base-gen__15750_18000metamath-hint-v5-qwen-32B__12250_14000metamath-groupedmetamath-grouped-jsonqwen3-metamathqa-rewardsmetamath-filtered-qwen-genmetamath-hint-v5-qwen-32B__3500_5250metamath-hint-v5-qwen-32Bmetamath-grouped-diverse-rouge05metamath-hint-v5-qwen-32B-base-gen__24750_27000metamath-hint-v5-qwen-32B-base-gen__22500_24750metamath-hint-v5-qwen-32B__7000_8750
