datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code_Opt_Triton
Overview
This dataset, TEEN-D/Code_Opt_Triton, is an extended version of the publicly available GPUMODE/Inductor_Created_Data_Permissive dataset. It contains pairs of original (PyTorch or Triton) programs and their equivalent Triton code (generated by torch inductor), intended for training models in PyTorch-to-Triton code translation and optimization.
The primary modification in this extended version is that each optimized Triton code snippet is paired with both its original source… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/Code_Opt_Triton.full-repo-coverage-diffCode_Opt_Triton_Shuffled
TEEN-D/Code_Opt_Triton_Shuffled
Overview
This dataset, TEEN-D/Code_Opt_Triton_Shuffled, is a shuffled version of the extended TEEN-D/Code_Opt_Triton dataset (which itself is an extension of GPUMODE/Inductor_Created_Data_Permissive). It provides a collection of pairs of original (PyTorch or Triton) programs and their corresponding optimized Triton code, designed for training machine learning models for code translation and optimization tasks targeting GPUs.
The key… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/Code_Opt_Triton_Shuffled.train_adv_mean_diff_code_sabotage_safedifficulty_sorting_high_seed_code_w_openthoughtsb2_code_difficulty_eval_636d
mlfoundations-dev/b2_code_difficulty_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
21.0
69.5
77.4
29.4
47.3
45.5
47.4
16.1
19.6
AIME24
Average Accuracy: 21.00% ± 2.00%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2
23.33%
7
30
3
13.33%
4
30
4
16.67%
5… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_eval_636d.open-code-reasoning-rlvr-difficultydifficulty_sorting_easy_seed_code_w_openthoughtsdifficulty_sorting_medium_seed_code_w_openthoughtsdifficulty_sorting_random_seed_code_w_openthoughtsb2_code_difficulty_10ktulu-3-personas-code-rlvr-difficultyDiffuEraser-finetune-prompt-codeb2_code_difficulty_3k_eval_636d
mlfoundations-dev/b2_code_difficulty_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
21.7
62.0
76.0
27.4
44.9
38.2
31.4
10.3
13.5
AIME24
Average Accuracy: 21.67% ± 1.35%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
20.00%
6
30
3
16.67%
5
30
4
23.33%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_3k_eval_636d.base_get_difficulty_seed_codeb2_code_difficulty_0.3k_eval_636d
mlfoundations-dev/b2_code_difficulty_0.3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
16.0
56.0
76.0
27.8
39.7
36.5
14.4
6.1
6.2
AIME24
Average Accuracy: 16.00% ± 1.55%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
20.00%
6
30
2
20.00%
6
30
3
23.33%
7
30
4
6.67%
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_0.3k_eval_636d.b2_code_difficulty_1k_eval_636d
mlfoundations-dev/b2_code_difficulty_1k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
16.7
54.0
74.0
29.6
40.0
35.0
26.7
7.4
8.8
AIME24
Average Accuracy: 16.67% ± 0.94%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
20.00%
6
30
3
20.00%
6
30
4
13.33%
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_1k_eval_636d.train_adv_mean_diff_code_sabotage_unsafeb2_code_difficulty_opencodereasoningb2_code_difficulty_10k_eval_636d
mlfoundations-dev/b2_code_difficulty_10k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
21.3
59.0
68.8
28.2
46.1
42.8
40.1
11.9
13.7
AIME24
Average Accuracy: 21.33% ± 1.26%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
26.67%
8
30
2
23.33%
7
30
3
30.00%
9
30
4
20.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_10k_eval_636d.mean_diff_code_deception_adv_trainset_safemean_diff_code_deception_adv_trainset_unsafeone_shot_mean_diff_code_sabotage_safeone_shot_mean_diff_code_sabotage_unsafesynthetic-code-understanding-v2-rust-test-difficultcommitpack_original_cm_old_code_to_diffdifficulty_sorting_high_seed_codedifficulty_sorting_easy_seed_codedifficulty_sorting_medium_seed_codedifficulty_sorting_random_seed_code
