CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01code-rl /full-repo-coverage-difftextn<1K0 likes28 downloads2mo agoHugging Face02candywal /train_adv_mean_diff_code_sabotage_safetextn<1K0 likes20 downloads1y agoHugging Face03mlfoundations-dev /difficulty_sorting_high_seed_code_w_openthoughtstabular100K<n<1M0 likes14 downloads2y agoHugging Face04mlfoundations-dev /b2_code_difficulty_eval_636d mlfoundations-dev/b2_code_difficulty_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 21.0 69.5 77.4 29.4 47.3 45.5 47.4 16.1 19.6 AIME24 Average Accuracy: 21.00% ± 2.00% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 16.67% 5 30 2 23.33% 7 30 3 13.33% 4 30 4 16.67% 5… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_eval_636d.tabular1K<n<10K0 likes13 downloads1y agoHugging Face05saurabh5 /open-code-reasoning-rlvr-difficultytext1K<n<10K0 likes13 downloads1y agoHugging Face06mlfoundations-dev /difficulty_sorting_easy_seed_code_w_openthoughtstabular100K<n<1M0 likes12 downloads2y agoHugging Face07mlfoundations-dev /difficulty_sorting_medium_seed_code_w_openthoughtstabular100K<n<1M0 likes11 downloads2y agoHugging Face08mlfoundations-dev /difficulty_sorting_random_seed_code_w_openthoughtstabular100K<n<1M0 likes11 downloads2y agoHugging Face09mlfoundations-dev /b2_code_difficulty_10ktabular10K<n<100K0 likes10 downloads1y agoHugging Face10saurabh5 /tulu-3-personas-code-rlvr-difficultytext10K<n<100K0 likes10 downloads1y agoHugging Face11mlfoundations-dev /b2_code_difficulty_3k_eval_636d mlfoundations-dev/b2_code_difficulty_3k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 21.7 62.0 76.0 27.4 44.9 38.2 31.4 10.3 13.5 AIME24 Average Accuracy: 21.67% ± 1.35% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 23.33% 7 30 2 20.00% 6 30 3 16.67% 5 30 4 23.33%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_3k_eval_636d.tabular1K<n<10K0 likes8 downloads1y agoHugging Face12mlfoundations-dev /base_get_difficulty_seed_codetabular100K<n<1M1 likes7 downloads2y agoHugging Face13mlfoundations-dev /b2_code_difficulty_0.3k_eval_636d mlfoundations-dev/b2_code_difficulty_0.3k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 16.0 56.0 76.0 27.8 39.7 36.5 14.4 6.1 6.2 AIME24 Average Accuracy: 16.00% ± 1.55% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 20.00% 6 30 2 20.00% 6 30 3 23.33% 7 30 4 6.67% 2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_0.3k_eval_636d.tabular1K<n<10K0 likes7 downloads1y agoHugging Face14mlfoundations-dev /b2_code_difficulty_1k_eval_636d mlfoundations-dev/b2_code_difficulty_1k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 16.7 54.0 74.0 29.6 40.0 35.0 26.7 7.4 8.8 AIME24 Average Accuracy: 16.67% ± 0.94% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 13.33% 4 30 2 20.00% 6 30 3 20.00% 6 30 4 13.33% 4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_1k_eval_636d.tabular1K<n<10K0 likes7 downloads1y agoHugging Face15candywal /train_adv_mean_diff_code_sabotage_unsafetextn<1K0 likes7 downloads1y agoHugging Face16mlfoundations-dev /b2_code_difficulty_opencodereasoningtabular10K<n<100K0 likes6 downloads1y agoHugging Face17mlfoundations-dev /b2_code_difficulty_10k_eval_636d mlfoundations-dev/b2_code_difficulty_10k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 21.3 59.0 68.8 28.2 46.1 42.8 40.1 11.9 13.7 AIME24 Average Accuracy: 21.33% ± 1.26% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 26.67% 8 30 2 23.33% 7 30 3 30.00% 9 30 4 20.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_code_difficulty_10k_eval_636d.tabular1K<n<10K0 likes6 downloads1y agoHugging Face18candywal /mean_diff_code_deception_adv_trainset_safetextn<1K0 likes6 downloads1y agoHugging Face19candywal /mean_diff_code_deception_adv_trainset_unsafetextn<1K0 likes6 downloads1y agoHugging Face20candywal /one_shot_mean_diff_code_sabotage_safetextn<1K0 likes6 downloads1y agoHugging Face21candywal /one_shot_mean_diff_code_sabotage_unsafetextn<1K0 likes6 downloads1y agoHugging Face22justus27 /synthetic-code-understanding-v2-rust-test-difficulttextn<1K1 likes6 downloads1y agoHugging Face23mlfoundations-dev /difficulty_sorting_high_seed_codetabular1K<n<10K0 likes5 downloads2y agoHugging Face24mlfoundations-dev /difficulty_sorting_easy_seed_codetabular1K<n<10K0 likes5 downloads2y agoHugging Face25mlfoundations-dev /difficulty_sorting_medium_seed_codetabular1K<n<10K0 likes5 downloads2y agoHugging Face26mlfoundations-dev /difficulty_sorting_random_seed_codetabular1K<n<10K0 likes5 downloads2y agoHugging Face27candywal /one_shot_mean_diff_code_rule_violation_safetextn<1K0 likes5 downloads1y agoHugging Face28candywal /train_adv_mean_diff_code_deception_safetextn<1K0 likes5 downloads1y agoHugging Face29mlfoundations-dev /b2_code_difficulty_code_golftabular10K<n<100K0 likes4 downloads1y agoHugging Face30mlfoundations-dev /b2_code_difficultytabular10K<n<100K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.