CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations-dev /phi_24K_qwq_6K_eval_2e29 mlfoundations-dev/phi_24K_qwq_6K_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 35.0 66.2 80.8 8.0 41.9 46.6 1.2 3.2 2.0 27.7 3.6 0.5 AIME24 Average Accuracy: 35.00% ± 2.22% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 36.67% 11 30 2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_24K_qwq_6K_eval_2e29.tabular10K<n<100K0 likes24 downloads1y agoHugging Face02roborovski /phi-2-labeled Dataset Card for "phi-1" More Information needed tabular10K<n<100K1 likes21 downloads3y agoHugging Face03onepaneai /phi2-spl-evaluationtabularn<1K0 likes17 downloads2y agoHugging Face04roborovski /phi-2-embeddings Dataset Card for "phi-2-embeddings" More Information needed tabular10K<n<100K0 likes15 downloads3y agoHugging Face05mlfoundations-dev /phi_27K_qwq_3K_eval_2e29 mlfoundations-dev/phi_27K_qwq_3K_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 31.3 57.0 63.4 7.4 42.6 47.1 1.2 0.5 0.1 23.7 2.9 0.6 AIME24 Average Accuracy: 31.33% ± 1.26% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 30.00% 9 30 2 26.67%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_27K_qwq_3K_eval_2e29.tabular10K<n<100K0 likes8 downloads1y agoHugging Face06open-llm-leaderboard /netcat420__MFANN-abliterated-phi2-merge-unretrained-detailsgated Dataset Card for Evaluation run of netcat420/MFANN-abliterated-phi2-merge-unretrained Dataset automatically created during the evaluation run of model netcat420/MFANN-abliterated-phi2-merge-unretrained The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/netcat420__MFANN-abliterated-phi2-merge-unretrained-details.tabular10K<n<100K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.