CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twinkle-ai /phi-4-eval-logs-and-scorestabular100K<n<1M0 likes152 downloads7mo agoHugging Face02atul10 /arm_o0_phi4_multi_full_14_alltabular10K<n<100K0 likes50 downloads1y agoHugging Face03open-llm-leaderboard /benhaotang__phi4-qwq-sky-t1-detailsgated Dataset Card for Evaluation run of benhaotang/phi4-qwq-sky-t1 Dataset automatically created during the evaluation run of model benhaotang/phi4-qwq-sky-t1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/benhaotang__phi4-qwq-sky-t1-details.tabular10K<n<100K0 likes47 downloads2y agoHugging Face04open-llm-leaderboard /EpistemeAI__DeepThinkers-Phi4-detailsgated Dataset Card for Evaluation run of EpistemeAI/DeepThinkers-Phi4 Dataset automatically created during the evaluation run of model EpistemeAI/DeepThinkers-Phi4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__DeepThinkers-Phi4-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face05open-llm-leaderboard /Triangle104__Phi4-RP-o1-Ablit-detailsgated Dataset Card for Evaluation run of Triangle104/Phi4-RP-o1-Ablit Dataset automatically created during the evaluation run of model Triangle104/Phi4-RP-o1-Ablit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Phi4-RP-o1-Ablit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face06open-llm-leaderboard /Quazim0t0__Phi4Basis-14B-sce-detailsgated Dataset Card for Evaluation run of Quazim0t0/Phi4Basis-14B-sce Dataset automatically created during the evaluation run of model Quazim0t0/Phi4Basis-14B-sce The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Phi4Basis-14B-sce-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face07atul10 /arm_o0_phi4_multi_smoketesttabularn<1K0 likes27 downloads1y agoHugging Face08mlfoundations-dev /Phi-4-reasoning-plus_eval_5554 mlfoundations-dev/Phi-4-reasoning-plus_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 76.0 96.2 84.0 14.6 83.5 66.8 0.8 2.4 3.5 7.1 53.0 68.0 0.5 AIME24 Average Accuracy: 76.00% ± 1.23% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Phi-4-reasoning-plus_eval_5554.tabular10K<n<100K0 likes20 downloads1y agoHugging Face09gx-ai-architect /numinamath-178k-phi4-bon-verified-dpo-trl-40ktabular10K<n<100K0 likes17 downloads2y agoHugging Face10if001 /physics_textbook_phi4tabular10K<n<100K0 likes16 downloads1y agoHugging Face11if001 /world_history_textbook_phi4tabular1K<n<10K0 likes16 downloads1y agoHugging Face12gx-ai-architect /numinamath-178k-phi4-bon-verified-dpo-trltabular10K<n<100K0 likes13 downloads2y agoHugging Face13mlfoundations-dev /Phi-4-reasoning-plus_eval_2693tabular1K<n<10K0 likes12 downloads1y agoHugging Face14if001 /psychology_textbook_phi4tabular10K<n<100K1 likes12 downloads1y agoHugging Face15math-extraction-comp /Undi95__Phi4-abliteratedtabular1K<n<10K0 likes11 downloads2y agoHugging Face16Lakshan2003 /Phi-4-Mini-customerservice-Human-evaluator_2_datatabularn<1K0 likes11 downloads8mo agoHugging Face17if001 /math_textbook_phi4tabular10K<n<100K1 likes10 downloads1y agoHugging Face18if001 /jp_history_textbook_phi4tabular1K<n<10K0 likes10 downloads1y agoHugging Face19Lakshan2003 /Phi-4-mini-customerservice-context-summarization-llm-judge-data customer-service-context-summarization-evaluation-data Phi-4-mini-customerservice-context-summarization-llm-judge-data Dataset updated with context summarization evaluation columns. This README refresh triggers Hugging Face metadata re-index. tabular1K<n<10K0 likes10 downloads6mo agoHugging Face20if001 /bunpo_phi4phi4で以下の53の文法パターン × 2364vocab を生成し、フィルタリングを行っています。 0, "です/だ (肯定文)", 1, "ではありません/じゃない (否定文)", 3, "〜ます (動詞の丁寧形)", 4, "〜ません (動詞の否定形)", 5, "〜たい (希望・願望)", 6, "〜ている (進行形)", 7, "〜てください (依頼)", 8, "〜てもいいですか (許可)", 9, "〜なければなりません/〜なきゃいけない (義務)", 10, "〜でしょう/〜だろう (推測)", 11, "〜が好きです/嫌いです (好み)", 12, "〜と思います (意見・思考)", 13, "〜から/〜ので (理由)", 14, "〜のが好きです/嫌いです (動作の好み)", 15, "〜でしょうか (丁寧な質問)", 16, "〜てしまう (完了・後悔)", 17, "〜ながら (同時進行)", 18, "〜ば/〜たら (仮定形)", 19, "〜ておく (準備)", 20, "〜ようにする (努力・習慣)", 21, "〜そうだ (伝聞・推量)", 22… See the full description on the dataset page: https://huggingface.co/datasets/if001/bunpo_phi4.tabular100K<n<1M0 likes9 downloads1y agoHugging Face21open-llm-leaderboard /mkurman__phi4-MedIT-10B-o1-detailsgated Dataset Card for Evaluation run of mkurman/phi4-MedIT-10B-o1 Dataset automatically created during the evaluation run of model mkurman/phi4-MedIT-10B-o1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mkurman__phi4-MedIT-10B-o1-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face22open-llm-leaderboard /prithivMLmods__Phi4-Super-detailsgated Dataset Card for Evaluation run of prithivMLmods/Phi4-Super Dataset automatically created during the evaluation run of model prithivMLmods/Phi4-Super The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Phi4-Super-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face23open-llm-leaderboard /Undi95__Phi4-abliterated-detailsgated Dataset Card for Evaluation run of Undi95/Phi4-abliterated Dataset automatically created during the evaluation run of model Undi95/Phi4-abliterated The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Undi95__Phi4-abliterated-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face24open-llm-leaderboard /Triangle104__Phi4-RP-o1-detailsgated Dataset Card for Evaluation run of Triangle104/Phi4-RP-o1 Dataset automatically created during the evaluation run of model Triangle104/Phi4-RP-o1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Phi4-RP-o1-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face25open-llm-leaderboard /Quazim0t0__Phi4.Turn.R1Distill_v1.5.1-Tensors-detailsgated Dataset Card for Evaluation run of Quazim0t0/Phi4.Turn.R1Distill_v1.5.1-Tensors Dataset automatically created during the evaluation run of model Quazim0t0/Phi4.Turn.R1Distill_v1.5.1-Tensors The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Phi4.Turn.R1Distill_v1.5.1-Tensors-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face26Lakshan2003 /Phi-4-Mini-customerservice-Human-evaluator_1_datatabularn<1K0 likes7 downloads8mo agoHugging Face27open-llm-leaderboard /hotmailuser__Phi4-Slerp4-14B-detailsgated Dataset Card for Evaluation run of hotmailuser/Phi4-Slerp4-14B Dataset automatically created during the evaluation run of model hotmailuser/Phi4-Slerp4-14B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hotmailuser__Phi4-Slerp4-14B-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face28gx-ai-architect /numinamath-178k-phi4-bon-verified-dpo-trl-40k-old-r1-formattabular10K<n<100K0 likes6 downloads2y agoHugging Face29R0bfried /RAGAS-INSTRUCT-phi4-evaltabularn<1K0 likes6 downloads1y agoHugging Face30R0bfried /RAGAS-RAFT-phi4-evaltabularn<1K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.