datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phi_27K_qwq_3K_lambda_0.9phi2_rejection_sampling
Phi-2 Rejection Sampling
The Phi-2 Rejection Sampling dataset is an English-language dataset consisting of 10 prompts and responses generated by Phi-2 and graded by the OpenAssistant's reward model.
Dataset Details
Dataset Description
The Phi-2 Rejection Sampling dataset is a small (n = 10) English-language dataset. This dataset was created with the purpose was to demonstrate a feedback pipeline where in which Phi-2 would interact with the OpenAssistant reward… See the full description on the dataset page: https://huggingface.co/datasets/BluefinTuna/phi2_rejection_sampling.phi2-i0-v2ActionRoutes_Phi2_ZeroShotmicrosoft-phi2-mental-healthphi_24K_qwq_6K_eval_2e29
mlfoundations-dev/phi_24K_qwq_6K_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
35.0
66.2
80.8
8.0
41.9
46.6
1.2
3.2
2.0
27.7
3.6
0.5
AIME24
Average Accuracy: 35.00% ± 2.22%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
36.67%
11
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_24K_qwq_6K_eval_2e29.phi-2-labeled
Dataset Card for "phi-1"
More Information needed
phi-2-symbol-100k-en-512phi_27K_qwq_3KMedQuad-phi2-1kphi-2-symbol-100k-en-2048phi2-spl-evaluationphi-2-embeddings
Dataset Card for "phi-2-embeddings"
More Information needed
10k_prompts_SPIN_iter1_phi2_top
Dataset Card for "10k_prompts_SPIN_iter1_phi2_top"
More Information needed
10k_prompts_SPIN_iter0_phi2_top
Dataset Card for "10k_prompts_SPIN_iter0_phi2_top"
More Information needed
phi2-alignment
Dataset Card for Phi-2 Alignment
Dataset Description
Creator: Mahdi Ranjbar
Notebook: Colab Notebook
Dataset Summary
This dataset was developed for the alignment internship take-home assignment with the goal of assessing alignment
across three dimensions: Helpfulness, Honesty, and Harmlessness (3H). I crafted 10 prompts covering various tasks
and generated 8 answers for each prompt using the new Microsoft Phi-2 model.
Subsequently, I employed the Open… See the full description on the dataset page: https://huggingface.co/datasets/mahdi-ranjbar/phi2-alignment.phi_27K_qwq_3K_lambda_0.9phi-2-symbol-100kPHI2-MEDQUAD-16407-CLphi2-PIIActionRoutes_Phi2_FewShotphi-2-symbol-100k-enphi_24K_qwq_6Kphi_27K_qwq_3K_eval_2e29
mlfoundations-dev/phi_27K_qwq_3K_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
31.3
57.0
63.4
7.4
42.6
47.1
1.2
0.5
0.1
23.7
2.9
0.6
AIME24
Average Accuracy: 31.33% ± 1.26%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
30.00%
9
30
2
26.67%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_27K_qwq_3K_eval_2e29.PHI2-MEDQUAD-16407phi2_finetune_datasetActionRoutes_Phi2_Instructphi2phi2-tool-selectionphi-2_DEVELOPER-CREDITS_FOR-CUSTOM
