datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magpie-Phi3-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/Magpie-Phi3-Pro-1M-v0.1.medical-logits-phi3.5-mini_medmcqa_pubmMagpie-Phi3-Pro-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Phi3-Pro-300K-Filtered.Magpie-Phi3-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Phi3-Pro-1M-v0.1.phi_30K_qwq_0K_eval_2e29
mlfoundations-dev/phi_30K_qwq_0K_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
26.7
42.2
44.8
8.1
32.0
47.1
0.8
0.3
0.1
20.7
2.3
0.3
AIME24
Average Accuracy: 26.67% ± 1.63%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
26.67%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_30K_qwq_0K_eval_2e29.Magpie-Phi3-Pro-300K-Filtered-ko-e3medical_dirichlet_phi3This dataset was generated using: https://github.com/sarus-tech/dp-llm-ft .
Patients, diseases, symptoms lists, treatments are all generated using a local Phi 3.5 (microsoft/Phi-3.5-mini-instruct).
The diseases are sampled independently for each patient, using a discrete distribution of diseases sampled from a Dirichlet distribution.
primekg-phi3-distilled-sft-seed7phi_30K_qwq_0K_temp2management_rule_phi3.5_unsupprimekg-phi3-distilled-sft-seed999economy_rule_phi3.5_unsupprimekg-phi3-distilled-sft-seed2024Edge-Industrial-Anomaly-Phi3
Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs
This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification.
🚀 Purpose
Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.convsersations_excitement_phi3.5-mini-instruct_largeprimekg-phi3-distilled-sft-seed2024primekg-phi3-distilled-sft-seed101primekg-phi3-distilled-sft-seed7accounting_rule_phi3.5_unsupprimekg-phi3-vanillakd-sft-seed2024primekg-phi3-vanillakd-sft-seed7Magpie-Phi3-Pro-300K-Filtered-koTranslate Magpie-Align/Magpie-Phi3-Pro-300K-Filtered using nayohan/llama3-instrucTrans-enko-8b.
This is a raw translation dataset. It needs to be filtered for repetitions generated by the model.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin},
year={2024},
eprint={2406.08464}… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/Magpie-Phi3-Pro-300K-Filtered-ko.phi3-arena-short-dpo
Dataset Summary
DPO (Direct Policy Optimization) dataset of normal and short answers generated from lmsys/chatbot_arena_conversations dataset using microsoft/Phi-3-mini-4k-instruct model.
Generated using ShortGPT project.
magnifi__Phi3_intent_v56_3_w_unknown_5_lr_0.002-details
Dataset Card for Evaluation run of magnifi/Phi3_intent_v56_3_w_unknown_5_lr_0.002
Dataset automatically created during the evaluation run of model magnifi/Phi3_intent_v56_3_w_unknown_5_lr_0.002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/magnifi__Phi3_intent_v56_3_w_unknown_5_lr_0.002-details.phi-3-vision-custom-datasetmagpie_phi3-mini_spanishScienticDatasetArxiv-phi3-Format
Scientific Dataset Repository
Este repositorio contiene un conjunto de datos de conversaciones estructuradas, diseñado para ser utilizado en modelos de deep learning. El conjunto de datos está organizado y preparado para su uso en la plataforma Huggingface.
Información del Conjunto de Datos
El conjunto de datos incluye las siguientes características:
conversations:
from: Tipo de dato: string. Indica el origen de la conversación.
value: Tipo de dato: string. Contenido de… See the full description on the dataset page: https://huggingface.co/datasets/ejbejaranos/ScienticDatasetArxiv-phi3-Format.function_call_orpo_sft_phi3_chatML_only
take dataset
hiyouga/glaive-function-calling-v2-sharegpt
load code.
from datasets import load_dataset
# unsloth_Phi_3_mini_4k_instruct_unsloth_oasst2_orpo_mix_tokenizer_phi_3_v1_0001
import os
dataset = load_dataset('NickyNicky/function_call_orpo_sft_phi3_only','default')
# sft
data_formater=data_formater.remove_columns(['prompt','chosen','rejected'])
# orpo-dpo
data_formater=data_formater.remove_columns(['Text'])
sft Column 'Text'.
You are a helpful AI… See the full description on the dataset page: https://huggingface.co/datasets/NickyNicky/function_call_orpo_sft_phi3_chatML_only.Phi3_intent_v37_2_wo_unknownruslanmv_medicalChat_phi3.5_instruct
Use this dataset
from datasets import load_dataset
data = load_dataset("syubraj/ruslanmv_medicalChat_phi3.5_instruct")
Source : ruslanmv/ai-medical-chatbot
