datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Phi-4-mini-reasoning-secalignphi-4-mini-instruct-ultrafeedbackPhi-4-Mini-customerservice-Human-evaluator_2_dataPhi-4-mini-instruct-generationsPhi-4-mini-customerservice-context-summarization-llm-judge-data
customer-service-context-summarization-evaluation-data
Phi-4-mini-customerservice-context-summarization-llm-judge-data
Dataset updated with context summarization evaluation columns.
This README refresh triggers Hugging Face metadata re-index.
local-code-arena-mbpp-phi4-mini
Local Code Arena Telemetry: MBPP Benchmark on Phi-4 Mini
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against Microsoft's Phi-4 Mini (3.8B) model.
This run establishes a vital cross-vendor reference point, documenting how high-density synthetic reasoning filtration scales relative to dedicated code-only specialists on local consumer hardware.… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-phi4-mini.Phi4-Mini-P2T-4B-TestingTesting Results for USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B)
Phi4-Mini-Prose2Tags-4B-Raw-Training-DataRaw data used to train USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B)
Phi-4-Mini-customerservice-Human-evaluator_1_dataPhi-4-mini-customerservice-LLM-as-a-judge-dataPhi-4-mini-customerservice-Human-eval-dataPhi-4-Mini-customerservice-Human-evaluator_3_dataPhi-4-mini-customerservice-evaldataPhi-4-Mini-cost-benchmark-resultsPhi-4-mini-customerservice-context-summarization-evaldatasumm_phi4mini_32k_8.3K_valcorejur_testeagente_phi4-mini
