CoolFace
Modelpublic

brittlewis12/Phi-3-mini-4k-instruct-GGUF

sourceHugging Facemitupdated 2y agoView on Hugging Face
2likes793downloads
Model Card

Phi 3 Mini 4K Instruct GGUF

*Updated with Microsoft’s [latest model changes](https://huggingface.co/microsoft/Phi-3-mini-4k-instruct/commit/4f818b18e097c9ae8f93a29a57027cad54b75304) as of July 21, 2024*

Original model: Phi-3-mini-4k-instruct

Model creator: Microsoft

This repo contains GGUF format model files for Microsoft’s Phi 3 Mini 4K Instruct.

The Phi-3-Mini-4K-Instruct is a 3.8B parameters, lightweight, state-of-the-art open model trained with the Phi-3 datasets that includes both synthetic data and the filtered publicly available websites data with a focus on high-quality and reasoning dense properties.

Learn more on Microsoft’s Model page.

What is GGUF?

GGUF is a file format for representing AI models. It is the third version of the format, introduced by the llama.cpp team on August 21st 2023. It is a replacement for GGML, which is no longer supported by llama.cpp. Converted with llama.cpp build 3432 (revision 45f2c19), using autogguf.

Prompt template

<|system|>
{{system_prompt}}<|end|>
<|user|>
{{prompt}}<|end|>
<|assistant|>

Download & run with cnvrs on iPhone, iPad, and Mac!

cnvrs.ai

cnvrs is the best app for private, local AI on your device:

  • —create & save Characters with custom system prompts & temperature settings
  • —download and experiment with any GGUF model you can find on HuggingFace!
  • —make it your own with custom Theme colors
  • —powered by Metal ⚡️ & Llama.cpp, with haptics during response streaming!
  • —try it out yourself today, on Testflight!
  • —follow cnvrs on twitter to stay up to date

Original Model Evaluation

Comparison of July update vs original April release:

BenchmarksOriginalJune 2024 Update
Instruction Extra Hard5.76.0
Instruction Hard4.95.1
Instructions Challenge24.642.3
JSON Structure Output11.552.3
XML Structure Output14.449.8
GPQA23.730.6
MMLU68.870.9
Average21.936.7

Original April release

As is now standard, we use few-shot prompts to evaluate the models, at temperature 0. The prompts and number of shots are part of a Microsoft internal tool to evaluate language models, and in particular we did no optimization to the pipeline for Phi-3. More specifically, we do not change prompts, pick different few-shot examples, change prompt format, or do any other form of optimization for the model. The number of k–shot examples is listed per-benchmark.
Phi-3-Mini-4K-In<br>3.8bPhi-2<br>2.7bMistral<br>7bGemma<br>7bLlama-3-In<br>8bMixtral<br>8x7bGPT-3.5<br>version 1106
MMLU <br>5-Shot68.856.361.763.666.568.471.4
HellaSwag <br> 5-Shot76.753.658.549.871.170.478.8
ANLI <br> 7-Shot52.842.547.148.757.355.258.1
GSM-8K <br> 0-Shot; CoT82.561.146.459.877.464.778.1
MedQA <br> 2-Shot53.840.949.650.060.562.263.4
AGIEval <br> 0-Shot37.529.835.142.142.045.248.4
TriviaQA <br> 5-Shot64.045.272.375.267.782.285.8
Arc-C <br> 10-Shot84.975.978.678.382.887.387.4
Arc-E <br> 10-Shot94.688.590.691.493.495.696.3
PIQA <br> 5-Shot84.260.277.778.175.786.086.6
SociQA <br> 5-Shot76.668.374.665.573.975.968.3
BigBench-Hard <br> 0-Shot71.759.457.359.651.569.768.32
WinoGrande <br> 5-Shot70.854.754.255.66562.068.8
OpenBookQA <br> 10-Shot83.273.679.878.682.685.886.0
BoolQ <br> 0-Shot77.6--72.266.080.977.679.1
CommonSenseQA <br> 10-Shot80.269.372.676.27978.179.6
TruthfulQA <br> 10-Shot65.0--52.153.063.260.185.8
HumanEval <br> 0-Shot59.147.028.034.160.437.862.2
MBPP <br> 3-Shot53.860.650.851.567.760.277.8