CoolFace
Modelpublic

InferenceIllusionist/Excalibur-7b

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
9likes31downloads
Model Card

Excalibur-7b

<img src="https://i.imgur.com/viIO4WT.png" width="550"/>

<b> Update: A fine-tuned version of this model is now publicly available, along with benchmark results. If you're looking for a more conversational, assistant-style exchange you won't want to miss it!</b>

<i>Image generated with Envoid's Model9 SDXL model </i>

GGUFs can be found here

Alternative GGUFs from bartowski can be found here.

EXl2 can also be found here again courtesy of bartowski!

Performance Comparison

NameAvg.ARCHellaSwagMMLUTruthfulQAWinograndeGSM8K
<b>Excalibur-7b</b><u><b>73.6</b></u><u><b>69.71</b></u><u><b>87.56</b></u><u><b>65.66</b></u><u><b>67.24</b></u><u><b>82.79</b></u><u><b>68.61</b></u>
Magic-Dolphin-7b67.4865.7885.6164.6458.0179.6451.18
merlinite-7b6463.6584.5264.9150.1579.7241.09

* Open LLM Leaderboard Dataset

Methodology

Magic-Dolphin-7b was an unexpected surprise. Profoundly satisfied with it as a first attempt. For this follow-up I wanted to target the MMLU benchmark specifically. The challenge this time was placing more weight on Merlinite-7b as an unknown quantity that hasn't been in the spotlight despite its novel LAB tuning method.

<b>Excalibur-7b</b> builds on past success and is the culmination of several learnings:

  • —Measuring KL-divergences for new quantization types brought a deeper understanding of benchmarking and assessing model performance
  • —This signifcantly sped up the testing process by using MMLU as a base, narrowing down over 10 candidate linear merges to 1: merliniteX-blockB1
  • —Reaching the limitations of linear merging necessitated a pivot to reviewing the viability of SLERP, DARE-TIES, and Passthrough methods
  • —Thus a competing candidate merge pool was tested between different merge algorithms. Once more the list was narrowed from 10 candidates to 1: merliniteX-blockF2
  • —merliniteX-blockF2 (SLERP of Magic-Dolphin-7B and jaskier-7b-dpo in unorthadox proportions) was originally planned for release earlier this week
  • —Instead -blockB1 and -blockF2 were merged and the results were placed head to head in a final round of tests. Ultimately a more conventional execution of SLERP showed the best results for the final step.

Sample Question

<img src="https://i.imgur.com/fdFYIhv.jpeg" width="550"/>

Bonus Question - Vision Capabilities

<b>Requires additional mistral-7b-mmproj-v1.5-Q4_1.gguf file for vision functionality</b> <img src="https://i.imgur.com/4wbUrjf.jpeg" width="550"/>

Select up the gguf file of your choice in Kobold as usual, then make sure to choose the mmproj file above in the LLaVA mmproj field of the model submenu: <img src="https://i.imgur.com/x8vqH29.png" width="550"/>

This is a merge of pre-trained language models created using mergekit.

Merge Details

Merge Method

This model was merged using the SLERP merge method.

Models Merged

The following models were included in the merge:

Configuration

The following YAML configurations were used to produce this model:

<b>merliniteX-blockB1</b>

yaml
models:
  - model: models/merlinite-7b
    parameters:
      weight: 1.0
  - model: models/Kunoichi-DPO-v2-7B
    parameters:
      weight: 0.2
  - model: models/jaskier-7b-dpo-v6.1
    parameters:
      weight: 0.6
  - model: models/Monarch-7b
    parameters:
      weight: 0.4
merge_method: linear
dtype: float16

<b>merliniteX-blockF2</b>

yaml
slices:
  - sources:
      - model: models/Magic-Dolphin-7b
        layer_range: [0, 32]
      - model: models/jaskier-7b-dpo-v6.1
        layer_range: [0, 32]
merge_method: slerp
base_model: models/Magic-Dolphin-7b
parameters:
  t:
    - filter: self_attn
      value: [0, 0.5, 0.3, 0.7, 0.5, 1]
    - filter: mlp
      value: [1, 0.5, 0.7, 0.3, 0.5, 0]
    - value: 0.5 # fallback for rest of tensors
dtype: float16

<b>merliniteX-blockH1 (Excalibur-7b)</b>

yaml
slices:
  - sources:
      - model: models/merliniteX-blockF2
        layer_range: [0, 32]
      - model: models/merliniteX-blockB1
        layer_range: [0, 32]
merge_method: slerp
base_model: models/merliniteX-blockF2
parameters:
  t:
    - filter: self_attn
      value: [1, 0.7, 0.3, 0.5, 0]
    - filter: mlp
      value: [0, 0.3, 0.7, 0.5, 1]
    - value: 0.5 # fallback for rest of tensors
dtype: float16