CoolFace
Modelpublic

ethicalabs/xLSTM-7b-Polymath

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes102downloads
Model Card

Model Card for xlstm-7b-instruct-phase-2

[!WARNING] This repository contains experimental models designed strictly for academic evaluation and research purposes. Critical Constraints: No Production Deployment: Experimental models must not be deployed in commercial, enterprise, or mission-critical environments under any circumstances. No Liability: Experimental models are provided "as-is" without warranties of any kind. The developers assume zero liability for downstream consequences, system integration failures, or regulatory non-compliance resulting from unauthorized deployment.

This model is a fine-tuned version of ethicalabs/xLSTM-7b-Instruct for task alignment.

It has been trained using TRL using SFT on assistant-only tokens.

The k_proj and v_proj matrices have been frozen to isolate and preserve the model's pre-trained knowledge base.

This fine-tuning focused only on the q_proj (query) and FFN matrices, adapting the model's reasoning and query-retrieval mechanisms without overwriting its core, frozen knowledge.

This experiment was designed to test the hypothesis that the model's reasoning capabilities (q_proj) could be specialized for math/code while its knowledge (k_proj, v_proj) remained intact.

Quick start

Work in Progress!

Training procedure

<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>

This model was trained with SFT.

Evaluation

This model has been loaded in 4-bit and evaluated with lighteval

TaskVersionMetricValueStderr
allacc0.5383±0.1476
acc:logprobnormalization=LogProbCharNorm(name='norm', ignorefirst_space=True)0.7000±0.1528
acc:logprobnormalization=LogProbCharNorm(name='norm', ignorefirst_space=False)0.8000±0.1333
truthfulqa_mc10.6000±0.1633
truthfulqa_mc20.7066±0.1481
em:normalizegold=<function gsm8knormalizer at 0x7c5d972c3ba0>&normalizepred=<function gsm8knormalizer at 0x7c5d972c3ba0>0.6000±0.1633
leaderboard:arc:challenge:25acc0.8000±0.1333
acc:logprobnormalization=LogProbCharNorm(name='norm', ignorefirst_space=True)0.7000±0.1528
leaderboard:gsm8k:5em:normalizegold=<function gsm8knormalizer at 0x7c5d972c3ba0>&normalizepred=<function gsm8knormalizer at 0x7c5d972c3ba0>0.6000±0.1633
leaderboard:hellaswag:10acc0.5000±0.1667
acc:logprobnormalization=LogProbCharNorm(name='norm', ignorefirst_space=False)0.8000±0.1333
leaderboard:mmlu:_average:5acc0.5316±0.1474
leaderboard:mmlu:abstract_algebra:5acc0.3000±0.1528
leaderboard:mmlu:anatomy:5acc0.3000±0.1528
leaderboard:mmlu:astronomy:5acc0.7000±0.1528
leaderboard:mmlu:business_ethics:5acc0.4000±0.1633
leaderboard:mmlu:clinical_knowledge:5acc0.7000±0.1528
leaderboard:mmlu:college_biology:5acc0.5000±0.1667
leaderboard:mmlu:college_chemistry:5acc0.4000±0.1633
leaderboard:mmlu:collegecomputerscience:5acc0.4000±0.1633
leaderboard:mmlu:college_mathematics:5acc0.2000±0.1333
leaderboard:mmlu:college_medicine:5acc0.5000±0.1667
leaderboard:mmlu:college_physics:5acc0.5000±0.1667
leaderboard:mmlu:computer_security:5acc0.9000±0.1000
leaderboard:mmlu:conceptual_physics:5acc0.4000±0.1633
leaderboard:mmlu:econometrics:5acc0.4000±0.1633
leaderboard:mmlu:electrical_engineering:5acc0.7000±0.1528
leaderboard:mmlu:elementary_mathematics:5acc0.3000±0.1528
leaderboard:mmlu:formal_logic:5acc0.3000±0.1528
leaderboard:mmlu:global_facts:5acc0.3000±0.1528
leaderboard:mmlu:highschoolbiology:5acc0.9000±0.1000
leaderboard:mmlu:highschoolchemistry:5acc0.5000±0.1667
leaderboard:mmlu:highschoolcomputer_science:5acc0.6000±0.1633
leaderboard:mmlu:highschooleuropean_history:5acc0.7000±0.1528
leaderboard:mmlu:highschoolgeography:5acc1.0000±0.0000
leaderboard:mmlu:highschoolgovernmentandpolitics:5acc0.8000±0.1333
leaderboard:mmlu:highschoolmacroeconomics:5acc0.6000±0.1633
leaderboard:mmlu:highschoolmathematics:5acc0.3000±0.1528
leaderboard:mmlu:highschoolmicroeconomics:5acc0.7000±0.1528
leaderboard:mmlu:highschoolphysics:5acc0.3000±0.1528
leaderboard:mmlu:highschoolpsychology:5acc0.9000±0.1000
leaderboard:mmlu:highschoolstatistics:5acc0.5000±0.1667
leaderboard:mmlu:highschoolus_history:5acc0.8000±0.1333
leaderboard:mmlu:highschoolworld_history:5acc0.9000±0.1000
leaderboard:mmlu:human_aging:5acc0.5000±0.1667
leaderboard:mmlu:human_sexuality:5acc0.4000±0.1633
leaderboard:mmlu:international_law:5acc0.6000±0.1633
leaderboard:mmlu:jurisprudence:5acc0.6000±0.1633
leaderboard:mmlu:logical_fallacies:5acc0.4000±0.1633
leaderboard:mmlu:machine_learning:5acc0.5000±0.1667
leaderboard:mmlu:management:5acc0.5000±0.1667
leaderboard:mmlu:marketing:5acc0.8000±0.1333
leaderboard:mmlu:medical_genetics:5acc0.9000±0.1000
leaderboard:mmlu:miscellaneous:5acc0.5000±0.1667
leaderboard:mmlu:moral_disputes:5acc0.7000±0.1528
leaderboard:mmlu:moral_scenarios:5acc0.1000±0.1000
leaderboard:mmlu:nutrition:5acc0.6000±0.1633
leaderboard:mmlu:philosophy:5acc0.5000±0.1667
leaderboard:mmlu:prehistory:5acc0.4000±0.1633
leaderboard:mmlu:professional_accounting:5acc0.3000±0.1528
leaderboard:mmlu:professional_law:5acc0.4000±0.1633
leaderboard:mmlu:professional_medicine:5acc0.2000±0.1333
leaderboard:mmlu:professional_psychology:5acc0.3000±0.1528
leaderboard:mmlu:public_relations:5acc0.3000±0.1528
leaderboard:mmlu:security_studies:5acc0.3000±0.1528
leaderboard:mmlu:sociology:5acc0.8000±0.1333
leaderboard:mmlu:usforeignpolicy:5acc0.7000±0.1528
leaderboard:mmlu:virology:5acc0.5000±0.1667
leaderboard:mmlu:world_religions:5acc0.8000±0.1333
leaderboard:truthfulqa:mc:0truthfulqa_mc10.6000±0.1633
truthfulqa_mc20.7066±0.1481
leaderboard:winogrande:5acc0.7000±0.1528

Framework versions

  • PEFT 0.17.1
  • TRL: 0.24.0
  • Transformers: 4.57.1
  • Pytorch: 2.8.0+cu126
  • Datasets: 4.2.0
  • Tokenizers: 0.22.1

Citations

Cite TRL as:

bibtex
@misc{vonwerra2022trl,
	title        = {{TRL: Transformer Reinforcement Learning}},
	author       = {Leandro von Werra and Younes Belkada and Lewis Tunstall and Edward Beeching and Tristan Thrush and Nathan Lambert and Shengyi Huang and Kashif Rasul and Quentin Gallou{\'e}dec},
	year         = 2020,
	journal      = {GitHub repository},
	publisher    = {GitHub},
	howpublished = {\url{https://github.com/huggingface/trl}}
}