ControlLLM/Llama-3.1-8B-SynE-FPT
015
Control-LLM-Llama3.1-8B-SynE-Full-Parameter-Tuning
This is a fine-tuned model of Llama-3.1-8B for muliligual-Chinese tasks on SynE dataset.
Linked Paper
This model is associated with the paper: Control LLM: Controlled Evolution for Intelligence Retention in LLM.
Linked Open Source code - training, eval and benchmark
This model is associated with the github: Control-LLM.
Evaluation Results
Here is an overview of the evaluation results and findings:
Benchmark Results Table
The table below summarizes evaluation results across Chinese tasks and original capabilities.
Explanation:
- CEval: Chinese Evaluation
- CEvalC: Chinese Evaluation (CoT - Chain of Thought)
- CMMLU: Chinese MMLU
- CMMLUC: Chinese MMLU (CoT)
- C-Avg: Chinese - Size Weighted Average across CEval, CEvalC, CMMLU, and CMMLUC
- BBH: BigBench Hard
- MLU: MMLU (Massive Multitask Language Understanding)
- MLUP: MMLU Pro
- O-Avg: Original Capability - Size Weighted Average across BBH, MLU, and MLUP
- Overall: Combined average across all tasks
Full Parameter Tuning on Chinese-SynE
The following plot illustrates the Catastrophic Forgetting of full parameter tuning in terms of hidden states alignment drift.
