CoolFace
Modelpublic

KISTI-KONI/KONI-4B-base-20250819

sourceHugging Facegemmaupdated 1y agoView on Hugging Face
3likes13downloads
Model Card

KISTI-KONI/KONI-4B-base-20250819

Model Description

KONI (KISTI Open Neural Intelligence) is a large language model developed by the Korea Institute of Science and Technology Information (KISTI). Designed specifically for the scientific and technological domains, KONI excels in both Korean and English, making it an ideal tool for tasks requiring specialized knowledge in these areas.


Key Features

  • —Bilingual Model: Supports both Korean and English, with a focus on scientific and technical texts.
  • —Continual Pretraining: The model is continual pretrained on a filtered, high-quality bilingual corpus that includes scientific data and publicly available resources. This ensures adaptability to evolving scientific and technological content.
  • —Base Model: Built upon google/gemma-3-4b-pt, KONI-4B-base undergoes continual pretraining for superior performance on both general LLM benchmark and scientific benchmarks.
  • —Training Environment: Trained on 24 H200 GPUs at the KISTI supercomputer, optimizing both speed and quality during development.
  • —Corpus: Utilizes a high-quality, filtered corpus of OOO B tokens, comprising scientific texts as well as publicly available bilingual data.
  • —Data Optimization: The continual pretraining process involved testing a variety of data distributions (balanced, reasoning-enhanced, knowledge-enhanced, minimal Korean settings, etc.), followed by selection of the optimal combination for training.
  • —Enhanced Performance: KONI-4B-base delivers excellent performance, especially compared to other 4B-sized pretrained models, even though it is not instruction-tuned. An instruction-tuned version is expected soon, which will further improve its performance.

Model Performance

KONI-4B-base has demonstrated strong performance on a variety of scientific benchmarks, outperforming several other 4B-sized pretrained models. Here is a comparison of KONI-4B-base’s performance across various benchmarks including scientific and technological benchmarks:

RankModelKMMLUKMMLU-HardKoBESTkormedmcqaMMLUARC_easyARC_challengeHellaswagScholarBench-MCAidaBench-MCaverage
1Qwen/Qwen3-8B0.55000.29000.78000.37500.74000.87000.64000.57000.70940.73140.6256
2kakaocorp/kanana-1.5-8b-base0.48000.25000.62000.59100.63000.83000.56000.60000.68000.75480.5996
3LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct0.47000.23000.59000.53100.65000.83000.59000.62000.69000.70570.5907
4kakaocorp/kanana-1.5-2.1b-instruct-25050.42000.21000.77000.52240.55000.80000.53000.51000.66300.66880.5644
5KISTI-KONI/KONI-4B-base-202508190.43000.21000.73000.48000.58000.82000.52000.57000.68000.61470.5635
6KISTI-KONI/KONI-Llama3.1-8B-Instruct-202410240.40000.20000.56000.49050.63000.83000.54000.61000.69800.67220.5631
7LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct0.43000.21000.74000.48420.59000.77000.50000.54000.69000.65110.5605
8google/gemma-3-4b-pt0.39800.19980.69900.47260.59640.83000.54350.57630.66700.58860.5571
9meta-llama/Llama-3.1-8B-Instruct0.40000.20000.70000.47890.65000.84000.54000.61000.69600.67090.5786
10google/gemma-3-4b-it0.39000.21000.72000.44000.58000.84000.56000.56000.69900.60130.5600
11saltlux/Ko-Llama3-Luxia-8B0.38000.21000.71000.43200.55000.80000.48000.56000.66500.61090.5398
12MLP-KTLim/llama-3-Korean-Bllossom-8B0.37000.22000.55000.41630.64000.84000.57000.59000.65250.58620.5435
13kakaocorp/kanana-1.5-2.1b-base0.39000.24000.62000.51380.47000.73000.44000.45000.65000.64780.5152
14mistralai/Mistral-7B-v0.30.37000.22000.63000.37350.62000.83000.55000.62000.54400.42570.5183
15naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B0.39000.24000.64000.35500.47000.73000.44000.45000.59500.54500.4855
16naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-0.5B0.37000.22000.62000.33830.44000.72000.39000.41000.56000.51730.4586
17google/gemma-3-1b-it0.30690.24000.35560.27610.39700.66200.34300.42040.57200.39720.3776
18google/gemma-3-1b-pt0.25820.24560.55690.19640.26410.71460.35410.47030.21920.19800.3477
19etri-lirs/eagle-3b-preview0.16000.21000.51000.18040.25000.57000.24000.37000.26780.22240.2981

As shown, KISTI-KONI/KONI-4B-base-20250819 is the top-performing model in the 4B-size pretrained model category, outstanding google/gemma-3-4b-pt and KISTI-KONI/KONI-Llama3.1-8B-Instruct. While this version is not instruction-tuned, it offers exceptional results in scientific and technological domains. An upcoming instruction-tuned version is expected to further enhance its performance. Stay tuned for updates!


Strengths & Use Cases

  • —Domain-Specific Excellence: KONI-4B-base excels at tasks involving scientific literature, technological content, and complex reasoning. It is ideal for research, academic analysis, and specialized problem-solving.
  • —Bilingual Advantage: The model’s bilingual nature enables handling diverse datasets and generating high-quality responses in both English and Korean, especially in bilingual scientific collaborations.
  • —Benchmark Performance: KONI-4B-base has shown superior performance in benchmarks such as KMMLU, kormedmcqa, and ScholarBench-MC, proving its robustness in knowledge-intensive tasks.

Citation

If you use this model in your work, please cite it as follows:

bibtex
@article{KISTI-KONI/KONI-4B-base-20250819,
  title={KISTI-KONI/KONI-4B-base-20250819},
  author={KISTI},
  year={2025},
  url={https://huggingface.co/KISTI-KONI/KONI-4B-base-20250819}
}

Acknowledgements

  • —This research was supported by the Korea Institute of Science and Technology Information (KISTI) in 2025 (No. (KISTI) K25L1M1C1), aimed at developing KONI (KISTI Open Neural Intelligence), a large language model specialized in science and technology.
  • —This work also benefited from the resources and technical support provided by the National Supercomputing Center (KISTI).

References

  • —https://huggingface.co/google/gemma-3-4b-pt