CoolFace
Modelpublic

JGOS-Model/JGOS-31B-Citizen

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
31likes38downloads
Model Card

JGOS-31B-Citizen

<p align="center"> <img src="https://huggingface.co/JGOS-Model/JGOS-31B-Citizen/resolve/main/k-ai.png" alt="#1 on the K-AI Leaderboard" width="780"/> </p>

<p align="center"> ๐Ÿ† <b>#1 on the K-AI Leaderboard</b> &middot; Korea's national Korean-language AI benchmark (<a href="https://leaderboard.aihub.or.kr/leaderboard">leaderboard.aihub.or.kr</a>) </p>

JGOS-31B-Citizen is a Korean, multimodal large language model specialized for administrative & public-sector AI services โ€” civil-complaint response, public-document understanding, and government-domain question answering.

Overview

JGOS-31B-Citizen is built on VIDRAFT's Darwin V8 platform.

  • โ€”Base + FFN transfer, breeding & evolution (Darwin V8). Starting from our in-house gemma4-31b base, the feed-forward network (FFN) blocks of multiple source models are extracted and grafted, then bred (merged) and evolved across multiple generations through the Darwin V8 pipeline to accumulate capability.
  • โ€”Korean administrative-domain fine-tuning. The evolved model is further trained on Korean-specialized datasets to strengthen Korean comprehension, reasoning, and administrative/public-sector domain performance.
The set of grafted source models, the number of evolution generations, the breeding strategy, dataset composition, and training configuration are proprietary and not disclosed.

Specifications

ItemValue
Parameters~31B (dense)
ModalityText + Image (multimodal)
Context lengthup to 256K tokens
Base familygemma4-31b (Gemma-compatible architecture)
FocusAdministrative & public-sector AI services

Highlights

  • โ€”๐Ÿ† #1 on the K-AI Leaderboard โ€” Korea's national Korean-language AI benchmark (KMMLU-Pro ยท CLIcK ยท HLE ยท MuSR ยท Com2)
  • โ€”GPQA Diamond: 84.34%

Evaluation

GPQA Diamond (198 questions)

Method (test-time compute)Score
maj@8 + tie-retry + DELPHI + near-miss maj@32-64 (weighted vote)84.34% (167/198)

Training Datasets

JGOS-31B-Citizen was trained using large-scale Korean corpora sourced from the Korean AI Hub (AIHub) โ€” Korea's national AI data repository operated by NIA. The following datasets were used to optimize performance on the K-AI Leaderboard benchmarks (KoMMLU-Pro, CLIcK, HLE, MuSR, Com2):

#Dataset NameAIHub Link
1Medical and Legal Professional Books Corpus71487
2Financial and Legal Document Machine Reading Comprehension71610
3Large-scale Web-based Korean Corpus624
4Large-scale Book-based Korean Corpus653
5National Records Large-scale AI Learning Corpus71788
6Korean Generation-based Common Sense Reasoning Dataset459
7Multi-session Dialogue Corpuspkg1
8Essential Medical Knowledge Data (142GB)71875
9Specialized Medical Knowledge Data (206GB)71874
10Korean Dialogue Dataset272
All datasets are publicly available via AIHub (registration required).

License

This model is built on a Gemma-family architecture and is distributed under the **Gemma Terms of Use**. By using this model, you agree to the Gemma license terms.