Abhayn01/arka-advanced-slm-sft-v1
0234
ARKA Advanced SLM — Mixed SFT v1
ARKA stands for Academic Research Knowledge Assistant.
Creator/developer: Abhay Kumar Rudrapaul.
Base
Abhayn01/arka-advanced-slm-scitech-stream-500m-v1
Architecture / lineage information
- Decoder-only causal language model
- Approximately 124.4M parameters
- GPT-2 tokenizer vocabulary: 50,257
- Context length used for this SFT: 1024
- Full-model SFT; transformer layers were not frozen
- SFT learning rate: 5e-06
SFT mixture
- MMLU auxiliary training questions
- SQuAD 2.0
- OpenBookQA
- AI2 ARC Challenge and Easy
- SciQ
- CounterFact factual/counterfactual relations (true target is supervised; false target is never used as the answer)
- ARKA identity/model information
- Creator profile examples explicitly requested by the creator
- ChatML-style, Alpaca-style, QA, and context-guided serializations
Train examples: 160,863 Validation examples: 4,000
Creator profile included by request
- Name: Abhay Kumar Rudrapaul
- B.Tech in Electrical Engineering
- Tripura Institute of Technology, Narsingarh
- Agartala, Tripura, India
- Interests include AI, machine learning, technology, language-model development, and applied AI systems
- Creator/developer of ARKA and TIT Campus Assistant–ARKA
No passwords, account credentials, private IDs, or health information are included.
Notes
The GPT-2 tokenizer vocabulary is kept unchanged. ChatML markers are learned as ordinary text sequences rather than adding new tokenizer special tokens.
Review each upstream dataset's license/terms before redistributing transformed data. This model repository records the source names but does not upload transformed copies of those source datasets.
