CoolFace
Modelpublic

punsaisuwan/frankenmoe-python-typescript

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes
Model Card

FrankenMoE: Multi-Expert LoRA System (Python + TypeScript + Reasoning)

Triple LoRA adapters ที่ fine-tune บน 4-bit quantized MLX version ของ mistralai/Mistral-Small-3.1-24B-Instruct-2503 แยกตามความเชี่ยวชาญ พร้อม Orchestrator ที่รองรับ Weighted Routing, LLM Fallback Classification, Streaming Output, และ Multi-layer Safety Guard สำหรับป้องกัน Repetition Loop และ Persona Drift

v2 Update: เพิ่ม Expert 3 (Reasoning + Thai Translation Pipeline), เปลี่ยนจาก frankenmoe_dispatcher.py เดิมมาเป็น moe_orchestrator.py ที่มี Streaming Support และ Safety Net ที่ผ่านการทดสอบ Multi-turn มากกว่า 15 Turns โดยไม่พบ Repetition Loop

Experts

ExpertLanguage/DomainRankTarget ModulesNotes
expert-1-pythonPython16q,k,v,o_projแก้ปัญหา rank/capacity ไม่พอจาก v1
expert-2-typescriptTypeScript8default (incl. MoE switch layers)แก้ NaN loss + Catastrophic Interference
expert-3-reasoningGeneral Reasoning (EN to TH)-defaultมี Auto-translate pipeline เป็นภาษาไทย

Key Features

  • —Weighted + LLM Fallback Routing: ระบบเลือก Expert ด้วย Keyword Scoring ก่อน ถ้ากำกวมจะใช้ Base Model เป็น Classifier ช่วยตัดสิน
  • —Streaming Output: Python/TypeScript Expert รองรับ Token-by-token Streaming ผ่าน query_stream() เพื่อลด Perceived Latency
  • —Repetition Guard: ตรวจจับ Generation Loop ด้วย Sliding Window (ทุก ~50 Token ระหว่าง Stream) และ Retry-once Policy สำหรับ Non-stream
  • —Persona Enforcement: บังคับสรรพนามและคำลงท้ายให้สม่ำเสมอ ("ผม" / "ครับ") ตลอดทุก Turn
  • —Auto Thai Translation: Reasoning Expert แปลผลลัพธ์เป็นไทยอัตโนมัติ พร้อม Code-block Preservation (ป้องกันโค้ดถูกแปลผิดเพี้ยน)

Usage

ต้องมี Base Model 4-bit อยู่ในเครื่อง:

pip install mlx-lm huggingfacehub[hfxet] huggingface-cli download --local-dir frankenmoe-python-typescript punsaisuwan/frankenmoe-python-typescript

Interactive Chat (Streaming)

python moe_orchestrator.py

Python API

from moe_orchestrator import MoEOrchestrator

orchestrator = MoEOrchestrator( basemodelpath="./output-moe-mlx-4bit", adapter_paths={ "python": "./adapters/expert-1-python", "typescript": "./adapters/expert-2-typescript", "reasoning": "./adapters/expert-3-reasoning", } )

Non-streaming (สำหรับ Automated Test / Batch)

response = orchestrator.query("เขียนฟังก์ชัน fibonacci ด้วย Python") print(response)

Streaming (สำหรับ Interactive Use)

for chunk in orchestrator.query_stream("เขียนฟังก์ชัน fibonacci ด้วย Python"): print(chunk, end="", flush=True)

Training Details

  • —Expert 1 (Python): อัปเกรดจาก v1 (rank=8, broken) เป็น rank=16, scale=8.0, ครอบคลุม 4 attention projections เพื่อแก้ปัญหา capacity ไม่พอสำหรับ code generation ที่ซับซ้อน
  • —Expert 2 (TypeScript): แก้ NaN loss ผ่าน SAFE_PAD_TOKEN_ID=198 และแก้ Catastrophic Interference ระหว่าง function/type guard bias ด้วย dataset rebalancing + ลด learning rate เป็น 1.5e-5
  • —Expert 3 (Reasoning): Train แยกจาก Code Experts เพื่อป้องกัน Cross-domain Interference พร้อม Pipeline แปลไทยที่แยก Code Block ออกก่อนแปล (ป้องกันโค้ดถูกแปลผิด) แล้วค่อยประกอบกลับ

Safety and Stability Notes

ระบบผ่านการทดสอบ Multi-turn Conversation 15 Turns โดยไม่พบ Repetition Loop และไม่ต้องพึ่ง Safety Net Fallback แม้แต่ครั้งเดียว (REPETITION_PENALTY=1.15 เป็นค่าที่ผ่านการ Tune แล้วว่าดีที่สุด)

ปัญหาเล็กน้อยที่ทราบอยู่ (Known Minor Issues):

  • —Code fence อาจปิดไม่ครบในบาง Edge Case
  • —พบ CJK character leak ในระดับต่ำมาก ไม่กระทบความเข้าใจ

Requirements

mlx-lm huggingfacehub[hfxet]

License

Apache 2.0 (ตาม Base Model)