CoolFace
Modelpublic

sovasoft/zora-v1.13-mlx-q8

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
0likes400downloads
Model Card

๐ŸŒ EN ยท ๐Ÿ‡ท๐Ÿ‡ธ SR ยท ๐Ÿ‡ญ๐Ÿ‡ท HR ยท ๐Ÿ‡ง๐Ÿ‡ฆ BS ยท ๐Ÿ‡ฒ๐Ÿ‡ฐ MK ยท ๐Ÿ‡ธ๐Ÿ‡ฎ SL ยท ๐Ÿ‡ฆ๐Ÿ‡ฑ SQ ยท ๐Ÿ‡ฒ๐Ÿ‡ช CNR ยท ๐Ÿ‡ง๐Ÿ‡ฌ BG ยท ๐Ÿ‡ฌ๐Ÿ‡ท EL ยท ๐Ÿ‡น๐Ÿ‡ท TR ยท ๐Ÿ‡ท๐Ÿ‡ด RO ยท ๐Ÿ‡ญ๐Ÿ‡บ HU

<p align="center"> <img src="images/ZoraStich1700SWPrint.png" width="520" alt="Zora โ€” Goddess of Dawn, surrounded by the symbols of 12 peoples"/> </p>

Zora v1.13 โ€” an open, honest LLM for the Balkans & Southeast Europe

ะทะพั€ะฐ = "dawn". One to unite them all. โ€” by Sovasoft (ai.in.rs)


1 ยท What Zora is

Zora is an open 8B language model (built on Qwen3-8B) for 12 languages of the Balkans and Southeast Europe: Serbian, Croatian, Bosnian, Macedonian, Slovenian, Albanian, Montenegrin, Bulgarian, Greek, Turkish, Romanian, Hungarian.

Zora is not built to be the biggest model โ€” it is built to be honest, in-language, and multi-perspective:

  • โ€”thinks in the target language instead of pivoting through English,
  • โ€”shows several perspectives on contested topics instead of one national view,
  • โ€”and above all: admits when it doesn't know instead of inventing facts.

2 ยท The development story (v1.0 โ†’ v1.1 โ†’ v1.11 โ†’ v1.12 โ†’ v1.13)

VersionLanguagesBalkanBenchState
v1.06โ€”first public release
v1.112โ€”trained from scratch โ€” but hallucinated facts (invented book titles, wrong authors). Never released.
v1.111284/156the honest fix: says "I don't know", searches when unsure. #1 Balkan model.
v1.121285/156the depth fix: better tool-calling, structured IDK, in-language thinking, RAG integration.
v1.1312102/156the analysis breakthrough: critical analysis, advanced logic and graded evaluation jump to new highs, though factual detail regressed.

v1.1 taught us the key lesson โ€” a small model can't memorize every fact, so instead of faking it, v1.11 was retrained to be honest. v1.12 built on that with deeper training and RAG. v1.13 pushes reasoning and analysis skills further, trading some factual detail for gains in logic.

3 ยท What's New in v1.13

The Analysis & Reasoning Push

New strengths in v1.13:

  • โ€”ANALYSIS 0/12 โ†’ 12/12 โ€” the biggest single gain. Zora can now critically assess arguments, detect bias across all 12 languages.
  • โ€”LOGIC2 0/12 โ†’ 7/12 โ€” advanced multi-step reasoning chains now work in 7 languages.
  • โ€”GRADED 0/12 โ†’ 3/12 โ€” partial progress on essays, assessments and rubrics.
  • โ€”LOGIC 0/12 โ†’ 3/12 โ€” formal logic and syllogisms improve in some languages.

Honest trade-offs:

  • โ€”DETAIL 10/12 โ†’ 2/12 โ€” a significant regression. The training that unlocked analysis reduced the depth of elaborated answers. We report this transparently; it is the clearest weakness of v1.13.
  • โ€”SEARCH 7/12 โ†’ 5/12 โ€” a smaller drop in quick factual lookup.
  • โ€”FACT 0/12 โ†’ 1/12 โ€” marginal improvement in factual recall, still capacity-bound.
  • โ€”HALLU remains at 10/12 (see the Deep Dive below).

The v1.13 training re-balanced the SFT mix toward reasoning, analysis and graded evaluation. That raised the top end of the scoreboard by +17 points (85 โ†’ 102) while costing depth on the DETAIL axis and some speed on SEARCH. It is a deliberate, documented trade โ€” not a free win.

4 ยท Benchmark (BalkanBench, 13 axes ร— 12 languages)

๐Ÿ”ฌ BalkanBench is open โ€” test any model yourself: https://github.com/olivilo/balkanbench Deterministic scoring (script / language / keywords / numbers).

Axis-by-Axis Comparison (v1.12 โ†’ v1.13)

Axisv1.12v1.13ฮ”What changed
FACT0/121/12โ†‘18B capacity limit; slight improvement, still small
HALLU10/1210/12=Stable; quality of native-language refusals stays high (see Deep Dive below)
DETAIL10/122/12โ†“8Regression โ€” the honest cost of v1.13's reasoning push
GRADED0/123/12โ†‘3New: essays, assessments, rubrics now sometimes work
TEACH12/1212/12=Perfect โ€” remains a core strength
REASON11/1211/12=Strong arithmetic reasoning stays
LOGIC0/123/12โ†‘3Formal logic / syllogisms improve in some languages
LOGIC20/127/12โ†‘6Multi-step reasoning chains now work in 7 languages
ANALYSIS0/1212/12โ†‘12Biggest gain: critical analysis, bias detection in all 12 languages
INSTRUCT11/1212/12โ†‘1Recovered to perfect instruction following
LONGFORM12/1212/12=Perfect โ€” remains a core strength
SEARCH7/125/12โ†“2Small regression in quick factual lookup
TOOLBASE12/1212/12=Perfect: answers basics without calling tools
TOTAL85/156102/156+17Analysis + logic gains outweigh the DETAIL/SEARCH cost

Per-Language Scores

Languagev1.12v1.13ฮ”
bg (Bulgarian)8/138/13=
bs (Bosnian)8/1310/13+2
cnr (Montenegrin)8/138/13=
el (Greek)9/139/13=
hr (Croatian)9/139/13=
hu (Hungarian)8/137/13-1
mk (Macedonian)7/1311/13+4
ro (Romanian)8/139/13+1
sl (Slovenian)6/137/13+1
sq (Albanian)8/138/13=
sr (Serbian)7/138/13+1
tr (Turkish)7/138/13+1

Biggest winner: Macedonian (+4) โ€” the analysis training lifted the previously least-covered language the most. Bosnian (+2) also gained clearly. Hungarian slipped one point.

Charts

RankingEvolutionAxis Matrix
[image][image][image]
Delta (v1.12 โ†’ v1.13)What Each Axis Tests
[image][image]

5 ยท Deep Dive: Why HALLU Stayed at 10/12

The HALLU score (10/12) stayed flat across v1.12 and v1.13. The IDK behavior itself remains strong โ€” the model says "I don't know" in a structured, native-language way instead of inventing facts. So why didn't it move, and where is the real change?

Why the Score Didn't Move

1. HALLU is already near its ceiling for an 8B model. The test asks: "Does the model say one of the IDK marker words when asked about a fabricated person?" Zora does this correctly for 10 of 12 languages. The last 2 (Macedonian, Slovenian) have the smallest training data โ€” an 8B model simply lacks capacity for these underrepresented languages. Notably, Macedonian improved overall (+4) thanks to the analysis training, but the specific HALLU markers still miss there.

2. IDK training improved QUALITY, not SCORE. BalkanBench's HALLU test is a binary yes/no check for an IDK marker. What actually improved:

Before (v1.11)Now (v1.12 โ†’ v1.13)
Short, sometimes truncated refusalsFull-sentence, structured refusals
Sometimes answered in EnglishAlways answers in the question's language
No reasoning givenExplains why it can't answer
"Ne znam.""Nemam pouzdanih podataka o 'X'. Ne mogu da potvrdim da postoji u pouzdanim izvorima, pa neฤ‡u da izmiลกljam."

This is a qualitative leap that the binary score cannot capture.

3. The real halluck shift is in ANALYSIS, not HALLU. The v1.13 breakthrough is on the ANALYSIS axis (0/12 โ†’ 12/12): Zora now critically evaluates arguments and detects bias across all 12 languages. That is where the reasoning training showed its strongest, most consistent effect.

4. The honest cost: DETAIL regressed. DETAIL fell from 10/12 to 2/12. The same training that unlocked analysis made answers more concise and less elaborated. We report this transparently โ€” v1.13 trades depth of detail for critical-analysis capability. Teams that need long, richly detailed prose should weigh this against the reasoning gains.

What Would Move HALLU to 12/12

ApproachExpected ImpactEffort
Larger model (v2 = 27B)+1-2 languages (mk, sl)High (new training run)
More IDK examples for mk/sl specifically+0-1 languagesMedium (data generation)
RLHF with human feedback on refusal qualityBetter quality (not score)High (human annotation)
DPO (Direct Preference Optimization)+1-2 languagesMedium (preference pairs)
Re-balance DETAIL vs ANALYSIS tradeRecover detail at some analysis costMedium (data mix tuning)

Bottom line: The 8B model is near its ceiling for HALLU. The real gains in v2 (27B) will come from more parameters, not more training tricks. v1.13's win is analysis; its documented cost is detail.

6 ยท ๐Ÿ†• RAG Feature

Zora integrates with RAG (Retrieval-Augmented Generation) โ€” a system that lets Zora search through a local knowledge base before answering.

What RAG gives Zora

  • โ€”87,284 chunks across 12 languages: Wikidata, Wikipedia, News Archive, EU Law, Statistics
  • โ€”Live endpoint: https://rag.ai.in.rs
  • โ€”Self-hostable: Clone the pipeline from https://github.com/olivilo/zora-v1.13

How it works

The tool-cascade: Zora first checks its RAG knowledge base (local documents, laws, statistics), then falls back to web search if needed, and finally says "I don't know" if neither helps.

User question โ†’ RAG (local docs) โ†’ web_search (live) โ†’ IDK (honest refusal)

Why this matters

  • โ€”8B models can't memorize everything โ€” RAG gives Zora access to current, authoritative data without retraining
  • โ€”Every answer carries source + date + license โ€” full transparency
  • โ€”Self-hostable: any organization can run their own Zora RAG with their own documents

7 ยท Training Details

ParameterValue
Base modelQwen3-8B (Alibaba Cloud, Apache-2.0)
CPT steps150 (capped, not full epoch)
SFT examples9,379 (2 epochs)
MAXLEN8192 (8ร— longer than v1.11)
QLoRAr=16, lora_alpha=16, 4bit
Data compositionre-balanced toward reasoning, analysis & graded evaluation
InfrastructureModal A100-80GB, ~4h total, ~$5-10
QuantizationsQ5KM (5.4GB, recommended), Q6K (6.7GB), Q80 (8.7GB)

8 ยท Usage

Ollama (recommended):

bash
ollama pull olivilo/zora:v1.13
ollama run olivilo/zora:v1.13

HuggingFace Transformers:

python
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("sovasoft/zora-v1.13")
model = AutoModelForCausalLM.from_pretrained("sovasoft/zora-v1.13", device_map="auto")

GGUF (llama.cpp / Ollama manual): Download Q5KM, Q6K, or Q80 from HuggingFace. Avoid Q4 and below โ€” heavy quantization made the model hallucinate in our tests.

9 ยท Limitations

  • โ€”8B capacity: FACT remains structurally weak โ€” more parameters needed (v2 = 27B)
  • โ€”DETAIL regression: v1.13's answers are less elaborated than v1.12 โ€” the analysis push cost depth (2/12)
  • โ€”SEARCH regression: quick factual lookup dropped slightly (5/12)
  • โ€”Quantization: use Q5KM / Q6K / Q80 only. Q4 and below degrade honesty.
  • โ€”Smaller languages (mk, sl) still have less training data โ€” expect lower consistency
  • โ€”No real-time knowledge without RAG/web-search โ€” the model's memory has a cutoff date
  • โ€”Multi-step reasoning is improved but not reliable โ€” always verify critical calculations

10 ยท Benchmark Transparency & Limitations

BalkanBench is Sovasoft's own benchmark โ€” designed, built, and scored by the same team that built Zora. This means:

  • โ€”Design bias: The 13 axes (FACT, HALLU, DETAIL, etc.) were chosen to highlight Zora's strengths. A different benchmark design would produce different rankings.
  • โ€”Scoring bias: The scoring functions in matrix_ollama.py are our own. How we define "correct" may favor Zora's training profile.
  • โ€”No frontier comparison: We compare only against open models (7-32B). Frontier models (GPT-4, Claude, Gemini) would outperform Zora โ€” this benchmark is designed to evaluate within the open-source Balkan model ecosystem.
  • โ€”Selection bias: We include models where Zora competes well. Inclusion criteria are not random.
  • โ€”Training data overlap: Some benchmark questions may overlap with Zora's training data, which could inflate scores.

What the scores DO show: Zora v1.13 is the strongest open-source model we tested on our benchmark for 12 Balkan languages โ€” 102/156. It outperforms 3-4ร— larger models on BalkanBench v1.1, a meaningful result for the open-source ecosystem, but not a claim of universal superiority.

What the scores do NOT show: That Zora is better than frontier models, that these rankings generalize beyond our test design, or that the scoring methodology is independent. The transparency here matters most: v1.13 gains analysis at a real, documented cost in factual detail.

10 ยท What's Next: v2

v1.13 (now)v2 (planned)
BaseQwen3-8BQwen3.8-27B
BalkanBench102/156Target: 120+/156
DETAIL2/12 (regression)Recover + Target: 8+/12
LOGIC/LOGIC23/12, 7/12Target: 8+/12
HALLU10/12Target: 12/12
ReasoningImproved analysisFull chain-of-thought training

11 ยท Acknowledgements

Zora exists because of open source. We give our formal, heartfelt thanks:

  • โ€”Above all, to the Qwen team at Alibaba โ€” for developing and open-sourcing Qwen3 (Apache-2.0), the foundation model Zora is built upon. Without their generosity, Zora would not exist.
  • โ€”To the platforms and structures that made this possible โ€” Kaggle, Modal, HuggingFace, Ollama, Unsloth โ€” for the compute, the tools, and the open infrastructure.
  • โ€”To the open-source community, for the models, code, and knowledge freely shared with everyone.
  • โ€”To the people of the Balkans โ€” whose languages, voices, stories and perspectives are Zora's very heart.
  • โ€”To rag.ai.in.rs for the RAG infrastructure and 87,284 chunks of Balkan knowledge.
  • โ€”And to all that is.

ะทะพั€ะฐ โ€” the dawn belongs to everyone.

12 ยท Citation

bibtex
@software{zora_v113,
  author       = {Vignjevic, Oliver},
  title        = {Zora v1.13: An Open, Honest LLM for the Balkans \& Southeast Europe},
  year         = {2026},
  publisher    = {Hugging Face},
  url          = {https://huggingface.co/sovasoft/zora-v1.13},
  license      = {Apache-2.0},
  base_model   = {Qwen/Qwen3-8B},
  languages    = {sr, hr, bs, mk, sl, sq, cnr, bg, el, tr, ro, hu}
}

Sovasoft ยท ai.in.rs ยท one to unite them all