CoolFace
Modelpublic

overthelex/qwen2.5-1.5b-edrsr-legal-uk

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes28downloads
Model Card

Qwen2.5-1.5B-EDRSR-Legal-UK

Ukrainian legal domain model obtained by continued pretraining (CPT) of Qwen/Qwen2.5-1.5B on the EDRSR corpus of Ukrainian court decisions.

Part of a scaling experiment (0.5B / 1.5B / 3B / 14B) for the PhD dissertation at Glushkov Institute of Cybernetics, NAS of Ukraine.

Training Data

  • —Corpus: Unified State Register of Court Decisions of Ukraine (EDRSR)
  • —Documents: 33.9M court decisions (after dedup + quality filtering from 38.5M)
  • —Tokens: 161.4B tokens (Qwen2 BPE tokenizer, fertility = 0.515 for Ukrainian legal text)
  • —Sequence length: 8,192 tokens
  • —Shards: 1,233 pre-packaged numpy shards

Training Details

  • —Hardware: 8x NVIDIA H100 SXM 80GB (NVIDIA Innovation Lab via Brev)
  • —Framework: HuggingFace Trainer + DeepSpeed ZeRO-3
  • —Precision: bfloat16
  • —Global batch size: 128 sequences (1.05M tokens/step)
  • —Total steps: 9,536 (10B tokens processed)
  • —Training time: 17.8 hours
  • —Throughput: 140K tokens/sec, 6.7 sec/step

Results

MetricValue
Initial loss (step 10)1.2216
Final loss (step 9,536)0.2391
Loss reduction-80%
Base perplexity4.61
CPT perplexity1.31
Perplexity reduction-71.5%

Scaling Law

All four models in the series converge to similar perplexity after CPT:

ModelBase PPLCPT PPLReduction
0.5B6.831.35-80%
1.5B4.611.31-72%
3B3.831.30-66%
14B2.841.28-55%

Intended Use

This is a base model (not instruction-tuned). It is intended for:

  • —Research on domain adaptation of LLMs for low-resource legal languages
  • —Downstream fine-tuning for Ukrainian legal NLP tasks
  • —Scaling law analysis of continued pretraining
  • —Perplexity evaluation on Ukrainian legal text

Limitations

  • —Not instruction-tuned; will not follow instructions or chat
  • —Trained on Ukrainian court decisions only; may not generalize to other legal systems

Related Resources