CoolFace
Modelpublic

overthelex/qwen2.5-0.5b-edrsr-legal-uk

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes14downloads
Model Card

Qwen2.5-0.5B-EDRSR-Legal-UK

Ukrainian legal domain model obtained by continued pretraining (CPT) of Qwen2.5-0.5B on the EDRSR corpus of Ukrainian court decisions.

Part of a scaling experiment (0.5B / 1.5B / 3B) for the PhD dissertation at Glushkov Institute of Cybernetics, NAS of Ukraine.

Training Data

  • —Corpus: Unified State Register of Court Decisions of Ukraine (EDRSR)
  • —Documents: 33.9M court decisions (after dedup + quality filtering from 38.5M)
  • —Tokens: 161.4B tokens (Qwen2 BPE tokenizer, fertility = 0.515 for Ukrainian legal text)
  • —Sequence length: 8,192 tokens
  • —Shards: 1,233 pre-packaged numpy shards

Training Details

  • —Hardware: 8x NVIDIA H100 SXM 80GB (NVIDIA Innovation Lab via Brev)
  • —Framework: HuggingFace Trainer + DeepSpeed ZeRO-3
  • —Precision: bfloat16
  • —Global batch size: 128 sequences (microbatch=4, gradaccum=4, 8 GPUs)
  • —Tokens per step: 1.05M
  • —Total steps: 9,536 (10B tokens processed)
  • —Learning rate: 2e-4, cosine schedule, 300-step linear warmup
  • —Training time: 8.4 hours
  • —Throughput: 262K tokens/sec, 3.2 sec/step

Results

MetricValue
Initial loss (step 10)1.5254
Final loss (step 9,536)0.2633
Loss reduction-83%
Train loss (avg)0.3117
Total FLOPs2.65e18

Intended Use

This is a base model (not instruction-tuned). It is intended for:

  • —Research on domain adaptation of LLMs for low-resource legal languages
  • —Downstream fine-tuning for Ukrainian legal NLP tasks
  • —Scaling law analysis of continued pretraining
  • —Perplexity evaluation on Ukrainian legal text

Limitations

  • —Not instruction-tuned; will not follow instructions or chat
  • —Trained on Ukrainian court decisions only; may not generalize to other legal systems
  • —Small model (0.5B) -- limited reasoning capacity compared to larger variants

Related Resources