overthelex/qwen2.5-0.5b-edrsr-legal-uk
014
Qwen2.5-0.5B-EDRSR-Legal-UK
Ukrainian legal domain model obtained by continued pretraining (CPT) of Qwen2.5-0.5B on the EDRSR corpus of Ukrainian court decisions.
Part of a scaling experiment (0.5B / 1.5B / 3B) for the PhD dissertation at Glushkov Institute of Cybernetics, NAS of Ukraine.
Training Data
- Corpus: Unified State Register of Court Decisions of Ukraine (EDRSR)
- Documents: 33.9M court decisions (after dedup + quality filtering from 38.5M)
- Tokens: 161.4B tokens (Qwen2 BPE tokenizer, fertility = 0.515 for Ukrainian legal text)
- Sequence length: 8,192 tokens
- Shards: 1,233 pre-packaged numpy shards
Training Details
- Hardware: 8x NVIDIA H100 SXM 80GB (NVIDIA Innovation Lab via Brev)
- Framework: HuggingFace Trainer + DeepSpeed ZeRO-3
- Precision: bfloat16
- Global batch size: 128 sequences (microbatch=4, gradaccum=4, 8 GPUs)
- Tokens per step: 1.05M
- Total steps: 9,536 (10B tokens processed)
- Learning rate: 2e-4, cosine schedule, 300-step linear warmup
- Training time: 8.4 hours
- Throughput: 262K tokens/sec, 3.2 sec/step
Results
Intended Use
This is a base model (not instruction-tuned). It is intended for:
- Research on domain adaptation of LLMs for low-resource legal languages
- Downstream fine-tuning for Ukrainian legal NLP tasks
- Scaling law analysis of continued pretraining
- Perplexity evaluation on Ukrainian legal text
Limitations
- Not instruction-tuned; will not follow instructions or chat
- Trained on Ukrainian court decisions only; may not generalize to other legal systems
- Small model (0.5B) -- limited reasoning capacity compared to larger variants
