kamaalg/azerbaijani-eval-benchmarks
Azerbaijani evaluation benchmarks (v0) Small, reproducible Azerbaijani benchmarks for evaluating base language models, part of an open Azerbaijani LLM stack. Built by build_benchmarks.py (rerun to regenerate deterministically). file task items format mmlu_az.jsonl multiple-choice knowledge 102 {question, choices[4], answer, subject} ner_az.jsonl named-entity recognition (BIO) 42 {tokens[], tags[]} mmlu_az.jsonl MMLU-style 4-way multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-eval-benchmarks.
This repository belongs to kamaalg on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
