premmm/nepali-bench
NepaliBench ๐๏ธ A rigorous evaluation benchmark for Nepali language models. Why this exists There is no standard, publicly reproducible benchmark for evaluating Nepali LLMs. This dataset was created after a systematic evaluation of himalaya-ai's NanochatGPT and Gemma fine-tune revealed that models claiming Nepali capability had no shared evaluation standard to measure against. Dataset 100 carefully curated evaluation examples across 8 categories:โฆ See the full description on the dataset page: https://huggingface.co/datasets/premmm/nepali-bench.
137
Upload README.md with huggingface_hub
Upload nepali_bench.json with huggingface_hub
initial commit
