serbian
Datasets
All datasets matching “serbian”serbian-llm-benchmark
Serbian LLM Evaluation Dataset
Welcome to the Serbian LLM Evaluation Dataset, your one-stop solution for evaluating Serbian Language Models (LLMs) like never before! This comprehensive toolkit empowers you to measure model performance across diverse domains in Serbian, ensuring your models are smarter, faster, and more intuitive. Whether you're a researcher, developer, or just an enthusiast—this dataset is tailor-made to help your LLM thrive.
🔍 What's Inside?
This… See the full description on the dataset page: https://huggingface.co/datasets/datatab/serbian-llm-benchmark.serbian-llm-eval-v1
Serbian LLM eval 🇷🇸
This dataset should be used for Serbian (and potentially also other HBS languages) LLM evaluation.
Here is the GitHub project used to build this dataset.
For technical report of the project see this in-depth Weights & Biases report. ❤️
I'll give a TL;DR here:
What is covered?
Common sense reasoning:
Hellaswag, Winogrande, PIQA, OpenbookQA, ARC-Easy, ARC-Challenge
World knowledge:
NaturalQuestions, TriviaQA
Reading comprehension:
BoolQ… See the full description on the dataset page: https://huggingface.co/datasets/gordicaleksa/serbian-llm-eval-v1.serbian-llm-eval-v0
Serbian LLM eval v0 🇷🇸
Please instead use the version 1 of the dataset here.
Weights & Biases report.
Project Sponsors
Platinum sponsors 🌟
Ivan (fizicko lice, anoniman)
Gold sponsors 🟡
qq (fizicko lice, anoniman)
Mitar Perovic
Nikola Ivancevic
Silver sponsors ⚪
psk.rs, OmniStreak, Marko Radojicic, Luka Vazic, Miloš Durković, Marjan Radeski, Marjan Stankovic (fizicko lice), Nikola Stojiljkovic, Mihailo Tomic, Bojan Jevtic, Jelena… See the full description on the dataset page: https://huggingface.co/datasets/gordicaleksa/serbian-llm-eval-v0.SerbianEmailsNER
SerbianEmailsNER Dataset
An LLM-generated synthetic dataset comprising of emails in Serbian language and corresponding NER annotations. The primary purpose of the dataset is to be used in evaluation of NER models and anonymization software for Serbian language.
📝 Summary
This dataset contains 300 synthetically generated emails written in both Latin and Cyrillic scripts, evenly split across four real-world correspondence types:
private-to-private
private-to-business… See the full description on the dataset page: https://huggingface.co/datasets/goranagojic/SerbianEmailsNER.NanoBEIR-srSerbian-PD
🇷🇸 Serbian Public Domain 🇷🇸
Serbian-Public Domain or Serbian-PD is a large collection aiming to aggregate all Serbian monographies and periodicals in the public domain. As of March 2024, it is the biggest Serbian open corpus.
Dataset summary
The collection contains 1,405 titles making up 156,712,807 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file has the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Serbian-PD.
