datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
serbian-llm-benchmark
Serbian LLM Evaluation Dataset
Welcome to the Serbian LLM Evaluation Dataset, your one-stop solution for evaluating Serbian Language Models (LLMs) like never before! This comprehensive toolkit empowers you to measure model performance across diverse domains in Serbian, ensuring your models are smarter, faster, and more intuitive. Whether you're a researcher, developer, or just an enthusiast—this dataset is tailor-made to help your LLM thrive.
🔍 What's Inside?
This… See the full description on the dataset page: https://huggingface.co/datasets/datatab/serbian-llm-benchmark.ultrafeedback_binarized_serbian
Dataset Card for UltraFeedback Binarized Serbian
Dataset Description
This dataset is a Serbian-translated version of the UltraFeedback dataset, utilized for training Zephyr-7Β-β. The original dataset comprises 64k English-language prompts, each paired with four completions from various models. In this Serbian version, the prompts and completions have been translated into Serbian. The dataset creation process remains the same: selecting the completion with the highest… See the full description on the dataset page: https://huggingface.co/datasets/datatab/ultrafeedback_binarized_serbian.guanaco-sharegpt-style-serbian
Guanaco Sharegpt-style Serbian
Dataset Description
This dataset is a Serbian-translated version of the philschmid/guanaco-sharegpt-style
Dataset Structure
Usage
To load the dataset in Serbian, run:
from datasets import load_dataset
ds = load_dataset("datatab/guanaco-sharegpt-style-serbian")
Data Splits
The dataset has one splits, suitable for:
Supervised fine-tuning (sft).
The dataset is stored in parquet format with each entry using… See the full description on the dataset page: https://huggingface.co/datasets/datatab/guanaco-sharegpt-style-serbian.serbian_qa
Dataset Card for "serbian_qa"
Dataset Summary
The "serbian_qa" dataset is a collection of context-query pairs in Serbian. It is designed for question-answering tasks and contains contexts from various Serbian language sources, paired with automatically generated queries of different lengths.
Supported Tasks and Leaderboards
Tasks: Question Answering, Information Retrieval
Languages
The dataset is in Serbian (sr).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/serbian_qa.
