datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ultrafeedback_binarized_serbian
Dataset Card for UltraFeedback Binarized Serbian
Dataset Description
This dataset is a Serbian-translated version of the UltraFeedback dataset, utilized for training Zephyr-7Β-β. The original dataset comprises 64k English-language prompts, each paired with four completions from various models. In this Serbian version, the prompts and completions have been translated into Serbian. The dataset creation process remains the same: selecting the completion with the highest… See the full description on the dataset page: https://huggingface.co/datasets/datatab/ultrafeedback_binarized_serbian.airoboros-3.0-serbian
airoboros-3.0-serbian
This dataset is a translation of the airoboros-3.0 datasets to Serbian Latin.
NOTE:I used various online translation APIs, so the quality of translations isn't perfect yet. However, I will try to refine them over time with the help of automated scripts and LLMs.
Huge thanks to Jondurbin (@jon_durbin) for creating the original dataset as well as the tools for creating it: https://twitter.com/jon_durbin.
Original dataset link:… See the full description on the dataset page: https://huggingface.co/datasets/draganjovanovich/airoboros-3.0-serbian.alpaca-cleaned-serbian-full
Serbian Alpaca Cleaned Dataset
Original Repository: https://github.com/gururise/AlpacaDataCleaned
Original HF Repository: https://huggingface.co/datasets/yahma/alpaca-cleaned
Dataset Description
This is a serbian cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the… See the full description on the dataset page: https://huggingface.co/datasets/datatab/alpaca-cleaned-serbian-full.russian_assistant_to_serbian
Dataset Card for Russian to Serbian Assistant
Dataset Description
Dataset Summary
Russian to Serbian Assistant је едукативни dataset намењен српским говорницима који уче руски језик. Dataset садржи руске фразе са фонетском транскрипцијом прилагођеном српском језику, превод на српски, и детаљним граматичким правилима за правилну изговор.
Supported Tasks
Учење руског језика: Помоћ српским говорницима у савладавању руског језика
Фонетска транскрипција:… See the full description on the dataset page: https://huggingface.co/datasets/mkrstic8/russian_assistant_to_serbian.guanaco-sharegpt-style-serbian
Guanaco Sharegpt-style Serbian
Dataset Description
This dataset is a Serbian-translated version of the philschmid/guanaco-sharegpt-style
Dataset Structure
Usage
To load the dataset in Serbian, run:
from datasets import load_dataset
ds = load_dataset("datatab/guanaco-sharegpt-style-serbian")
Data Splits
The dataset has one splits, suitable for:
Supervised fine-tuning (sft).
The dataset is stored in parquet format with each entry using… See the full description on the dataset page: https://huggingface.co/datasets/datatab/guanaco-sharegpt-style-serbian.orca_math_world_problem_200k_serbian
Kartica Podataka
Ovaj skup podataka sadrži ~200K tekstualnih matematičkih zadataka za osnovnu školu. Svi odgovori u ovom skupu podataka su generisani korišćenjem Azure GPT4-Turbo. Molimo vas da pogledate Orca-Math: Unlocking the potential of SLMs in Grade School Math za detalje o konstrukciji skupa podataka.
Opis Skupa Podataka
Kreirao: Microsoft
Jezik(i) (NLP): Engleski
Licenca: MIT
Izvori Skupa Podataka
Repozitorijum:… See the full description on the dataset page: https://huggingface.co/datasets/datatab/orca_math_world_problem_200k_serbian.wiki-serbian-croatian
Wiki Serbian-Croatian
A cleaned Wikipedia corpus combining Serbian (sr) and Croatian (hr) Wikipedia articles,
with Croatian text transliterated to Cyrillic script.
Processing
Removed wiki markup, infoboxes, tables, templates
Removed calendar stub articles
Filtered articles with >40% Latin characters
Croatian Latin transliterated to Serbian Cyrillic
License
Source data: CC BY-SA 4.0 (Wikimedia Foundation)
Corpus compilation: CC BY 4.0 — Alexei… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/wiki-serbian-croatian.Roleplay-Serbian
RolePlay-Serbian
Roleplay-Serbian Dataset is a dataset for roleplaying in the Serbian language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this github… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Serbian.
