datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SerbianEmailsNER
SerbianEmailsNER Dataset
An LLM-generated synthetic dataset comprising of emails in Serbian language and corresponding NER annotations. The primary purpose of the dataset is to be used in evaluation of NER models and anonymization software for Serbian language.
📝 Summary
This dataset contains 300 synthetically generated emails written in both Latin and Cyrillic scripts, evenly split across four real-world correspondence types:
private-to-private
private-to-business… See the full description on the dataset page: https://huggingface.co/datasets/goranagojic/SerbianEmailsNER.airoboros-3.0-serbian
airoboros-3.0-serbian
This dataset is a translation of the airoboros-3.0 datasets to Serbian Latin.
NOTE:I used various online translation APIs, so the quality of translations isn't perfect yet. However, I will try to refine them over time with the help of automated scripts and LLMs.
Huge thanks to Jondurbin (@jon_durbin) for creating the original dataset as well as the tools for creating it: https://twitter.com/jon_durbin.
Original dataset link:… See the full description on the dataset page: https://huggingface.co/datasets/draganjovanovich/airoboros-3.0-serbian.alpaca-cleaned-serbian-full
Serbian Alpaca Cleaned Dataset
Original Repository: https://github.com/gururise/AlpacaDataCleaned
Original HF Repository: https://huggingface.co/datasets/yahma/alpaca-cleaned
Dataset Description
This is a serbian cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the… See the full description on the dataset page: https://huggingface.co/datasets/datatab/alpaca-cleaned-serbian-full.uner_llm_inst_serbian
Dataset Card for Universal NER v1 in the Aya format - Serbian subset
This dataset is a format conversion for the Serbian data in the original Universal NER v1 into the Aya instruction format and it's released here under the same CC-BY-SA 4.0 license and conditions.
The dataset contains different subsets and their dev/test/train splits, depending on language. For more details, please refer to:
Dataset Details
For the original Universal NER dataset v1 and more details… See the full description on the dataset page: https://huggingface.co/datasets/universalner/uner_llm_inst_serbian.alpaca-cleaned-serbian-10000wiki-serbian-croatian
Wiki Serbian-Croatian
A cleaned Wikipedia corpus combining Serbian (sr) and Croatian (hr) Wikipedia articles,
with Croatian text transliterated to Cyrillic script.
Processing
Removed wiki markup, infoboxes, tables, templates
Removed calendar stub articles
Filtered articles with >40% Latin characters
Croatian Latin transliterated to Serbian Cyrillic
License
Source data: CC BY-SA 4.0 (Wikimedia Foundation)
Corpus compilation: CC BY 4.0 — Alexei… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/wiki-serbian-croatian.alpaca-cleaned-serbian-30000
