datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bioleaflets-biomedical-ner
Dataset Card for BioLeaflets Dataset
Dataset Summary
BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website.
Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately.
This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.RusLang-edu-1000
RusLang-Edu-1000 — an educational Russian-language QA dataset
RusLang-Edu-1000 is an expert-curated dataset of 1,000 instruction-format records ("question — detailed educational answer") covering the Russian language and linguistics: from phonetics and orthography to dialectology and theoretical linguistics. Every record contains a detailed answer (on average ≈1,100 characters), a short reference answer, a concise statement of the rule, and rich annotation (subject area, task… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/RusLang-edu-1000.RusLang-Edu-100
Russian Linguistics & Grammar Instruction Dataset (v2.0)
A curated and academically verified instruction dataset for evaluation (Eval/Benchmarking), Supervised Fine-Tuning (SFT), and Alignment (RLHF / DPO) of Large Language Models (LLMs) on Russian grammar, orthography, punctuation, morphology, syntax, stylistics, and linguistic analysis.
📌 Key Highlights
Language: Standard Russian (ru)
Volume: 100 expert-curated and linguistically verified instruction cards… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/RusLang-Edu-100.
