datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bioleaflets-biomedical-ner
Dataset Card for BioLeaflets Dataset
Dataset Summary
BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website.
Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately.
This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.hotel-multimodalRusLangNet
RusLangNet
tags: Natural Language Processing, Russian Text, Network Analysis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'RusLangNet' dataset is a curated collection of Russian text data extracted from various online platforms, with the aim of facilitating research in natural language processing (NLP) and network analysis. The texts are sourced from social media, news articles, forums, and other websites that have… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/RusLangNet.
